# What is fail-open authentication?

Fail-open authentication is an access check that lets a request through when the check itself errors or times out, instead of denying the request.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Define fail-open and fail-closed behavior
-   Explain how auth code fails open by accident
-   Identify tests that catch a fail-open check

## Related content

-   [How to test a command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications)
-   [How to run a two-account access test](https://specstory.com/learning/cli-and-web/two-account-access-test)
-   [What are exit codes?](https://specstory.com/learning/cli-and-web/exit-codes)

## What is fail-open authentication?

Fail-open authentication is a security flaw in which a login check that errors or times out allows the request instead of refusing it. The opposite behavior is failing closed, also called failing secure, which refuses the request. A fail-open check works normally while its dependencies work, so the flaw shows only when something breaks.

MITRE's Common Weakness Enumeration lists the flaw as [CWE-636](https://cwe.mitre.org/data/definitions/636.html), "Not Failing Securely," a weakness in which a product falls back to a less secure state after an error. The flaw can sit in any code that makes an access decision, e.g. the API server behind a [command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications), which decides whether the tool's user may export orders.

[METR disclosed](https://metr.org/blog/2026-08-31-security-update/) that a dashboard whose authentication silently failed open let an attacker reach an AI agent behind the login. In METR's account, the attacker prompted that agent for a model API key and ran up about $600,000 in model usage over three weeks. METR's post says the app was built by [vibe coding](https://specstory.com/learning/ai-coding/vibe-coding) and does not say how the check failed open.

## What causes fail-open authentication?

An access check fails open when its error path ends without a refusal. Many web frameworks let a check refuse a request by raising an error and allow it by returning. A request moves through the check in four steps:

1.  The request reaches the access check, e.g. middleware before each admin route.
2.  The check calls a login service to validate the token.
3.  The call fails, e.g. the login service does not answer within its timeout.
4.  The error branch runs. Code that fails closed refuses the request. Code that fails open returns, so the request reaches the route's handler.

Diagram: Where an access check fails open

The check reaches no decision when its call fails. What the error branch does next determines whether the request is refused or served.

MITRE says the weakness typically comes from a wish to keep a product working after an error. In code, it usually takes one of four forms:

-   **A catch-all exception handler.** Python's `except Exception` catches any error. If the handler logs it and returns, a timeout or a token parse error becomes an allowed request.
-   **A check that looks only for a refusal.** Code that refuses only on an explicit "no" allows any other result, e.g. `if result == "deny"` lets an error result through.
-   **A setting that turns the check off.** Code that skips the check when a setting is empty, e.g. a signing key, runs without authentication after a deploy that leaves it out.
-   **A fail mode set to allow.** Some proxies and login services let an administrator choose what happens when the checking service is unreachable. The [Envoy proxy's authorization filter](https://www.envoyproxy.io/docs/envoy/latest/api-v3/extensions/filters/http/ext_authz/v3/ext_authz.proto) accepts requests in that case when `failure_mode_allow` is `true`, and rejects them by default.

Authorization, the check of what a user may do, fails open for the same causes, e.g. a permission lookup that times out.

## What does a fail-open bug look like?

Here is an illustrative example. Acme Co. sells furniture online. Its command-line tool, `acme`, exports orders through an admin API, and the API asks a login service whether each token belongs to an admin. A developer asks a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) to "Stop the admin API from crashing when the login service is slow." The agent wraps the call in a handler:

```python
def require_admin(request):
    try:
        user = login_service.verify(request.token, timeout=2)
    except Exception as err:
        log.warning("token check failed: %s", err)
        return  # the request continues to the handler
    if user is None or not user.is_admin:
        raise Forbidden()
```

The `except` branch returns, so any error in the call allows the request. Four things follow:

1.  The agent's tests send an admin token, a customer token, and no token. The login service answers each time, so the tests pass.
2.  A week later, the login service hangs for 10 minutes under heavy load.
3.  In that window, anyone who can reach the API can export the orders. A developer at Acme shows it by running the export with no token.
4.  The call times out after 2 seconds, the `except` branch returns, and the API sends the orders.

The terminal and the API's log show the result:

```text
$ ACME_TOKEN="" acme export --format csv > orders.csv
$ echo $?
0
$ head -n 2 orders.csv
order_id,total
A-1042,$240.00

# API server log
WARNING token check failed: Read timed out. (read timeout=2)
INFO "GET /admin/orders.csv HTTP/1.1" 200
```

The [exit code](https://specstory.com/learning/cli-and-web/exit-codes) is 0, and only the server log shows that the token was never validated. This example is simplified. A real API would also need an alert when the login service hangs.

## What changes when a coding agent writes the code?

A coding agent can edit an access check while it fixes something else, e.g. a crash when a service is slow. A handler that catches the error stops the crash, and the agent's [happy path](https://specstory.com/learning/testing/happy-path-and-edge-cases) tests pass, because none of them makes the login service fail.

A prompt like Acme's does not say what should happen when a check has no answer, and the agent's code and tests come from the same prompt. That missing rule is a gap in [AI-generated code security](https://specstory.com/learning/verification/ai-generated-code-security). In the diff, a warning and a `return` read as careful error handling. Login checks that fail open are common [bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs) that a demo misses, because the app works until a dependency fails.

Add the rule to the prompt, e.g. "Deny the request when the access check errors." Before the agent starts, write a test that makes the check fail, and keep it in a file the agent does not edit.

## How do you fix or prevent fail-open authentication?

The fix is to make each path that does not reach an explicit allow end in a refusal. The Open Worldwide Application Security Project (OWASP) [Authorization Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html) calls this deny by default, and says to handle each exception and failed access check, however unlikely. These changes apply it:

-   **Refuse in the error branch.** In the Acme check, replace `return` in the `except` branch with `raise ServiceUnavailable()`, so an outage returns `503` instead of orders.
-   **Allow in one place.** End the check with one line that allows a request after a positive test, e.g. `if user and user.is_admin: return`, and raise on each other path.
-   **Stop at startup when a setting is missing.** A server that cannot load its signing key should exit with a nonzero code instead of running without the check. Continuing to run is the weakness MITRE lists as CWE-455, "Non-exit on Failed Initialization."
-   **Limit what a bypass reaches.** Keep [secrets](https://specstory.com/learning/environments/agent-secrets) out of reach of any agent behind a login, because a bypass hands that agent to the attacker, as METR's incident shows.

A test catches a fail-open check only when it makes the check fail. This pytest test makes the login call time out and expects a refusal:

```python
def test_export_is_refused_when_login_service_times_out(client, monkeypatch):
    def time_out(*args, **kwargs):
        raise TimeoutError("login service timed out")

    monkeypatch.setattr(login_service, "verify", time_out)
    res = client.get("/admin/orders.csv",
                     headers={"Authorization": "Bearer x"})
    assert res.status_code == 503
    assert b"A-1042" not in res.data
```

On the original code, the test fails with a `200`. Other tests cover the rest of the error path:

-   **Negative tests.** Send requests that should fail, e.g. one with a malformed token, because each kind of bad input can raise a different error.
-   **Outage tests.** Stop the login service in a test copy of the app and send requests with no token. At Acme, `acme export` should then exit with a nonzero code and write nothing.
-   **Scans.** A dynamic application security testing ([DAST](https://specstory.com/learning/glossary#dast)) scan probes a running app from outside and finds a fail-open check only if its requests make the check error.
-   **Static analysis.** [Static analysis](https://specstory.com/learning/code-review/static-analysis) can flag the handler, e.g. Pylint's `broad-exception-caught` rule warns on `except Exception` unless the handler raises. The rule checks for a `raise`, not for what the handler means.

A [two-account access test](https://specstory.com/learning/cli-and-web/two-account-access-test) misses this bug, because both accounts hold valid sessions while the login service answers.

## How is fail open different from fail closed?

Fail closed means that a check refuses a request when it has no answer, and fail open means that it allows the request. Both terms describe only the error path, because a working check gives the same answer either way.

The choice trades security for availability. During a login service outage, code that fails closed locks out every user, and code that fails open lets anyone in. Checking a signed token's signature locally lowers that cost without failing open, because fewer requests then depend on the login service. Failing open can suit a check whose wrong allow costs little, e.g. a rate limiter, but an access check should fail closed.

## FAQs

### When is failing open the right choice?

Failing open is the right choice only when a wrong allow costs less than an outage, e.g. a rate limiter. A login or permission check should fail closed, because a request let through by mistake can expose data or money.

### Can static analysis find fail-open code?

Static analysis can find some fail-open code by flagging handlers that catch any exception. A general rule sees the handler's code, not whether it allows or refuses the request, so a test that makes the check fail is still needed.

### What is a catch-all exception handler?

A catch-all exception handler is a block of code that catches any error, whatever its type. In an access check, a catch-all handler that logs the error and returns turns any failure, e.g. a timeout, into an allowed request.

### Does fail-open apply to authorization too?

Fail-open applies to authorization as well as authentication, because a permission lookup can time out the same way a login check can. A two-account access test misses it, since each lookup in that test succeeds. A test that makes the lookup fail catches it.

---

Source: [Fail-open authentication | Fail open vs. closed | SpecStory](https://specstory.com/learning/cli-and-web/fail-open-authentication)
