# What is a false pass (false green test)?

A false pass is a test or check that reports success while the software is broken, because the check missed the defect or the failure was lost.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define a false pass and a false negative
-   Explain the common causes of false passes
-   Identify checks that expose a false pass

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [Why do AI-written tests always pass?](https://specstory.com/learning/verification/ai-written-tests-always-pass)
-   [Why do coding agents say "done" when the code doesn't work?](https://specstory.com/learning/verification/coding-agent-done-claims)
-   [What is runtime verification?](https://specstory.com/learning/verification/runtime-verification)

## What is a false pass (false green test)?

A false pass is a test or check result that reports success on broken code, because the check missed the defect or lost the failure. It is also called a false green test, after the green status that a [continuous integration (CI) pipeline](https://specstory.com/learning/ci-cd/ci-cd) shows for a passing run. The tests pass, but the app is broken.

The International Software Testing Qualifications Board (ISTQB) calls this a [false negative result](https://glossary.istqb.org/en_US/term/false-negative-result) and lists "false-pass result" as a synonym. A false negative is a test result that misses a defect that does exist. Some developers say false positive for the same event, because they read a pass as the positive result. ISTQB uses false positive for the opposite error, a reported defect that does not exist.

A false pass hides the defect that the check should have caught, and a green run draws no attention. The change merges, and users find the defect. When a coding agent writes the code and its tests, a green run is often the main evidence behind its report that the work is "done." Checking what that run left out is part of how teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing).

## What causes a false pass?

A test run turns green when the runner finishes, finds no failed [assertion](https://specstory.com/learning/glossary#assertion), and returns [exit code](https://specstory.com/learning/cli-and-web/exit-codes) 0 to the pipeline. A false pass happens when one step in that chain reports success without doing its work. The usual causes are these:

-   **The test checks nothing useful.** The test runs code but has no assertion, or it is a [tautological test](https://specstory.com/learning/test-quality/tautological-test) whose expected value comes from the code itself. [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass) often have this problem.
-   **A substitute replaces the broken part.** A [mock](https://specstory.com/learning/testing/mocks-vs-stubs) or stub returns a canned answer where the real component fails, so the test checks the substitute instead.
-   **The test never ran.** A [skipped test](https://specstory.com/learning/glossary#skipped-test) is left out of the run, and the summary still counts the run as passed. A file that the runner never collects is left out with no mark in the summary.
-   **The test ended before its check.** An asynchronous test that does not await its promise can finish, and pass, before its assertion runs.
-   **The pipeline dropped the result.** A step that ends in `|| true` exits 0 whatever the tests return. A step that pipes test output into another command can report that command's exit code instead of the runner's. In GitHub Actions, `continue-on-error: true` lets the job pass when its test step fails.
-   **The test ran the wrong code.** A stale build or a cached artifact makes the tests check a version that the change did not touch.

Diagram: Where a false pass enters a test run

Each stage can report success without doing its work. The CI status shows only that no stage reported a failure.

[A 2026 study](https://arxiv.org/abs/2606.18168) of 86,156 changes to test files by five coding agents found that 80.2% had weak or no explicit oracle signals. Many of those tests ran code while checking little or nothing about its output.

## What does a false pass look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes the checkout code, writes two tests, and marks the failing cart test with `test.skip`. The pipeline runs the suite with a verbose reporter and prints this:

```text
 ✓ tests/checkout.test.js > customer can change the delivery address
 ↓ tests/checkout.test.js > cart keeps its items after an address change

 Test Files  1 passed (1)
      Tests  1 passed | 1 skipped (2)
```

The runner exits with code 0, so the pipeline turns green. The agent's summary says the change is "done." A customer then edits an address at checkout. The address saved, but the cart emptied.

The only test that checks the cart never ran, and the summary reports it as `1 skipped`. Removing `.skip` gives the true result on the same code:

```text
 FAIL  tests/checkout.test.js > cart keeps its items after an address change
AssertionError: expected [] to have a length of 2 but got +0
```

The code did not change between the two runs. Only the skip marker changed. This example is simplified. A real suite often prints hundreds of results, and a reader can miss one skipped test among them.

## What is a ghost feature?

A ghost feature is a feature that exists in the code, and may pass its tests, but never runs for users, because nothing in the app calls it. It is a false pass at the level of a whole feature.

Its unit tests import the function and call it directly, so they never go through the entry point users reach, e.g. a button. At Acme, a coding agent could write `saveDeliveryAddress()` and a passing test for it while the checkout page's Save button still calls the old handler. The run is green, and customers get no change.

The agent's summary then lists the function as finished. That report is a [false completion claim](https://specstory.com/learning/verification/coding-agent-done-claims), and a reader of the diff can miss it, because each file looks complete. A run of the app through its real entry points shows that the change has no effect. A search for the function's callers outside the tests finds none.

## How do you find or prevent a false pass?

Each of these checks closes one of the causes above:

-   **Make each test fail once.** Break the behavior a test covers, and confirm that the test goes red.
-   **Read the counts, not only the color.** Treat an unexplained skip or an empty run as a failure. Some runners already fail an empty run, e.g. pytest exits with code 5 when it collects no tests.
-   **Lint the test files.** A lint rule can flag a test that has no assertion, e.g. the `expect-expect` rule in the ESLint plugins for Jest and Vitest.
-   **Keep the exit code.** Remove `|| true` and `continue-on-error: true` from test steps, and set `pipefail` when test output goes through a pipe.
-   **Run mutation testing.** [Mutation testing](https://specstory.com/learning/test-quality/mutation-testing) makes small deliberate bugs, called mutants, and runs the tests against each one. A mutant that survives usually marks behavior that no test checks.
-   **Run the software itself.** A [smoke test](https://specstory.com/learning/testing/smoke-testing) or an end-to-end test drives the app through the entry points users reach, which can expose ghost features.

[Code coverage](https://specstory.com/learning/test-quality/code-coverage) shows which lines no test ran, but it cannot show whether a test checked a result. [Researchers have found](https://arxiv.org/abs/2506.02954) test suites with 100% code coverage that caught only 4% of deliberately injected bugs, which is why coverage alone is a weak signal.

Mutation testing is slow on a large codebase, and a run of the app covers only the paths it took. A run that stopped early or skipped tests is incomplete, not a pass.

## How is a false pass different from a flaky test?

A [flaky test](https://specstory.com/learning/debugging/flaky-tests) both passes and fails on the same code, so its result is noisy. A false pass is often consistent, and rerunning it returns the same green result while the code stays broken. The two overlap when a retry hides a real intermittent failure, e.g. a race condition that fails once and then passes. The green result after that retry is itself a false pass.

## How does RunStory help with false passes?

A green test run shows only what the tests checked. An agent saying "done" is a claim that needs to be verified. RunStory independently runs the software against your change and returns evidence to the coding agent. It runs the software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Why is CI green when the app is broken?

CI is green when each step in the pipeline exits with code 0. That status covers only the checks that ran and what they asserted. A skipped test or a test with no assertion leaves the pipeline green while the app is broken.

### Is a false pass the same as a false negative?

A false pass and a false negative name the same result, a test result that misses a defect that is present. The ISTQB glossary lists the two terms as synonyms. Some developers call the same result a false positive, a term that ISTQB uses for the opposite error.

### Can a skipped test cause a false pass?

A skipped test causes a false pass when it is the only test that checks the broken behavior. Test runners usually list skipped tests in the summary and still exit with code 0, so the pipeline turns green. A team can treat an unexplained skip as a failure instead.

### How does mutation testing find false passes?

Mutation testing finds false passes by breaking the code on purpose and running the tests again. A mutant that survives is usually a controlled false pass, because the tests report success on code that was changed to be wrong. The survivors show where a real bug could pass the same way.

---

Source: [What is a false pass test? | False green tests | SpecStory](https://specstory.com/learning/verification/false-pass)
