# What is a fail-before, pass-after check?

A fail-before, pass-after check runs a new test on old code, where it must fail, then on the changed code, where it must pass, to show it catches the bug.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define a fail-before, pass-after check
-   Explain why a test must fail on the old code
-   Apply the check to a bug fix or agent change

## Related content

-   [What is code coverage?](https://specstory.com/learning/test-quality/code-coverage)
-   [How to check that AI-written tests catch bugs](https://specstory.com/learning/test-quality/checking-agent-written-tests)
-   [What are catching tests (catching JiTTests)?](https://specstory.com/learning/test-quality/catching-tests)
-   [What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

## What is a fail-before, pass-after check?

A fail-before, pass-after check is a test of a test that must fail on the code before a change and pass on the changed code. The check runs the same test file on both versions. Such a test is often called a fail-to-pass test, and some developers call the check a red-green check.

The failing run is the evidence that the test reaches the bug. A test can pass on the changed code whether or not the bug is gone, e.g. when it checks a value the bug never touches. Such a test can run every changed line and raise [code coverage](https://specstory.com/learning/test-quality/code-coverage) while it checks nothing the bug affects. Its green result says nothing about the fix.

The International Software Testing Qualifications Board (ISTQB) [defines confirmation testing](https://glossary.istqb.org/en_US/term/confirmation-testing) as testing after a fix, to confirm that the failure the defect caused no longer happens. That check is also called [retesting](https://specstory.com/learning/debugging/bug-fix-verification). A fail-before, pass-after check adds a run of the same test on the code without the fix, which shows that the test can detect the failure at all.

## How does a fail-before, pass-after check work?

The check runs one test on two versions of the source code, in this order:

1.  The author of the change writes a test that encodes the bug's [reproduction steps](https://specstory.com/learning/debugging/reproduction-steps) and the expected result.
2.  A person or a script restores the source files from the [commit](https://specstory.com/learning/glossary#commit) before the change and keeps the added test file. Checking out the whole old commit would also remove a test committed with the change.
3.  The test runner runs the test on the old source, where it fails.
4.  A person reads the failure message and checks that it shows the bug's symptom, not a missing import.
5.  The person or script restores the changed source, and the same test passes.
6.  The test joins the suite as a [regression test](https://specstory.com/learning/testing/regression-testing), so the suite fails if the same bug returns.

Diagram: One test, two runs

The test file stays the same in both runs, and only the source changes. A pass on the changed source counts only after a failure on the old source.

The two runs have four possible outcomes:

| Old code | Changed code | What it shows |
| --- | --- | --- |
| Fails | Passes | The test detects the bug, and the change removes the failure it checks |
| Passes | Passes | The test does not reach the bug, or the old code never had it |
| Fails | Fails | The change did not fix what the test checks, or the test is wrong |
| Passes | Fails | The change altered behavior that the test checks |

Only the first row passes the check. A [catching test](https://specstory.com/learning/test-quality/catching-tests) expects the opposite pattern, a pass on the old code and a failure on a bad change. [Git bisect](https://specstory.com/learning/debugging/git-bisect) uses binary search over commits to find where one check's result changes, e.g. the commit that fixed a bug.

## What is an example of a fail-before, pass-after check?

Here is an illustrative example. Acme Co. sells furniture online. A customer reports a checkout bug after changing the delivery address. The address saved, but the cart emptied.

A developer asks a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) to fix it. The agent edits `change_address` in `src/acme/checkout.py`, adds this test, and reports that it passes:

```python
from acme.checkout import Order, change_address

def test_change_address_keeps_cart():
    order = Order(id="A-1042", items=["sofa", "lamp"])
    updated = change_address(order, "12 Elm Street")
    assert updated.address == "12 Elm Street"
```

The developer commits the work, then runs the added test on the source from the commit before the fix:

```bash
git restore --source=HEAD~1 -- src/
pytest tests/test_checkout.py::test_change_address_keeps_cart
git restore -- src/
```

The run reports `1 passed`. The test passes on the buggy code, so it is no evidence for the fix. Its name mentions the cart, but it checks only the address.

The developer adds one line to the test, `assert len(updated.items) == 2`, and repeats the run. On the old source, the test now fails with the bug's symptom:

```text
>       assert len(updated.items) == 2
E       AssertionError: assert 0 == 2
E        +  where 0 = len([])
```

On the fixed source, the same test passes, and it stays in the suite. This example is simplified. A real review would also add a deliberate bug to the changed line, one of the steps for [checking agent-written tests](https://specstory.com/learning/test-quality/checking-agent-written-tests).

## What changes when a coding agent writes the code?

A coding agent that fixes a bug often writes the fix and the test in one session. If it runs the test only on the changed code, the only run anyone sees is the passing one.

Some [AI test case generation](https://specstory.com/learning/test-quality/ai-test-case-generation) tools drop generated tests that fail on the current code. That filter cannot show that a kept test would fail on a bug. The check also needs the same test in both runs, but an agent can loosen a failing test, which is [test tampering](https://specstory.com/learning/verification/test-tampering).

[SWE-bench](https://arxiv.org/abs/2310.06770), a benchmark built from GitHub issues, keeps a task only when at least one test from the project's own fix changes from fail to pass. The model gets the issue and the codebase but not those tests, and its patch is scored on them plus tests of existing behavior.

A team can do the same by committing a test from the bug report before the fix and checking that it fails there. A [continuous integration](https://specstory.com/learning/ci-cd/continuous-integration) (CI) job can repeat the check with a script. The script lists the added test files with `git diff --name-only --diff-filter=A origin/main...HEAD -- tests/`, runs them on the base branch's source, and fails the job if they pass there.

## What can a fail-before, pass-after check miss?

The check still leaves these gaps:

-   **The cause of the failure.** A test can fail on the old code for a reason unrelated to the bug, e.g. a missing import. Only the failure message shows whether the first run reached the bug.
-   **New features and refactors.** A test for a new feature fails on the old code only because the feature is missing, so a deliberate bug in the new code is the stronger check. That idea is behind [mutation testing](https://specstory.com/learning/test-quality/mutation-testing). A test for a refactor should pass on both versions, because a refactor must not change behavior.
-   **Other inputs.** The check covers the inputs in the test. The fix can still fail on inputs the report did not show, e.g. a cart with one item.
-   **Intermittent bugs.** A bug that appears only on some runs can let the test pass on the old code by chance. One run in each direction is then not enough.

A test that fails before and passes after shows that the change removed one failure, not that the software works.

## How is fail-before, pass-after different from test-driven development?

[Test-driven development](https://specstory.com/learning/testing/test-driven-development) (TDD) is a way to write code. Its red step runs a test before the code that passes it exists, so the test fails because the behavior is missing. A fail-before, pass-after check judges a test that already exists by running it on the source from an older commit. It works on a test written after the fix, by a person or an agent.

The two overlap when a developer writes a bug's test before the fix, because the red run on the unfixed code is then the fail-before run. Both rest on one rule. A test counts as evidence only after it has failed for the right reason.

## How do you show that the fix made the test pass?

The fix made the test pass only if the same test, unchanged, failed on the code without the fix. Keep a record of both runs, with the version of the code, the command, and the failure message. A summary that says the test failed first is not a record of a run, so rerun the check before the merge.

A passing test covers only what it checks, and the workflow around the fix can still fail. RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### What does it mean when a new test passes on the old code?

A new test that passes on the old code does not detect the bug the change was meant to fix. The test may check a value the bug never touches, or the old code may never have had the bug. Its pass on the changed code shows little about the fix.

### How do you run a new test against the previous commit?

To run a new test against the previous commit, restore the source folder from that commit with Git and leave the test folder as it is. Run the test, then restore the changed source. A full checkout of the old commit would also remove a new test committed with the fix.

### Does every new test need a fail-before check?

A fail-before check suits a new test that claims to catch a bug or a change in behavior. A feature test fails on the old code only because the feature is missing, and a refactor's tests should pass on both versions.

### Can CI run fail-before, pass-after checks automatically?

A CI job can run fail-before, pass-after checks with a script. The script finds the tests a change adds, runs them on the base branch's source, and fails the job if they pass there. A person still reads each failure message to confirm that it shows the bug.

---

Source: [Fail-before, pass-after | Test fails on old code | SpecStory](https://specstory.com/learning/test-quality/fail-before-pass-after)
