Key points
- Save the evidence from the first failure before anyone reruns the test.
- A pass on retry does not cancel a failure on the same code.
- Compare failure counts on the changed code and on the code before the change.
How do you tell a flaky test from a real bug?
Rerun the failed test 30 or more times on the same code with retries off, and keep the evidence from every failure. Then run it as often on the code before the change. The counts show whether the failure started with the change. The saved evidence shows whether the fault is in the test or in the software.
A pass on retry is a test result that fails first and then passes when the same test runs again on the same code. It is a sign of nondeterminism, from a flaky test or from an intermittent bug in the software. The second result does not cancel the first.
Telling the two apart is often the first step in debugging a failed run in a continuous integration and delivery (CI/CD) pipeline. At Google in 2016, about 84% of the times a test went from pass to fail involved a flaky test rather than a real break. Those odds favor a flake but cannot classify any one failure.
What do you need before you start?
Gather these before the first rerun:
- The first failure's record. Keep the error message and the trace of the failed attempt, not the retry.
- The exact code. Note the commit the test ran on and the commit before the change.
- A rerun command. Use runner options that repeat one test with retries off and keep a trace of each failure, e.g. Playwright's
--repeat-eachwith--retries=0and--trace=retain-on-failure. - A matching machine. Rerun on the CI runner or on a machine set up the same way, since poor environment parity can hide a failure.
How do you tell a flaky test from a real bug step by step?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asked a coding agent to "Let customers edit their delivery address during checkout." After the change, the checkout's existing end-to-end test failed once in CI and passed on a second run.
1. Save the first failure
Before any rerun, the developer downloads the failed attempt's trace and copies its error:
tests/checkout.spec.ts:14:5 › customer can place an order
Error: expect(received).toHaveLength(expected)
Expected length: 2
Received length: 0
Received array: []
The run used commit a41c9e2 and 4 parallel workers.
2. Rerun the test on the same code
The developer runs only this test 30 times with retries off, keeping each failure's trace:
npx playwright test tests/checkout.spec.ts:14 \
--repeat-each=30 --retries=0 --workers=4 \
--trace=retain-on-failure --output=runs/change
Three of the 30 runs fail with the same error. The failure reproduces, but its cause is still open.
3. Rerun the test on the code before the change
git switch --detach a41c9e2~1
npx playwright test tests/checkout.spec.ts:14 \
--repeat-each=30 --retries=0 --workers=4 \
--trace=retain-on-failure --output=runs/before
git switch -
All 30 runs pass. If the older code failed 1 run in 10, a clean streak of 30 would happen fewer than 1 time in 20. The failure most likely started with the change.
4. Change one condition at a time
Back on the changed code, 30 more runs with --workers=1 pass, so the failure is likely tied to parallel load. That fits a test fault, e.g. data shared between workers. It also fits a race condition in the checkout.
5. Read the evidence from a failing run
A saved trace shows the address save and a shipping recalculation running at once. The address save read the checkout while the recalculation had cleared the items, and wrote that empty copy back last. The address saved, but the cart emptied. The test's check is correct, so the fault is in the software.
6. Report the failure as a bug
The developer files a bug with the failure counts, the rerun command, and the trace, which form reproduction steps that another person can run. This example is simplified. A real investigation would also compare browsers and CI machines, e.g. a runner with less memory.
How many reruns confirm a failure?
One failure on unchanged code confirms that the result is nondeterministic. To rule out a failure that happens once in N runs, rerun about 3N times with retries off, e.g. 30 times for a failure of 1 run in 10.
A failure that happens on every run needs few reruns. When a handful of reruns all fail, the failure is consistent, and the next step is to check whether the change or the environment causes it.
Showing that a failure is rare, or gone, takes far more runs. If a failure happens on 1 run in N, the chance that R reruns all pass is (1 - 1/N)^R. This table gives the reruns needed to see a failure at each rate at least once, 19 times in 20:
| Failure rate | Reruns needed |
|---|---|
| 1 run in 2 | 5 |
| 1 run in 10 | 29 |
| 1 run in 20 | 59 |
| 1 run in 50 | 149 |
| 1 run in 100 | 299 |
Three clean reruns still miss a failure of 1 run in 10 most of the time. The math also assumes that runs are independent. Back-to-back reruns on one idle machine may not be, so vary the load and the test order across the reruns.
What are common mistakes?
These mistakes hide an intermittent bug or its cause:
- Retrying until green. A pass on retry that hides a bug in the software is one form of false pass. Some runners let the build pass, e.g. Playwright, which labels the test "flaky" and exits with success unless
--fail-on-flaky-testsis set. - Rerunning before saving the evidence. A green rerun can replace the failed run's files, and the rare timing may not return.
- Diagnosing from the log alone. A 2026 study found that for 58% of flaky end-to-end tests, the code and CI logs were not enough to find the cause. Finding those causes would need more evidence from running the tests.
- Rerunning the test only on its own. A failure that appears only in the full run often comes from state an earlier test left behind, which test isolation removes. Rerun the full suite too.
How do you check that it worked?
The classification holds when a bug fix aimed at that cause stops the failure. Rerun the same command as many times as the table gives for the old failure rate. The test should also still fail on the broken commit at about the old rate, which shows that it still detects the bug. Keep retries off for this test, so a return of the failure stays visible.
What changes when a coding agent writes the code?
A coding agent that is told to make the checks pass can rerun a red test until it passes and then report the task as "done." Its summary shows a green run, and the first failure drops out of view.
A pass on retry is not always a test fault, so a fix needs more than one green run. In one team's cleanup of 63 flaky tests with Claude Code, a fix counted only after the test passed 15 times in a row.
A practical adjustment is to have the agent run the rerun command with retries off on its branch and on its starting commit, and report both failure counts. Treat any pass on retry in its report as a failure to investigate.
What should a report on a failing test include?
A useful report names the revision the test ran on and the failure counts there and on the code before the change. It gives the rerun command and the worker count, and it attaches the error and trace from the first failure. A coding agent that fixes the failure needs the same report.
RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps.
FAQs
How do you tell flaky from real failures after a retry passes?
A retry that passes does not settle the question. Rerun the test many times with retries off, on the changed code and on the code before it. Then read a failing run's evidence to find the fault.
What evidence should you keep when a test passes on retry?
The evidence to keep is the first failure's record, including its error message, trace, code version, and runner settings. A green rerun can replace those files, and the rare timing behind the failure may not return.
If it fails once in three runs, did it pass?
A test that fails once in three runs did not pass. One failure on the same code shows that the result is nondeterministic, and the two passes do not cancel it. The failure needs more reruns and a look at its evidence.
Does automatic retry remove the nondeterminism?
Automatic retry does not remove nondeterminism. A retry runs the same code again, so the timing or shared state behind the failure is still there. The retry changes only the reported result, which can turn an intermittent bug into a passing build.