Skip to content

Do test retries hide real bugs?

Automatic test retries turn a failing run green when a later attempt passes, which hides flaky tests and can also hide real intermittent bugs.

Last updated , 9 min read

Do test retries hide real bugs?

Test retries, also called automatic reruns, can hide real bugs. A retry runs a failed test again, and the run passes if a later attempt passes. An intermittent bug in the software then looks the same as a flaky test, and the build turns green. A bug that fails the test on every attempt still fails the run.

A result that fails first and then passes on the same code is a pass on retry. Test runners let it pass the run, and some also label the test "flaky." In one team's cleanup of 63 flaky tests, 5 turned out to be real product bugs that retries had hidden for months.

Retries save time when the fault is in the test. When it is in the software, the bug can ship with a green check, and debugging starts later from a customer's report.

What causes a retry to hide a bug?

Three conditions let a retry turn a real failure into a pass.

Why does the check show only the last attempt?

A runner with retries on uses the last attempt to pass or fail each test. Each runner below keeps retries off until a setting turns them on, although the Playwright starter config turns them on in continuous integration (CI):

RunnerRetry settingWhat a pass on retry reports
Playwrightretries in the config, or --retriesThe run passes, and the test is listed as "flaky"
Cypressretries in the configThe test counts as passed
Jestjest.retryTimes() in the test fileThe test counts as passed
Vitestretry in the config, or --retryThe test passes, and verbose output adds (retry x1)
pytest--reruns, from the pytest-rerunfailures pluginThe test passes, the summary counts a rerun, and the earlier tracebacks are dropped

A CI check usually reads only the job's exit code. Playwright exits with success when its only problem is a flaky test, unless failOnFlakyTests or --fail-on-flaky-tests is set. The pull request then shows a green check.

How a Playwright retry turns a failed test green Attempt 1 fails Retry new worker Attempt 2 passes Exit code 0 run passes CI check green Test report first error kept, test listed as flaky the pull request shows only the green check
In Playwright, the last attempt sets the exit code, and the exit code sets the check. The first failure survives only in the report.

Why does an intermittent bug pass on the next attempt?

An intermittent bug fails only when an input the test does not control takes a bad value, a source of nondeterminism. A race condition is a common case, where two operations go wrong only when they overlap.

A retry does not repeat the conditions of the first attempt. Playwright retries a failed test in a new worker process with a new browser. Jest retries it after the other tests in its describe block or file finish, unless retryImmediately is set. The load and the timing change, so the bad overlap often does not return.

Overlaps are more common on a busy CI runner, which is one reason tests that pass locally fail in CI.

Why do teams stop reading the flaky label?

Flaky results are common, so a pass on retry reads as noise. At Google in 2016, about 84% of the times a test went from pass to fail involved a flaky test.

Google also let teams mark a test as flaky, so it reported a failure only after several failed runs in a row. The post says this cuts false alarms but "encourages developers to ignore flakiness in their own tests." It also delays the report of a real break until the extra runs finish.

What does a hidden bug look like after a retry?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Acme's Playwright config keeps these lines from the starter file of npm init playwright:

  retries: process.env.CI ? 2 : 0,
  use: {
    trace: 'on-first-retry',
  },

The agent changes the address handler and opens a pull request. In CI, the end-to-end test "customer can place an order" fails on its first attempt and passes on the retry. The output includes these lines:

  1) [chromium] › tests/checkout.spec.ts:14:5 › customer can place an order

    Error: expect(received).toHaveLength(expected)

    Expected length: 2
    Received length: 0
    Received array:  []

  1 flaky
    [chromium] › tests/checkout.spec.ts:14:5 › customer can place an order
  41 passed (2.1m)

The job exits with code 0, and a reviewer approves the change. After the release, a customer changes the delivery address during checkout. The address saved, but the cart emptied.

The test was right on its first attempt. The address save had overlapped with a shipping recalculation and written back an empty cart, and on the retry the two did not overlap. The on-first-retry trace mode recorded only the passing retry.

This example is simplified. A real report often lists several flaky tests at once.

What changes when a coding agent writes the code?

A run with a pass on retry exits with code 0 and counts no failed tests. A coding agent that checks both finds no failure, so its output can report that the tests passed. That is one reason "no bugs found" does not mean no bugs.

An agent asked to make a red check pass can also add retries. One line, e.g. jest.retryTimes(3) at the top of a test file, can turn an intermittently failing check green without changing the code under test. The green check is then a false pass, and the agent's output can say "done" while the bug remains. A self-healing test can hide a real break by another route, since it changes its own steps after a failure instead of failing.

A practical adjustment is to review retry settings as part of the tests. Reject an agent's change that adds or raises retries, and ask for evidence from each attempt, not only the final status.

How can teams keep retries honest?

A flaky test can still be retried if the pass on retry stays visible. These practices keep the flaky label without the green check:

  • Fail the run on a pass on retry. One retry still labels a failure as flaky or consistent, and failOnFlakyTests fails the run either way.
  • Turn retries off for new and changed tests. Those tests have no flakiness history in their current form, so a pass on retry points at the change. In Playwright, a separate CI step can run npx playwright test --only-changed=origin/main --retries=0.
  • Keep retries few and narrow. Each retry of a slow end-to-end test adds its full run time. Vitest's retry.condition retries only errors that match a pattern, e.g. ECONNREFUSED, so a failed assertion still fails.
  • Keep the first failure. The retain-on-first-failure trace mode keeps a trace of the failed attempt after the retry passes.
  • Count it in the gate. A quality gate that counts only failed tests lets a pass on retry through. Count it as a failure, and send it to CI failure triage to find whether the test or the change is at fault.

This Playwright config keeps the flaky label but not the green check:

export default defineConfig({
  retries: process.env.CI ? 1 : 0,
  failOnFlakyTests: !!process.env.CI,
  use: { trace: 'retain-on-first-failure' },
});

Known flaky tests then need quarantine, or they fail the run too.

How are retries different from quarantine?

A retry reruns a failed test inside the same run. It applies to any failing test, including one that the change broke. Test quarantine moves one known flaky test out of the tests that can fail a build. The test can keep running and reporting, and a named owner fixes it.

Both leave the cause in place. Only a fix that removes the hidden input ends the flakiness, and it counts once repeated runs pass with retries off.

How do you find the failures that retries hid?

Search past runs for tests that passed on retry. Playwright's JSON reporter gives such a test the status flaky and lists each attempt, including the failed one. Rerun those tests on the code where they first failed, with retries off and tracing on. Check the trace of any failure that returns for a fault in the software.

An old report shows that a test failed once, not the steps that make the failure happen again. RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. It is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

Which test frameworks can retry failed tests?

Test frameworks with a retry setting include Playwright, Cypress, Jest, and Vitest, and pytest gets one from a plugin. Retries are off by default, though Playwright's starter config turns them on in CI. Each of these runners lets a pass on retry pass the run.

Should retries be off for newly added tests?

Retries should be off for newly added tests, because a new test has no history of flakiness. A new test that passes only on retry is flaky from its first run, or it found an intermittent bug in the change.

Should a test that passed on retry block a merge?

A test that passed on retry should block a merge, because its first failure may come from the change. A known flaky test in unrelated code belongs in quarantine, where it can keep reporting without blocking while its owner fixes it.

Do retries cost CI time?

Retries cost CI time, because each retry runs the failed test again from the start, often in a new worker with a new browser. A slow end-to-end test adds its full run time for each retry.