What is a flaky test?
A flaky test is a test that passes on some runs and fails on others while the code and the test stay the same. Engineers also call it a flake or a nondeterministic test. Because the same code produces both results, one run of a flaky test says little about whether the code works.
Google labels a single result flaky, while researchers label the test itself flaky. A failure that does not repeat is hard to trace, so flakiness is a debugging problem as well as a testing problem. One flaky test makes a whole test suite harder to trust.
At Google in 2016, about 1.5% of test runs were flaky, and almost 16% of tests showed some flakiness. About 84% of the times a Google test went from pass to fail involved a flaky test. The same post reports that engineers often ignore real failures in flaky tests, because false alarms are so common.
How does a test become flaky?
A test becomes flaky when its result depends on an input that the test does not control. That hidden input is a source of nondeterminism. A flaky failure happens in this order:
- The test runs the code with the inputs it set up.
- The result also depends on an input the test did not set, e.g. how long a page takes to load.
- When that input has the value the test expects, the test passes.
- When the input has another value, the test fails, although the code did not change.
The first large study of flaky tests (2014, 51 projects) found the most common causes were asynchronous waits, concurrency, and dependence on test order. The hidden input usually comes from one of these sources:
- Asynchronous waits. The test checks a result before an operation finishes, e.g. after a fixed sleep. A check that waits for the result fixes it.
- Concurrency. Two threads or requests change the same data, and their order changes the result. A race condition in the product can show up this way.
- Test order. The test depends on state that an earlier test left behind, e.g. a database row. Test isolation removes this cause.
- Environment. The test relies on something outside the code, e.g. a live network service. A gap in environment parity can make a test pass locally but fail in continuous integration (CI).
- Unordered results. The test expects an order that the code does not promise, e.g. rows from an unsorted query. Sorting the results before the check fixes it.
- Time and randomness. The test reads the clock or a random value, e.g. a date check that fails near midnight in another time zone. A fixed clock or seed removes it.
What is an example of a flaky test?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds an end-to-end test that drives the checkout in a headless browser, e.g. with Playwright:
test("customer can change the delivery address", async ({ page }) => {
await page.goto("https://shop.example.com/checkout");
await page.getByLabel("Delivery address").fill("12 Elm Street");
await page.getByRole("button", { name: "Save address" }).click();
await page.waitForTimeout(200);
expect(await page.locator("#confirmation").isVisible()).toBe(true);
});
The test waits 200 milliseconds and checks once for the confirmation. On the developer's laptop, the confirmation page usually loads within 200 milliseconds, and the test passes. On a busy CI runner, the page sometimes takes 250 milliseconds, so the check runs too early and the test fails:
Error: expect(received).toBe(expected) // Object.is equality
Expected: true
Received: false
Someone reruns the job, and the test passes. Nothing changed between the runs, so the test is flaky. The fix replaces the fixed wait with a check that retries until the confirmation appears or a timeout expires:
await expect(page.locator("#confirmation")).toBeVisible();
The test now waits for the page, not for the clock. This example is simplified. A real suite would have other hidden inputs, e.g. shared test accounts.
What is test quarantine?
Test quarantine is a policy that moves a flaky test out of the tests that can fail a build while someone fixes it. Its failures no longer stop a merge in the CI pipeline.
A 2011 article by Martin Fowler recommends quarantine so that healthy tests keep giving clear feedback. It also warns that teams can forget quarantined tests. Google's post warns that quarantine can hide a real race condition. A quarantine policy can set these rules:
- Owner. A named person or team fixes each quarantined test.
- Tracking issue. Each quarantined test links to an open issue with the failure and its evidence.
- Limit. A cap on the number of quarantined tests or on their time in quarantine forces a fix or a deletion.
- Visible results. The test still runs in a job that cannot fail the build, so the team can tell when the flakiness stops.
What changes when a coding agent writes the code?
A 2026 study of four database systems found that tests generated by a large language model (LLM) were slightly more likely to be flaky than existing tests. The flaky generated tests most often relied on an ordering the system did not guarantee.
When a flaky test fails during a task, a coding agent can make the run pass by raising a timeout or skipping the test. Its output can then report the task as "done" while the cause remains. Weakening a test so that it passes is test tampering. One green run cannot show that a flake is fixed, because a flaky test passes on most runs anyway.
A practical adjustment is a review rule for tests that an agent writes. The rule rejects a fixed sleep, e.g. waitForTimeout, and an assertion on an order the code does not promise. It also rejects a larger timeout or a skip that comes without a stated cause.
What are the limits of retrying flaky tests?
Automatic retries rerun a failed test and let the build pass if a later attempt passes. That also turns an intermittent product bug green. In one team's cleanup of 63 flaky tests, 5 turned out to be real product bugs that retries had hidden for months. Retries have these limits:
- A retry hides the cause. A pass on retry shows that the result changed, not whether the test or the product is at fault. Some test runners, e.g. Playwright, label such a test "flaky" and by default let the run pass. Playwright's
--fail-on-flaky-testsoption fails the run instead. - A retry leaves the nondeterminism in place. The hidden input is still there and can fail the test again later.
- Retries cost time. Each retry of a slow end-to-end test adds its full run time.
The same team counted a fix only after the test passed 15 times in a row.
How is a flaky test different from a broken test?
A broken test fails on every run of the same code. A flaky test fails on some runs only, so each failure is harder to reproduce and to debug. Running a test many times on the same code, e.g. with npx playwright test --repeat-each=50 --retries=0, shows which kind it is.
A flaky result can also come from a real bug. When the product fails intermittently, e.g. because of a race condition, the test is correct and the flakiness is a product bug. Teams disagree on whether flakiness is mostly a test problem or a hidden product bug.
How do you show that a failure is real?
A failure is real when it comes from the product, not from an input the test does not control. Keep the log from the first failing run, and write down the reproduction steps. Repeat those steps on the same code until the failure returns, and rule out the test's hidden inputs, e.g. the network speed. Then repeat them after the fix.
RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. Your coding agent receives the actions RunStory took and evidence of the unexpected result. That is the record the steps above ask you to keep. It is in private alpha for CLIs and web apps.
FAQs
How should a team deal with tests that fail intermittently?
A team with tests that fail intermittently should first confirm that each test is flaky by rerunning it on the same code. It can then quarantine the test with an owner and fix the hidden input.
Can a tool detect whether a test is flaky?
A tool can confirm that a test is flaky only when it records the test passing and failing on the same code. Rerunning a test many times on purpose finds a rare flake sooner.
Is a flaky test the same as a nondeterministic test?
A flaky test and a nondeterministic test are two names for one thing, a test that passes on some runs and fails on others on the same code.
Would you trust an AI-written fix for a flaky test?
An AI-written fix for a flaky test is worth trusting when it names the hidden input, removes it, and then passes many repeated runs. A fix that only raises a timeout or skips the test usually hides the flake.