Skip to content

How to report and fix an intermittent bug

Reporting an intermittent bug means recording how often it happens, under which conditions, and with what evidence, so others can make it fail on purpose.

Last updated , 9 min read

Key points

  • Report the number of failures out of a stated number of tries, with retries off.
  • Compare failing runs with passing runs to find the condition the bug needs.
  • Force that condition so the bug fails on purpose, then check the fix with repeated runs.

How do you report and fix an intermittent bug?

To report and fix an intermittent bug, record how often it fails, the conditions of failing and passing runs, and the evidence from a failure. Compare the runs to find the condition the bug needs, then force that condition so the bug fails on purpose. Fix the cause, then rerun enough times to rule out the old failure rate.

An intermittent bug, also called an intermittent failure, is a defect that makes the same steps fail in some runs and pass in others on the same code. An input that the steps do not control, a source of nondeterminism, decides the result. A race condition is a common cause. Many other intermittent bugs fail every time under one condition that nobody recorded.

An intermittent bug slows debugging, which usually starts by reproducing the failure. Reproduction steps that end in "sometimes" give the fixer no baseline. A count, e.g. 3 failures in 20 tries, gives one.

From a rare failure to a checked fix Record rate and evidence Compare failing and passing Force the condition Fix remove the cause Check with repeated runs fails sometimes fails on purpose forced runs pass still fails, so compare the runs again
Forcing the condition turns a rare failure into one that happens on purpose. A fix that still fails sends the work back to comparing runs.

What do you need before you start?

If a test fails in continuous integration (CI), first check whether it is a real bug or a fault in the test. Then record these for each failure and for some passing runs:

  • The rate. Count failures and tries of the same steps with retries off.
  • The conditions. Note the version, platform, data, time, and load of each run.
  • The evidence. Save the output of a failing run, e.g. its log lines, before a rerun replaces it.
  • A rerun command. Keep a command that repeats the steps, so anyone can count the rate again.

A 2026 study found that for 58% of the flaky end-to-end tests it studied, the test code and CI logs were not enough to find the cause.

How do you report and fix an intermittent bug step by step?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asked a coding agent to "Let customers edit their delivery address during checkout."

After the release, a few customers report the same failure. The address saved, but the cart emptied. The developer tries the steps 10 times on staging, and each passes.

1. Save the evidence

The server log records each address edit and where the cart came from. The developer saves failing and passing lines before the log rotates:

09:14:02 PATCH /api/checkout/A-1042 req=7f3a cart=miss items=0
09:15:40 PATCH /api/checkout/A-1043 req=81c2 cart=hit  items=2

2. Count failures and tries

The log holds 480 address edits from one week, and 12 of them emptied a cart, or 1 edit in 40. The developer files the report now, with that count and with any suspected condition marked as a guess.

3. Compare failing and passing runs

Your own tries are passing runs, and the difference from the failures points at the condition. Each failure logged cart=miss, and the passing edits logged cart=hit. The cart cache drops an entry 30 minutes after the cart last changes, so the developer's tries, on carts changed under a minute earlier, never missed the cache.

4. Make the bug fail on purpose

The developer writes an end-to-end test in which a hypothetical helper, present only in the test build, expires the cache entry:

import { test, expect } from "@playwright/test";
import { addItemsToCart, expireCartCache, editAddress } from "./helpers";

test("an address edit after a cache miss keeps the cart", async ({ page }) => {
  await addItemsToCart(page, ["oak-chair", "pine-table"]);
  await expireCartCache(page); // test build only
  await editAddress(page, "12 Elm Street");
  await expect(page.getByTestId("cart-item")).toHaveCount(2);
});

Run 20 times with retries off, the test fails every time with 0 cart items:

npx playwright test tests/repro/address-cache-miss.spec.ts \
  --repeat-each=20 --retries=0

Without the expireCartCache line, it passes 20 of 20 runs.

5. Update the report

The updated report uses the fields of a bug report and adds the failing and passing conditions:

Title: Editing the address empties the cart after a cart cache miss
Environment: production, build 4.12.0
Rate: 12 of 480 address edits in one week, from the server log
Fails when: the cart last changed over 30 minutes before the edit (cache miss)
Passes when: the cart changed within 30 minutes (cache hit)
Tried: 10 edits on staging, each within a minute of a cart change, 0 failures
Observed: The address saved, but the cart emptied.
Expected: The address changes, and the cart keeps its items.
Evidence: log lines and request IDs of the 12 failures, attached
Repro (fails 20 of 20): npx playwright test tests/repro/address-cache-miss.spec.ts \
         --repeat-each=20 --retries=0
Owner: a developer on the checkout team

The named owner, usually on the team that owns the failing workflow, keeps the report open and adds each later failure to its rate.

6. Fix the cause and rerun

The coding agent receives the report and the test. Its change loads the cart from the database after a cache miss, and the forced test passes 20 of 20 runs. The next week's log shows 0 emptied carts in about 480 address edits, against about 12 at the old rate.

This example is simplified. A real intermittent bug often needs several rounds of comparing runs.

How do you raise the reproduction rate?

These techniques make the failure show up more often, up to a forced run that fails every time:

  • Repeat the steps. Playwright's command-line options include --repeat-each, which runs each test N times, and a shell loop repeats any other test command. Turn retries off, because a retry can pass the run after the bug fails.
  • Add load. More parallel workers, or load testing, make operations overlap more often.
  • Change one condition at a time. Vary the data, the test order, or the machine. A failure that appears only after other tests points at shared state that test isolation removes.
  • Match the failing environment. A busy runner changes timing, one reason tests pass locally and fail in CI.
  • Force the condition. A test build can trigger the condition directly, e.g. a helper that expires a cache entry, as in step 4.
  • Record the run. On Linux, the rr debugger replays a recorded run exactly, as often as needed. Its chaos mode, rr record --chaos, randomizes scheduling to make intermittent bugs appear more often.
  • Run a race detector. A race detector, e.g. Go's go test -race, reports unsynchronized memory access even in runs where the bug does not show.

What changes when a coding agent writes the code?

A coding agent's own tests usually run each step once, quickly, on fresh data, where the conditions an intermittent bug needs rarely occur.

A single green rerun is also weak evidence of a fix. A bug that fails 12 times in 480 passes a single rerun almost every time, so the agent's output can say "done" while the bug stays in the code. An agent asked to stop the failure can also make it rarer, e.g. by raising the cache life to a day.

A practical adjustment is to judge the fix by the forced test, not by the agent's summary. Ask for the forced test's result on the old code and on the change, with the evidence from both runs.

What are common mistakes?

These mistakes keep an intermittent bug hidden or open:

  • Writing "sometimes" in place of a count. Nobody can then tell a fix from a lucky run.
  • Closing a report that nobody could reproduce. A few passing tries cannot rule out a bug that fails 1 run in 40.
  • Trusting a quiet debugging session. A debugger or extra logging can shift timing enough to hide a timing bug. A bug that changes or disappears when someone observes it is called a heisenbug. A forced condition does not depend on that timing.
  • Making the bug rarer. A sleep, a longer timeout, or a retry lowers the rate and leaves the cause in the code.

How do you check that it worked?

The fix holds when the forced reproduction fails on the old code and passes on the fix. Then run the unforced steps with retries off, as many times as the old failure rate calls for. A rough rule from bug fix verification is 3 divided by the failure rate, e.g. about 120 clean runs for a bug that failed 1 run in 40. Keep the forced test in the suite, and watch for the symptom after the release.

How does RunStory help with intermittent bugs?

An intermittent bug is hard to fix without the evidence of a run that failed. RunStory runs your software in a separate environment. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is the same kind of evidence that step 1 saves by hand.

After your agent's change, RunStory repeats the failing workflow. It is in private alpha for CLIs and web apps. For a bug that fails at a rate, the forced test and the repeated runs above still decide whether the fix holds.

Join the RunStory alpha →

FAQs

Is an intermittent bug the same as a heisenbug?

An intermittent bug is not always a heisenbug. An intermittent bug fails in some runs of the same steps, while a heisenbug changes or disappears when someone observes it, e.g. under a debugger. An intermittent timing bug that stops failing under extra logging is both.

What should a report say about a bug that happens half the time?

A report about a bug that happens half the time should go out once the count is known, not after the cause is found. The updated report in the steps above shows what it contains.

Who should own an intermittent bug that users report?

An intermittent bug that users report should have one named owner, usually on the team that owns the failing workflow. The owner keeps the report open while nobody can reproduce the bug, and adds each later user report to its rate.

Can extra logging hide an intermittent bug?

Extra logging can hide an intermittent bug that depends on timing, because writing log lines shifts when operations run. A bug that needs a data condition, e.g. a cache miss, still fails, and the log can show that condition.