Key points
- Run the exact steps that exposed the bug on the fixed code, and confirm the failure is gone.
- Watch the same steps fail on the old code first, or a pass on the fixed code means little.
- Check nearby behavior with regression tests, then keep the failing steps as a test in the suite.
How do you verify a bug fix?
To verify a bug fix, run the exact steps that exposed the bug on the fixed code and confirm that the failure no longer happens. This check is called retesting, or fix verification. Run the same steps on the old code first, because a pass means little unless they failed there. Then check nearby behavior with regression testing.
The International Software Testing Qualifications Board (ISTQB) calls retesting confirmation testing. It defines the term as testing done after a defect is fixed, to confirm that the failure it caused does not happen again.
Retesting comes right after the fix in debugging. A fix that has never been run against the original failure is a claim, not evidence. The change may miss the cause, or the check may never reach the bug.
What do you need before you start?
A retest needs five things:
- The original steps. The reproduction steps give the input, the data, and the environment that exposed the bug.
- The expected result. A record of what the software should have done gives the retest its pass condition.
- The old code. The commit before the fix is where the steps must first fail.
- The fixed code. The commit or build with the fix runs in the same kind of environment as the failing run.
- Nearby behavior. A short list names the workflows that share code or data with the fix, e.g. discount codes next to the cart.
How do you verify a bug fix step by step?
Here is an illustrative example. Acme Co. sells furniture online. A customer reports a checkout bug, and the report says "The address saved, but the cart emptied." A developer asks a coding agent to fix it, and the agent's output says the fix is "done."
1. Reproduce the failure on the old code
The developer writes the customer's steps as a test for a test runner, e.g. Playwright. The startCheckout helper is hypothetical:
import { test, expect } from "@playwright/test";
import { startCheckout } from "./helpers";
test("changing the address keeps the cart", async ({ page }) => {
const order = await startCheckout(page, { items: 2 });
await order.setAddress("12 Elm Street");
expect(await order.itemCount()).toBe(2);
});
On the commit before the fix, e.g. after git switch --detach <fix-commit>~1, the test fails the way the customer's cart did:
1) checkout.spec.ts:4:5 › changing the address keeps the cart
Error: expect(received).toBe(expected) // Object.is equality
Expected: 2
Received: 0
2. Read what the fix changed
git diff --stat main...HEAD lists one changed file, src/checkout/address.js, and no edits to existing tests. A fix that also edits or deletes existing tests needs a closer look.
3. Run the same test on the fixed code
On the commit with the fix, the same test passes with the same data. The developer also repeats the steps once by hand, because a test can pass for the wrong reason.
4. Repeat the run to rule out chance
The command turns retries off, because a pass on retry can turn the run green while the failure remains. The cart bug failed on every run of the old code, so a short repeat is enough:
npx playwright test checkout.spec.ts --repeat-each=20 --retries=0
20 passed (4.2s)
An intermittent bug needs more runs. Raise its failure rate on the old code first, e.g. with added load. Then run the fixed code many more times than it took to see one failure.
If the bug still failed 1 run in 10, a streak of 29 clean runs would happen fewer than 1 time in 20. Clean runs make chance unlikely, not impossible. A rough rule is 3 divided by the failure rate, e.g. 30 runs for a bug that fails 1 run in 10.
5. Check the behavior next to the fix
The fix reloads the cart after the address saves, which could drop a discount code. So the developer runs the checkout tests, and the full suite runs in the team's continuous integration and delivery (CI/CD) pipeline. The developer also reruns the test with 1 item and with 5, because a fix can miss cases the report did not show.
6. Keep the test and record the evidence
The test joins the suite as a regression test for this bug. The bug report records the fix commit, the failing run, and the passing runs. This example is simplified. A real project would also retest in an environment that matches the customer's.
What is the difference between retesting and regression testing?
Retesting checks that one known failure is gone. Regression testing checks that what worked before a change still works after it. ISTQB defines regression testing as testing after a change to detect whether defects were introduced or uncovered in unchanged areas of the software. A fix usually needs both.
The two checks differ on these points:
| Attribute | Retesting | Regression testing |
|---|---|---|
| What it runs | The steps that exposed one bug | Tests for existing behavior |
| Expected result on the old code | Fails | Passes |
| Expected result on the fixed code | Passes | Passes |
| After the fix | Its test joins the regression suite | Keeps running on later changes |
What are common mistakes?
These mistakes produce a pass that says nothing about the bug:
- Retesting different steps. A similar path, e.g. a different product in the cart, can pass while the original steps still fail.
- Skipping the failing run. Steps that pass on the old code did not trigger the bug, so their pass on the fixed code confirms nothing.
- Retesting with no code change. A request to "retest, maybe it's gone" asks for a pass without a fix, so ask what changed first. If nothing changed, a pass shows only that the failure is intermittent or depends on the environment.
- Retesting a build without the fix. A build that lacks the fix commit says nothing about the fix, so check the build's commit first.
- Treating a green CI run as the retest. A CI run checks only the tests it has, so a green run says nothing about a bug that no test reproduces.
How do you check that it worked?
The verification is complete when these statements are true:
- The test that reproduces the bug failed on the old code and passes on the fixed code.
- Repeated runs pass, and their number fits how often the bug appeared.
- The nearby tests and the full regression suite pass.
- The bug report names the fix commit and links the runs.
If the bug returns later, the kept test lets Git bisect find the commit that brought it back.
What changes when a coding agent writes the code?
A coding agent that fixes a bug usually ends with a message that says the fix is "done." That message comes from the process that wrote the fix, and it is not a retest. A false completion claim arrives with no passing run of the original steps behind it.
The agent can also turn a check green without removing the bug, e.g. by loosening the assertion that exposed it. That is test tampering, and it often shows in the diff as an edit to the failing test. When an agent writes the test after the fix, nobody sees it fail on the old code unless someone runs it there.
The adjustment is to set the reproduction before the agent starts. Write the failing test from the report, run it on the old code, and keep that file out of the agent's edits. Then a person or a separate test runner runs it on the fix.
How does RunStory help with verifying bug fixes?
A bug fix is often checked only by the coding agent that wrote it. RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. Your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
How can you be sure a rarely reproduced issue is fixed?
A rarely reproduced issue cannot be shown fixed with certainty. Raise its failure rate on the old code first. Then run the fixed code until the clean streak is long enough to be unlikely if the bug were still there.
How should you respond to "retest, maybe it's gone"?
A request to retest with no code change should get a question back about what changed. When nothing changed, a pass shows only that the failure is intermittent or depends on the environment. That pass is not evidence of a fix.
Should you write a test after fixing a bug?
A test for a bug is best written before the fix, from the report, and checked to fail on the old code. A test written after the fix is still worth keeping when it fails on the old code. Either way, the test stays in the suite as a regression test.
Is a green CI run evidence that the bug is fixed?
A green CI run is evidence only that the tests in the suite passed. If no test reproduces the bug, the run says nothing about the fix. The run counts as evidence when the suite includes a test that failed on the old code.