What is an AI bug-fix loop, and how do you break it?
An AI bug-fix loop is a cycle in which each fix a coding agent makes causes another failure or brings back an earlier one. Developers also call it "whack-a-mole" fixing. Fixes stop undoing each other when each failure becomes a test that fails on the code before its fix, and each later fix must pass all of those tests.
Each round of the loop is a short pass through debugging, in which the agent receives a failure, changes the code, and checks the result. The round often ends when the check of the most recent failure passes, so a fix can undo an earlier one without any check failing.
An undone fix is a software regression, so the loop is one way that agents break working features. A finding is one reported failure, e.g. a customer's report that the cart is empty.
What causes an AI bug-fix loop?
Four causes are common, and a loop often has several.
Why does each fix check only the most recent failure?
After a fix, the agent's check often reruns only the one failure it was sent. Earlier findings may exist only as chat messages, not as tests. When two findings depend on the same code, a fix for the second can undo the fix for the first, and no check fails. The fixes then oscillate between two broken states.
Why do fixes land on the symptom?
An agent's fix often lands where the error shows, not where the wrong value starts. Such a symptom fix leaves the cause in place, so the failure returns elsewhere as the next finding.
Why does a long session make the loop worse?
The model behind a coding agent receives the session so far, or a summary of it, on each turn, so failed attempts stay in its context. The best practices guide for Claude Code says that after repeated corrections on the same issue, "the context is cluttered with failed approaches." It advises clearing the session after 2 failed corrections and starting again with a better prompt.
Why does an AI review keep the loop going?
Research by Huang and colleagues found that large language models struggle to correct their own reasoning without outside feedback. On their reasoning tasks, results sometimes got worse after the models tried.
A review by the agent that wrote the fix is not outside feedback. A separate AI reviewer has a limit too. Claude Code's guide warns that a reviewer prompted to find gaps will usually report some, even when the work is sound. Changes made to answer such reports can break code that the last review passed. The guide advises asking the reviewer to flag only gaps that affect correctness or the stated requirements.
What does an AI bug-fix loop look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer asked a coding agent to "Let customers edit their delivery address during checkout." Later, the developer pastes a customer's report into the same session, and the loop runs in 6 steps:
- The report says "The address saved, but the cart emptied."
- The agent changes the address handler to copy the existing checkout record first. It reruns the address steps, the cart keeps 2 items, and its output says "done."
- A reviewer changes the address to 12 Elm Street, farther from the warehouse. The shipping cost stays at $10.00 instead of $25.00, because the copied record keeps the old cost.
- The agent fixes the cost by rebuilding the checkout record from the address. Its shipping check passes, and its output says "done" again.
- The developer repeats the customer's steps. The cart is empty again, because the rebuilt record holds no items.
- The agent's third fix copies the record again, and the shipping cost goes back to $10.00.
Each fix passed the agent's check, and each fix after the first undid the one before it. The first finding lived only in the chat, so nothing reran it. Each "done" was a false completion claim, because no run checked the whole address change.
This example is simplified. A real loop often adds more findings, e.g. a broken discount code.
How can teams break the loop?
Fixes stop undoing each other when the check after each fix covers every finding so far. These steps turn findings into checks that stay:
- The developer stops the agent and returns the code to a known state, e.g. the version from before the loop. Reverting alone does not end the loop, because the next attempt can make the same change. In Claude Code, rewinding restores only edits made through its file editing tools, so Git is the safer record.
- The developer writes one test per finding from its reproduction steps and runs it on the code where the finding appeared. The test must fail there and pass after the fix, which makes it a fail-before, pass-after check.
- The developer keeps these tests in files the agent does not edit, because an agent can make a failing test pass by weakening it.
- The agent receives the failing tests together, with the observed results, and a request for one change at the root cause.
- A person or a test runner outside the agent reruns the finding tests after each fix, which is retesting, and then the regression suite. The fix counts only when all of them pass. The developer then commits it, which gives the next round a known state to return to.
- After 2 failed attempts on the same finding, the developer clears the session and starts a fresh one. Its first prompt lists the findings, their tests, and what each attempt broke.
At Acme, the two findings become two tests for a test runner, e.g. Playwright. The startCheckout helper is hypothetical, and the cost is in dollars:
test("changing the address keeps the cart", async ({ page }) => {
const order = await startCheckout(page, { items: 2 });
await order.setAddress("12 Elm Street");
expect(await order.itemCount()).toBe(2);
});
test("changing the address updates shipping", async ({ page }) => {
const order = await startCheckout(page, { items: 2 });
await order.setAddress("12 Elm Street");
expect(await order.shippingCost()).toBe(25);
});
The shipping test fails on the code from step 2, and the cart test fails on the code from step 4. The fix that passes both copies the record and recalculates only the shipping cost. When no change can pass both tests, the findings conflict, and a person decides which behavior the software should have. Human-in-the-loop review makes the agent wait for that decision.
A kept test catches a known failure when it comes back. It cannot catch a failure that nobody has found, and it checks only the result it names. An attempt limit caps the time a loop costs, but it does not fix the bug.
How is a bug-fix loop different from a flaky test?
A flaky test passes on some runs and fails on others while the code stays the same. In a bug-fix loop, the code changes each round, and each failure follows from a change. Rerunning the failing test many times on the same code usually tells the two apart. A result that changes between runs points to a flaky test, and a loop failure repeats until the code changes.
The two can mix. A flaky test inside a loop can send the agent after a failure its change did not cause.
How does RunStory help with AI bug-fix loops?
A fix loop is hard to see when the agent that made a fix is the only one that checks it. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
Should you start a fresh session when a fix loop stalls?
A fresh session often helps when a fix loop stalls, because the old session holds each failed attempt in its context. Start the fresh session with a short prompt that lists the findings, their failing tests, and what each earlier attempt broke.
Does reverting the agent's last change help?
Reverting the agent's last change helps when that change undid a fix that worked, because it returns the code to a known state. Reverting alone does not end the loop. The next attempt needs a kept test for each finding, or it can make the same change again.
How many fix attempts should an agent get?
The number of fix attempts an agent gets is a team's choice. The best practices guide for Claude Code advises a fresh session after 2 failed corrections on the same issue. When fixes keep failing, a person reads the failing tests and decides whether the findings conflict.
Why does an AI review loop keep finding new problems?
An AI review loop keeps finding problems because a reviewer asked to find gaps usually reports some, even when the work is sound. Asking it to flag only gaps that affect correctness or the stated requirements helps, and ending each round on passing tests gives the loop a stopping point.