Why do coding agents fix symptoms instead of root causes?
AI coding agents often fix symptoms instead of root causes because the error names the line where a failure showed, not the code that caused it. The agent edits that line until its check passes, and the passing check ends the task. The code that produced the bad value stays as it was.
In debugging, a symptom is the failure that someone observes, e.g. a crash. The root cause is the condition whose removal stops that failure from happening again. A symptom fix is a change that stops a failure where it shows and leaves the defect that caused it in the code. Root cause analysis is the method that traces a failure back to that cause.
The failure can then return in another form, often on a path the fix never touched. When each fix moves the failure somewhere else, the session can turn into an AI bug-fix loop, which developers often call "whack-a-mole" fixing.
What causes symptom fixes?
Four conditions pull an agent's edit toward the line where the error shows.
Why does the error point to the wrong place?
A stack trace lists the function calls that were active when the error happened. The function that created the bad value has often returned already, so the trace does not name it. An agent given only the error output starts from the files the trace names, and the cause often sits in a file the trace never names.
Why does a passing check end the work?
A coding agent works in a loop that edits code, runs a check, and reads the result. The loop usually stops when the failing check passes. A guard at the failing line passes any check that looks only for the error, e.g. a test that asserts the page renders without throwing.
A 2025 study of AI patches on SWE-bench, a benchmark built from real GitHub issues, found that 29.6% of "passing" patches behaved differently from the real fix. Those patches passed SWE-bench's tests, as the real fix did, so those tests could not tell them apart.
Why is the edit at the error the smaller change?
A guard at the failing line changes one line in a file the agent has already read. A fix at the cause can mean reading callers and changing shared code. A model rewarded for passing tests earns the same for a guard as for a fix at the cause. Taken further, that pressure is reward hacking.
Why does the bug report leave out the expected result?
A bug report that holds only an error message says what went wrong, not what should have happened. The rule that the bad value broke is often written nowhere, e.g. that changing the address keeps the cart. Without that rule, the error is the only target.
What does a symptom fix look like?
Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Let customers edit their delivery address during checkout." After that change, the order summary page crashes, and the developer sends the agent the error:
TypeError: Cannot read properties of undefined (reading 'length')
at renderSummary (/srv/acme/src/checkout/summary.js:18:32)
at handleSummary (/srv/acme/src/checkout/routes.js:42:10)
The agent reads summary.js, the top file in the trace, and changes line 18:
- const count = checkout.items.length;
+ const count = checkout.items?.length ?? 0;
It then adds a test that checks for the crash:
test("summary renders after an address change", () => {
const checkout = setAddress(cartWithItems(2), "12 Elm Street");
expect(() => renderSummary(checkout)).not.toThrow();
});
The test passes, and the agent's summary says "Fixed the crash on the order summary page." The page now loads and shows an empty order. The address saved, but the cart emptied.
The address handler, setAddress(), still replaces the whole checkout record, and the trace never named it. The first change was one way agents break working features, and the guard hides that break from any check that looks only for errors.
Symptom fixes from coding agents often share these signs:
- The session stays inside the trace. The agent reads and changes only the files the stack trace names and never opens the code that made the value.
- The error becomes a normal return. A default, e.g.
?? 0, or a catch block that returns an empty result makes the check stop failing in one edit. - The agent's own test checks only the crash. The test arrives with the fix and asserts that nothing throws, not the result a customer needs.
- The summary names the error, not the cause. It says the crash stopped, not why the value was wrong.
- A new failure is reported as old. When the fix breaks another path, the agent's summary can describe that failure as a gap that was already there.
This example is simplified. A real session often buries the guard among other changes.
How can teams push the agent toward the cause?
These practices each close one of the gaps above:
- Send the steps and the expected result. A report with reproduction steps, the observed result, and the expected result names the behavior to restore, not only the error to remove.
- Ask where the value was created before any edit. In plan mode, which blocks file edits, a prompt can say "Name the line that left
checkout.itemsundefined and how the record reached line 18." Check that answer against the code before accepting it. - Name the shortcuts that do not count. Instructions can say that a guard or a catch block at the failing line needs a stated cause. Instructions alone are a weak control, so they belong next to the checks below.
- Write the check from the expected result. A test that checks the end result, e.g. the items left in the cart, fails on a symptom fix. Keep it out of the agent's edits, because changing the test to pass is test tampering.
- Review the diff against the trace. In code review, a fix whose changes sit only in the files the trace names needs a written cause. The reviewer asks where the value was created and reads that code.
None of these practices shows that the cause is gone. They make a fix at the cause more likely and a symptom fix easier to spot.
How is a symptom fix different from a workaround?
A workaround is a deliberate change that avoids a known, recorded problem while the real fix waits, e.g. hiding the address form until the handler is fixed. A symptom fix stops the failure without anyone recording the problem, so nothing tracks the defect that remains.
The same code can be either one. A guard at line 18 is a workaround when a ticket records that the handler replaces the record. It is a symptom fix when the agent's summary says only that the crash is fixed.
A workaround is a reasonable choice to stop harm while the real fix is made. It also fits when the cause sits in code the team cannot change, e.g. a third-party library.
How do you show that a fix reached the cause?
Use a check that the symptom fix alone would fail. The agent's crash test fails on the old code and passes with the guard, so a fail-before, pass-after check on that test says nothing about the cause.
A test that expects 2 items in the cart after the address changes fails on the old code and with the guard alone. It passes only once the handler is fixed. Then verify a bug fix in the running software.
A symptom fix often shows up when someone reruns the failed workflow from its first step and checks the result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
Why do agents wrap failing code in error handlers?
Agents wrap failing code in error handlers because a catch block turns a thrown error into a normal return in one edit. The check that reported the error then passes, and the loop stops. The caller receives an empty result instead of an error.
Is a workaround ever the right fix?
A workaround is the right fix when harm has to stop before the real fix is ready. A workaround also fits when the cause sits in code the team cannot change, e.g. a third-party library. Either way, the team records the problem, so the workaround does not become a hidden symptom fix.
How do you ask a coding agent for the root cause?
You ask a coding agent for the root cause by requesting, before any edit, the line that created the wrong value and the path it took to the error. Then check that answer against the code and a test of the expected result.
Why do symptom fixes pass the tests?
Symptom fixes pass the tests when the tests check that the error is gone, not that the result is right. A test that asserts the page renders without throwing passes once a guard hides the crash, while the wrong value still reaches the page.