# Why do AI coding agents break working features?

Coding agents break working features when edits reach beyond the task, fix symptoms in the wrong place, or change code the tests do not cover.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Explain the main ways agents break existing features
-   Apply checks that protect behavior outside the diff
-   Identify out-of-scope edits and their blast radius

## Related content

-   [What is debugging?](https://specstory.com/learning/debugging/debugging)
-   [What is git bisect?](https://specstory.com/learning/debugging/git-bisect)
-   [How to write a bug report a coding agent can fix](https://specstory.com/learning/debugging/bug-report-for-coding-agents)
-   [How to verify a bug fix](https://specstory.com/learning/debugging/bug-fix-verification)

## Why do AI coding agents break working features?

[Coding agents](https://specstory.com/learning/ai-coding/coding-agent) break working features when their changes reach code that other features depend on, and the agent's checks do not cover those features. The change is often an edit the task never asked for, or a fix placed where an error showed instead of where it started. The existing feature stops working, and often no test fails.

A feature that worked before a change and fails after it is a [software regression](https://specstory.com/learning/testing/regression-testing). On the [SWE-CI benchmark](https://arxiv.org/abs/2603.03823), most models kept a codebase free of regressions in fewer than a quarter of long-running maintenance tasks.

Finding the change behind a broken feature is a [debugging](https://specstory.com/learning/debugging/debugging) task. Regressions are also one of the common [bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs), because a demo tries only the requested feature.

## What causes agents to break working features?

Four causes are common.

### Why do agents edit code outside the task?

An [out-of-scope edit](https://specstory.com/learning/glossary#out-of-scope-edit) is a change the task did not ask for, e.g. an unrequested refactor. A run of such edits is often called scope creep. The [Secure Coding with AI](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Coding_with_AI_Cheat_Sheet.html) Cheat Sheet from the Open Worldwide Application Security Project (OWASP) says agents "routinely touch files beyond the scope of the requested change."

A coding agent often stops once its own checks pass, and those checks do not measure how far the diff spreads. Each extra edit widens the [blast radius](https://specstory.com/learning/glossary#blast-radius), the set of features, users, and systems that could break because of the change.

### Why do fixes land where the error shows?

An error often shows up far from the code that caused it. An agent's fix can land on the line in the error output instead, e.g. a null check. The bad value is still created upstream, and other features now receive it with no error. Loosening a shared type to pass a type check works the same way. [Root cause analysis](https://specstory.com/learning/debugging/root-cause-analysis) traces a failure back to where the wrong state starts.

### Why does a change reach features the task never named?

Features share code, e.g. a checkout record that the cart page and the address form both read. A change made for one caller reaches the others too. A coding agent often reads only the files its search returns, not every caller of a function it changes. Many rules that callers depend on are also written nowhere, e.g. that changing the address keeps the cart.

### Why do breaks pile up over a long session?

A coding agent can make many changes before anyone runs the software. Later changes build on a broken one, so reverting the session removes valid work too. [Git bisect](https://specstory.com/learning/debugging/git-bisect) finds the first commit where the feature fails, if the session's changes were committed in steps that can each be tested.

## What does a broken working feature look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's session runs in five steps:

1.  The agent adds an address form and a handler in `src/checkout/address.ts`.
2.  The handler fails a type check, because the shared `saveCheckout()` function expects the whole checkout record.
3.  The agent changes `saveCheckout()` in `src/checkout/store.ts` to accept part of a record.
4.  The agent also renames a helper in `src/cart/discounts.ts`, which the task never mentioned.
5.  The agent adds an address test, the suite passes, and its summary says "done."

The diff lists 4 files:

```text
 src/cart/discounts.ts          |  6 +++---
 src/checkout/address.ts        | 24 ++++++++++++++++++++++++
 src/checkout/store.ts          |  4 ++--
 tests/checkout/address.test.ts | 12 ++++++++++++
 4 files changed, 41 insertions(+), 5 deletions(-)
```

A developer then adds 2 items to the cart and changes the address. The address saved, but the cart emptied. The changed `saveCheckout()` now replaces the whole stored record with the address alone. The cart page and the order summary read their items from that record, and neither is in the diff.

Diagram: What the diff shows and what the change reaches

The diff names three source files. The change to the shared store reaches three features, and the agent's test checks only one of them.

The out-of-scope rename in `discounts.ts` broke nothing. The cart broke through a type fix inside the task, placed in shared code instead of in the handler. This example is simplified. A real codebase would have more callers of `saveCheckout()`.

## Why do the tests stay green?

In [a 2026 study](https://arxiv.org/abs/2601.15195) of 33,596 pull requests opened by coding agents, failing continuous integration (CI) or tests was the most common reason for rejection at the code level. The study counted it in 17.6% of the rejected pull requests it analyzed.

A suite can still stay green after the kind of break Acme had, for four common reasons:

-   **The agent's test checks the task.** The agent generated the test from the same prompt as its code, so the test checks the address and not the cart.
-   **No older test covers the broken path.** Acme's cart tests never change the address with items in the cart.
-   **Test selection skips unlinked tests.** [Test impact analysis](https://specstory.com/learning/ci-cd/test-impact-analysis) runs only the tests its map links to the changed files, so a browser test that imports no source file can be skipped.
-   **A failing test can be edited to pass.** The agent can change an older test instead of the code, which is [test tampering](https://specstory.com/learning/verification/test-tampering).

A green suite is evidence about the tests it contains, not about the features that no test runs.

## How can teams protect behavior outside the diff?

Checks protect other features best when the agent's edits cannot change them. These practices each cover one gap:

-   **Name the paths the task may change.** State the directories in the task, e.g. `src/checkout/`. Some agents also take path rules, e.g. an `Edit(src/cart/**)` deny rule in Claude Code's settings, but such rules do not cover every way a process writes a file.
-   **Keep each change small.** A limit on [pull request size](https://specstory.com/learning/code-review/pull-request-size) keeps each diff short enough to read in full. Run the software after each change and commit it once it works, before the next change builds on it.
-   **Record behavior before the change.** [Golden file testing](https://specstory.com/learning/test-quality/golden-file-testing) saves what a feature does now, so a test fails when a later change alters it.
-   **Protect the older tests.** Run the regression tests from the main branch, and send any change to a test file to a person.
-   **Review each changed file against the task.** In [code review](https://specstory.com/learning/code-review/code-review), each file needs a reason, and shared code needs a look at its callers. Listing the changed files is also the first step to [review a pull request](https://specstory.com/learning/code-review/reviewing-agent-pull-requests) from an agent.
-   **Run the features next to the change.** At Acme, that means adding 2 items and then changing the address. When a feature breaks, send the agent the failing steps and the observed result.

OWASP also advises a CI check that flags changes outside the requested scope:

```bash
# Fail when the branch changes files outside the task's paths
changed=$(git diff --name-only --no-renames origin/main...HEAD) || exit 1
outside=$(printf '%s\n' "$changed" \
  | grep -v -E '^(src/checkout/|tests/checkout/)' || true)
if [ -n "$outside" ]; then
  printf 'Changed outside the task:\n%s\n' "$outside"
  exit 1
fi
```

On Acme's branch, the step flags `src/cart/discounts.ts` but not `src/checkout/store.ts`, the file that broke the cart.

None of these practices covers a feature that no test, rule, or run reaches.

## How is a broken feature different from an out-of-scope edit?

An out-of-scope edit is a change in the diff that the task did not ask for, and reading the diff finds it. A broken feature is a failure in the running software, and only running the feature finds it. The features that read what a change touches make up its blast radius, and those are the ones to run. At Acme, that means the cart page and the order summary, not only the address form.

## How does RunStory help with features an agent broke?

A feature an agent broke is often one that nobody ran after the agent's edit. What passed before needs to remain true, unless someone explicitly decides otherwise. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### How do you tell an agent which files it may change?

You tell an agent which files it may change by naming the paths in the task and, where the agent supports them, adding path rules to its settings. A CI step that checks the changed paths confirms the result.

### Do smaller tasks reduce broken features?

Smaller tasks often reduce broken features, because each diff stays short enough to read in full and each change can run before the next one builds on it. A small change to shared code can still break a feature elsewhere.

### Why does Git show what the agent touched but not what it broke?

Git shows what the agent touched but not what it broke, because a diff records changed lines and not the behavior of code that depends on them. A broken feature often lives in a file the diff never lists, e.g. a cart page that reads a changed record.

### Should the same agent fix the feature it broke?

The same agent can fix the feature it broke when it receives the failing steps and the observed result, not only a report that something broke. The fix counts only after those steps pass in a check that its edits cannot change.

---

Source: [Why AI coding agents break existing features | SpecStory](https://specstory.com/learning/debugging/agents-break-working-features)
