# What is verification debt?

Verification debt is the growing gap between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define verification debt
-   Explain how it builds up as agents speed up output
-   Compare it with technical and comprehension debt

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [What is AI slop in code?](https://specstory.com/learning/verification/ai-slop)
-   [What is the difference between reading code and running code?](https://specstory.com/learning/verification/reading-vs-running-code)
-   [What is a definition of done for coding agents?](https://specstory.com/learning/verification/definition-of-done-for-coding-agents)

## What is verification debt?

Verification debt is the backlog of AI-written code that a team produces faster than anyone can check it by running or testing it. It is also called AI code verification debt. The code may pass its own tests and get approved, but nobody has run it to check that it does what was asked.

The debt grows when [coding agents](https://specstory.com/learning/ai-coding/coding-agent) produce changes faster than people can run them. The same gap can form with code that people write, but agents widen it by raising the volume of changes. A failure found weeks later is harder to trace, because more changes sit on top of its cause.

The term is young and has no settled definition. Some writers use it for code not checked against quality standards, and call the gap in checks against users' needs validation debt. In this learning center, [verification](https://specstory.com/learning/verification/verification-vs-validation) includes running the software, so the debt is a gap in evidence about behavior.

More developers distrust the accuracy of AI tools (46%) than trust it (33%), and only 3% trust it highly ([Stack Overflow, 2025](https://survey.stackoverflow.co/2025/ai)). Distrust is not a check, so the debt shrinks only when teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing) and keep the evidence.

## How does verification debt build up?

Verification debt builds up one change at a time, in these steps:

1.  A developer asks a coding agent for a change.
2.  The agent writes the code and often its tests, from the same prompt.
3.  The tests pass, and the agent's output says the task is "done."
4.  A reviewer or an [AI code review](https://specstory.com/learning/code-review/ai-code-review) tool reads the diff, and [continuous integration](https://specstory.com/learning/ci-cd/continuous-integration) (CI) runs the test suite.
5.  The change merges, although nobody has started the app and used the workflow it touched.
6.  The next change builds on the unchecked one.

Diagram: Where verification debt builds up

Tests from the same prompt and a read diff can miss what the changed workflow does. Each change that merges without a run adds to the debt, and later changes build on it.

Steps 3 and 4 produce evidence about the code as written. Tests from the same prompt tend to share the code's blind spots, and a diff shows what changed, not what the running app does. When unchecked output is also bloated or duplicated, reviewers call it [AI slop](https://specstory.com/learning/verification/ai-slop).

The debt grows fastest when output rises. [GitHub says](https://github.blog/news-insights/company-news/the-august-17-outage-and-the-work-ahead/) monthly commits on its platform roughly doubled, from 1.4 billion to 2.9 billion, between April and August 2026. The post does not say how many of those commits came from coding agents.

Testing was rarely an agent session's main task in one large sample. In [Anthropic's analysis](https://www.anthropic.com/research/claude-code-expertise) of about 400,000 Claude Code sessions, only about 5% were mainly about testing and orchestrating code. In the same analysis, about 25% of sessions were mainly about writing code and about 26% about fixing it.

When agent changes arrive faster than people can read them, a [review bottleneck](https://specstory.com/learning/code-review/review-bottleneck) forms. Teams with many agent changes often find [CI becoming the bottleneck](https://specstory.com/learning/ci-cd/ci-for-coding-agents) too. Under that pressure, teams skim reviews or cut slow checks, so more changes merge without a run.

## What is an example of verification debt?

Here is an illustrative example. Acme Co. sells furniture online, and its developers hand most changes to coding agents. Acme counts its verification debt as the merged changes to customer workflows that nobody has run since the merge. One week goes as follows:

1.  Coding agents make 30 changes, and all 30 merge.
2.  Each change passes its tests in CI, and the agents wrote most of those tests.
3.  Reviewers read each diff. One change handles the request "Let customers edit their delivery address during checkout."
4.  Nobody starts the store and uses checkout after that change.
5.  At the end of the week, a developer finds 9 changes that touched a customer workflow, and 2 of them were run in the app.
6.  The developer runs checkout in a copy of the store, adds 2 items to the cart, and changes the address to 12 Elm Street. The address saved, but the cart emptied.
7.  The developer sends the steps to the agent. After the fix, the same steps keep 2 items in the cart.

After the checkout run, the list reads:

```text
Changes merged this week:                   30
Changes that touched a customer workflow:    9
Of those, changes run in the app:            3
Verification debt (changes not run):         6
```

The checkout run paid down one item and found a bug that the tests and the review missed. This example is simplified. A real team would also count changes that reach a workflow through shared code, e.g. a cart helper.

## What are the limits of the idea?

The idea names a real gap, but it has limits:

-   **No standard measure exists.** Teams count proxies, e.g. merged changes whose workflows nobody has run. [Code coverage](https://specstory.com/learning/test-quality/code-coverage) is a weak proxy, because it measures the code that tests ran, not the results they checked.
-   **Checked is not the same as correct.** A run covers only the workflows someone tried, and no findings is not the same as complete coverage.
-   **Unchecked changes carry different risks.** An unchecked change to checkout costs more than one to an internal script, so a plain count overstates some debt and understates the rest.
-   **The data on the gap is indirect.** The GitHub and Anthropic figures count commits and session topics. Neither counts the changes that merged without a run.
-   **The debt does not reach zero.** Software has more inputs and states than any team can run, so the aim is a small debt on the workflows customers use.

## How is verification debt different from technical debt?

[Technical debt](https://specstory.com/learning/code-review/technical-debt) is the future cost of shortcuts in code or design. Verification debt is missing evidence about what the code does. A clean design can still carry verification debt if nobody has run it. The two overlap when a team skips a check to merge sooner, so some writers count verification debt as a kind of technical debt.

Comprehension debt is code that no one on the team understands. It hides verification debt, because it is harder to list the workflows a change touched when nobody understands the change.

The three debts differ in what is missing:

| Debt | What is missing | How a team pays it down |
| --- | --- | --- |
| Technical debt | A design that is cheap to change | Cleaning up the shortcuts |
| Comprehension debt | People who understand the code | Reading and explaining the code |
| Verification debt | Evidence of what the code does | Running and testing the changed behavior |

## How do teams pay down verification debt?

Teams pay down verification debt by checking behavior as each change is made. These habits help:

-   **Write the checks before the agent starts.** A [definition of done](https://specstory.com/learning/verification/definition-of-done-for-coding-agents) lists the evidence a change needs, e.g. a run of the changed workflow, before the change counts as finished.
-   **Keep changes small.** [Pull request size](https://specstory.com/learning/code-review/pull-request-size) sets how closely reviewers read a change, and a small change touches fewer workflows.
-   **Review against the request.** When teams [review a pull request](https://specstory.com/learning/code-review/reviewing-agent-pull-requests) from a coding agent, they compare it with what was asked, not only with the diff.
-   **Run the changed workflows.** Start the app in a separate environment and use the workflow the change touched. [Running code](https://specstory.com/learning/verification/reading-vs-running-code) shows failures that reading can only predict.
-   **Keep a list of changes nobody has run.** Record each one when it merges, and start with the workflows customers use most.
-   **Keep each failure as a test.** A failure found in a run becomes a test, so the suite checks that behavior after later changes.

## How does RunStory help with verification debt?

Verification debt grows when changes merge faster than anyone runs them. Coding agents outpace testing, even with AI review.

[RunStory](https://specstory.com/runstory) runs your software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures back to your coding agent while you keep working. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### How do you measure verification debt?

Verification debt has no standard unit, so teams measure it with proxies. A common proxy is a list of merged changes whose workflows nobody has run. Teams then rank that list by how many customers use each workflow.

### Does AI code review reduce verification debt?

AI code review does not pay down verification debt on its own, because it usually reads the diff instead of running the changed workflow. It can catch problems visible in the code. A review that runs the project's tests checks only what those tests check, so the workflow stays unchecked until someone runs it.

### Is verification debt only a problem with AI-written code?

Verification debt usually refers to AI-written code, but the same gap can build up with any code that merges faster than people check it. Coding agents make the gap more common because they raise the volume of changes, while the time people have for review and testing grows more slowly.

### Can a team have verification debt with high test coverage?

A team can have verification debt with high test coverage, because coverage measures the code that tests ran, not the results they checked. Tests that a coding agent wrote from the same prompt as the code can run most lines and still miss a broken workflow.

---

Source: [What is verification debt? | AI-generated code | SpecStory](https://specstory.com/learning/verification/verification-debt)
