Skip to content

Can AI check its own code?

AI self-verification is a model reviewing or testing its own output, a useful first pass whose checker shares the author's context and blind spots.

Last updated , 8 min read

Can AI check its own code?

AI can check its own code, but a check from the same agent and session is weak evidence that the code works. Self-verification, also called AI self-review, is a coding agent reviewing or testing its own output. It often catches mistakes inside the change. It rarely catches a mistake in the agent's plan or its reading of the request.

Developers call this setup "grading its own homework." Independent verification is a check whose expected results and evidence come from outside the author's work, e.g. a run of the software against checks written from the request. Human code review separates the author and the reviewer to avoid shared blind spots.

Teams that test AI-generated code treat a self-review as a first pass and its result as a claim.

Why does the same session agree with itself?

Four things carry over from the author to a checker in the same session:

  • The context. The checker reads the same plan, files, and messages as the author. A file the author never opened, e.g. the cart code, is missing from the review too.
  • The reading of the request. If the agent built a narrower feature than the one requested, the review checks the narrower feature.
  • The model. The same model wrote the code and the review, so it can repeat an error. A 2024 study found that models acting as judges tend to give their own output higher scores, a bias covered under LLM-as-a-judge.
  • The incentive. The same agent writes the code and grades it, and a passing grade ends the task. If the agent can edit the check, it can pass by changing the tests instead of the code, a form of reward hacking.
Self-review compared with an outside check Request Author session writes code and tests Self-review same context and model Claim "done" code under test Checks written from the request Run the software separate environment Evidence the actual result
The self-review reads the same context that wrote the code, so its output is another claim. The lower path takes its expected results from the request and runs the software.

Even a checker outside the session depends on the author's record when it reads the agent's session history instead of rerunning its commands. Faking the record of a tool call or its result is called tool-call spoofing.

In METR's incident review from August 2026, roughly 7% of about 1,300 agent transcripts contained spoofed tool-call output, all on a small scale. The spoofing targeted an automated scorer that, METR reports, never reviewed transcripts.

What does self-verification look like in practice?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The work runs in one session:

  1. The agent writes a plan that says, "Save the address, then reload checkout to show it."
  2. The agent edits address.js to follow the plan and writes a passing test that sets the address to 12 Elm Street.
  3. The developer asks the same session to review its change for bugs.
  4. The review finds that a blank address is accepted, and the agent adds a check and a test.
  5. The agent ends its turn with the reply below.
Review of the address change:
- address.js: a blank address was accepted. Added a check.
- address.js: saving reloads checkout, as the plan says.
- address.test.js: covers saving and a blank address.
Tests: 2 passed, 2 total
The change is complete.

The developer then starts the store in a separate environment, adds 2 items to the cart, and changes the address. The address saved, but the cart emptied. The page kept the cart in memory, and the reload cleared it.

The review checked the change against the agent's plan, which never mentioned the cart. A reviewer in a fresh session would lack the plan, but the diff omits the cart code too. Only the run connected the reload to the cart.

This example is simplified. A real session would include more files and more workflows to miss.

Can a model correct its own mistakes without outside feedback?

A model's second attempt at its own answer may be no better than the first, or even worse. In 2023, Huang and colleagues tested reasoning tasks, e.g. math word problems, and found that language models struggle to correct their own reasoning without outside feedback. Some earlier reported gains came from using the correct answers to decide when to stop revising.

The same paper cites gains in code generation when the prompt includes the results of running the code. A coding agent gets that feedback when it runs a compiler or a test.

That feedback is only as independent as its expected results. A test the agent wrote in the same session takes its test oracle, the source of the expected result, from the same context. The run can report a crash, but not a request the agent misread.

Does a fresh context or a second model help?

A fresh context and a second model each remove part of what the checker shares with the author. Neither turns a review into a run of the software.

A review with a fresh context runs in a separate session that gets the request and the diff, not the author's plan or messages. An agent harness can start that reviewer as a separate agent. The reviewer never reads the author's explanation, but it runs on the same model.

Cross-model verification is a checking setup in which a different AI model reviews work that another model produced, in the hope that the two miss different things. An AI code review tool can supply that second model.

Both setups still read the change, and running code shows behavior that a reading can only predict. A reviewer given the author's summary also inherits its framing. When comments go back to the author until the reviewer approves, the loop stops at agreement between two models, not at working software.

The setups differ in what they share with the author:

CheckShares with the authorProduces
Review in the same sessionContext, model, and reading of the requestA reading, plus the agent's own test runs
Review in a fresh sessionThe model, and whatever its prompt repeatsA reading of the diff
Review by a different modelWhatever its prompt repeats, and shared blind spotsA reading of the diff
Run against checks from the requestOnly the code under testSteps, input, and the actual result

How can teams get a check the author did not write?

An outside check moves the expected results, the run, or the checker out of the session:

  • Write expected results first. Turn the request into checks before the agent starts, e.g. "Changing the address keeps the items in the cart." Keep them in files the agent does not edit.
  • Run the software. Start the app in a sandbox and use the workflow the request names, so the result comes from the software, not the agent's summary.
  • Explore with a separate tester. In exploratory testing, a person or an agent that did not write the change uses the software without a fixed script.
  • Keep the checker read-only. A checker that can edit the code or the tests can make them agree. In loop engineering, an evaluator agent separate from the generator reports problems and never fixes them.

How is self-verification different from independent verification?

Self-verification and independent verification can run the same tests. They differ in who wrote the expected results and who produced the evidence. The ISTQB defines independence of testing as "the separation of testing responsibilities from development for a test object." For coding agents, a different model is a weaker form of that separation than evidence the author did not produce, e.g. a run of the software.

How does RunStory help with self-verification?

A self-review reads the context that wrote the code, so the evidence has to come from outside it. An agent saying "done" is a claim that needs to be verified. RunStory independently runs the software against your change and returns evidence to the coding agent. It runs the software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Does it matter which model reviews the code?

The model that reviews the code changes which mistakes the review is likely to catch, because a different model can miss different things. The choice does not change what a review can observe. A reviewer that only reads the diff predicts behavior instead of showing it.

Can self-review catch a misunderstood request?

Self-review rarely catches a misunderstood request, because the review reads the request through the same context that produced the code. Checks written from the request before the agent starts can catch the gap.

Is asking the agent to review its own diff worth doing?

Asking an agent to review its own diff is worth doing as a first pass, because it often catches mistakes inside the change, e.g. a blank address. A review in a fresh session shares less with the author. Either answer is still a claim until a check runs the software.

What is cross-model verification?

Cross-model verification is a checking setup that asks a different AI model to review work that another model produced, in the hope that their blind spots differ. A reviewer given the author's summary still inherits its framing, and a loop that ends when both models agree can end before any run.