Skip to content

How to test AI-generated code

Testing AI-generated code means checking what the software does, not who wrote it, by reading the change, running its tests, and running the app itself.

Last updated , 11 min read

Key points

  • A coding agent's "done" is a claim, so check the change with evidence the agent did not write.
  • Collect three kinds of evidence by reading the diff, running the tests, and using the running app.
  • Send each failure back to the agent with reproduction steps, and repeat those steps after the fix.

How do you test AI-generated code?

AI-generated code is tested with three kinds of evidence, from reading the diff, running its tests, and using the app itself. The practice is also called AI code verification. It checks what the code does, not who or what wrote it. A coding agent reporting "done" is making a claim until that evidence backs it up.

Code from an agent often compiles, reads well, and passes the tests the agent wrote, and still breaks when someone uses it. In Stack Overflow's 2025 survey, the most common problem developers reported with AI tools was answers that are almost right but still wrong. Of the developers who answered that question, 66% had run into it.

The same checks apply when nobody reads the code closely, as in vibe coding. For an app built that way, the steps below form a basic quality assurance (QA) process.

What is AI code verification?

AI code verification is the practice of checking what AI-generated software does, by reading it, running its tests, and running the software itself. It replaces trust in the output with evidence, and that trust is low. More developers distrust the accuracy of AI tools (46%) than trust it (33%), and only 3% trust it highly (Stack Overflow, 2025).

In this learning center, "verification" means any check that produces evidence about software, including running it. The International Software Testing Qualifications Board (ISTQB) defines verification more narrowly, as confirming that a work product fulfills its specification.

How does verifying AI-generated code work?

Verification compares a change with the request, using evidence from outside the agent's session. A change comes back from the agent with code, tests, and a "done" claim. Three checks read or run it, and each failure goes back to the agent with reproduction steps.

Three checks on a change from a coding agent Request Coding agent Change code and tests "done" claim Read the diff Run the tests Run the software Evidence not a claim failure with reproduction steps
The agent's "done" is one input to the checks, not their result. Each check reads or runs the change, and failures return to the agent.

Each check needs a test oracle, the source of the expected result, that the agent did not write. A test that the agent wrote from the same prompt tends to repeat its mistakes.

What are the types of evidence that code works?

Evidence that code works comes in three kinds. NIST's Secure Software Development Framework treats reviewing code (PW.7) and testing executable code (PW.8) as separate practices for finding vulnerabilities, and suggests running tests in a sandboxed environment.

Reading the change

Reading the change means inspecting the diff without running it, by a person or an AI code review tool. It shows what the agent changed, including tests it deleted and edits the request did not ask for. It cannot show what the code does at runtime. Static analysis tools, e.g. a type checker, also check code without running it and share that limit.

Running the tests

Running the tests executes the project's automated checks and compares each result with an expected value. The result counts only when a test runner produces it, not when the agent's summary reports it.

Running the software

Running the software starts the app, uses it the way a customer would, and checks what happened. The ISTQB calls testing that executes the software dynamic testing. It gives the most direct evidence of what a customer gets, e.g. whether the cart keeps its items.

What is an example of testing AI-generated code?

Here is an illustrative example. A developer at Acme Co., which sells furniture online, asks a coding agent to "Let customers edit their delivery address during checkout." The agent edits the checkout code, adds one test, and reports "done." The developer then collects the evidence:

  1. The diff changes 2 checkout files and adds one test, and nothing in it looks wrong.
  2. The tests run in a clean copy of the repository, and 48 of 48 pass.
  3. The developer starts the store in a separate environment, adds 2 items to the cart, starts checkout, changes the address to 12 Elm Street, and continues to payment.
  4. The address saved, but the cart emptied.
  5. The developer sends the agent the steps and results below, and the agent changes the code.
  6. The same steps on the fixed code leave 2 items in the cart.
Steps:    add 2 items to the cart, start checkout,
          change the delivery address to 12 Elm Street,
          continue to payment
Expected: 2 items in the cart on the payment page
Actual:   0 items in the cart on the payment page

Only running the software found the bug. The agent's test checked the address but not the cart. This example is simplified. A real checkout would need more checks, e.g. one for the order total.

How do you test a change a coding agent made, step by step?

These steps apply to one change or to a whole feature.

1. Write down what "done" means before the agent starts

Turn the request into checks a person could run, e.g. "Changing the address keeps the items in the cart." Keep them in files the agent does not change, and review any edit to them.

2. Read the diff and every changed test

Compare the diff with the request, file by file. Look for tests that were deleted, skipped, or loosened, and for edits the request did not ask for. Read the agent's record of the commands it ran, and treat its summary as a claim.

3. Run the tests yourself

Run the full suite in a clean copy of the repository, outside the agent's session, e.g. with npm test. Check that each added test asserts a result, not only that the code ran.

4. Start the app and use the changed workflow

Start the software in a separate environment, e.g. a fresh container, and use the workflow the request named. Compare the results with the checks from step 1. For a command-line interface (CLI), run the changed commands and check their output and exit code. A fresh environment also exposes packages that only the agent's machine had.

Script the workflows customers use most as end-to-end tests, e.g. with Playwright, so they rerun after later changes without a person.

5. Try the workflows next to the change

Agents often edit shared code, so a change to checkout can break the cart or the order history. Rerun the workflows that passed before the change.

6. Send each failure back with reproduction steps

Give the agent the actions you took, the expected result, and the actual result. Keep each failure as a test, so the suite grows where bugs appear.

7. Rerun the failing check after the fix

Repeat the exact steps that failed, on the fixed code, and then repeat step 5. Together, the two runs are how teams verify a bug fix.

What are the limits of each kind of check?

Each check catches some agent failures and misses others:

Agent failureCaught byMissed by
A test edited or deleted so the suite passesReading the diffRunning the tests
Tests that assert almost nothingReading the tests, or mutation testingRunning the tests
A "done" report with no passing runRunning the tests yourselfReading the agent's summary
A broken workflow next to the changeRunning the softwareTests that do not cover it
A package that only the agent's machine hadA fresh environmentThe agent's own session
An edit the request did not ask forReading the diffRunning the software
An added dependency that is not the intended packageChecking each added package in the registryRunning the tests

A 2026 study of 86,156 changes to test files by five coding agents found that 80.2% had weak or no explicit oracle signals. Many of those tests ran code while checking little or nothing about its output.

A reviewer can approve code that is almost right. Running the software covers only the workflows someone tried, and no findings is not the same as complete coverage. None of the three checks is a security review or establishes that code is correct, but together they leave fewer gaps.

How is testing AI-generated code different from detecting it?

AI code detection estimates whether a model wrote a piece of code, usually from the text of the code. Testing checks what the code does when it runs, whoever wrote it. A detector's answer says nothing about whether the code works, and code that a detector flags can be correct. Detection serves policy questions, e.g. whether a contributor disclosed AI use.

What does verification cover?

The verification articles each cover one part of checking code that an agent wrote:

  • False completion claims happen when a coding agent reports "done" before any run shows that the work succeeds.
  • A definition of done lists the evidence a change needs before it counts as finished.
  • Reward hacking is an agent satisfying the check it is graded on without doing the intended task.
  • Test tampering is an agent editing, skipping, or deleting a failing test so the suite passes.
  • AI-written tests often pass because they assert what the code already does.
  • A false pass is a check that reports success while the software is broken.
  • AI slop is output from AI tools that looks finished but is bloated or unchecked.
  • Vibe coding bugs are recurring failures in apps built mostly by AI.
  • Runtime verification checks a running program against stated properties by observing it.
  • Reading code finds problems in the source, while running code finds problems that appear only at runtime.
  • Verification and validation check software against its specification and against its users' needs.
  • AI hallucinations in code are references to APIs or packages that do not exist.
  • LLM-as-a-judge uses a language model to grade output against written criteria.
  • Verification debt is the backlog of AI-written code that nobody has run or tested.
  • AI self-verification is an agent checking its own output, which is weak evidence that the code works.

How does RunStory help with testing AI-generated code?

A coding agent's summary and its own tests cannot stand in for running the software. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results.

Your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

How do you verify an agent's changes without testing by hand?

Verifying an agent's changes without testing by hand means writing the checks once as automated tests and running them after each change. Write them from the request, script the main workflows as end-to-end tests, and run them against the app in a separate environment.

Do you still need to read every line an agent writes?

Reading every line an agent writes is one kind of evidence, and it slows down as agents write more code. Reviewers can focus on the changed tests and on edits outside the request, then run the software to check behavior that reading cannot show.

How do you test a large AI-generated app?

Testing a large AI-generated app starts with the workflows customers use most, not with the whole codebase. Run those workflows in a separate environment after each change, and keep each failure you find as a test, so the suite grows where bugs appear.

How do you check what an agent did after a prompt?

Checking what an agent did after a prompt starts with the diff and the agent's record of the commands it ran, not with its summary. Then rerun the tests yourself and try the changed workflow in the running app.

Is there a QA process for vibe-coded software?

A QA process for software built by vibe coding uses the same three kinds of evidence as any AI-generated code. Write down what the app must do, run the tests, and use the main workflows in the running app after each change.