Skip to content

Code review checklist for AI-generated code

A code review checklist for AI-generated code lists what to check when a coding agent wrote the change, from scope and test edits to a record of a run.

Last updated , 9 min read

What should a code review checklist for AI-generated code include?

A code review checklist for AI-generated code covers the ways an agent's change can fail. This one has 9 checks in 3 groups. Intent and scope checks compare the change with the task. Test and dependency checks look for weakened tests and invented packages. Runtime and failure checks ask for a run record and read error and access handling.

A classic code review checklist covers design and readability. Most of the items below apply to any change. They weigh more when a coding agent wrote the change, because the change may reach review before any person has read or run it.

Reviewing agent pull requests sets out the review steps in order. On GitHub, a team can save the list as pull_request_template.md in the repository root, docs/, or .github/. Once the file is on the default branch, its text fills the body of a pull request opened in the browser, but not one opened through the API:

Intent and scope
- [ ] Each changed file serves the task, or has a stated reason.
- [ ] The change covers each part of the task, checked against the request.
- [ ] A named person owns the change and can explain each file.

Tests and dependencies
- [ ] No test was deleted, skipped, or weakened, and no check was turned off.
- [ ] Each added test fails if the behavior it names breaks.
- [ ] Each added package exists, is needed, and passes an audit.

Runtime and failure
- [ ] A record shows the changed workflow running after the last edit.
- [ ] Errors stop the action or are reported, not swallowed.
- [ ] Access checks run on the server, and no secret is in the diff.

The items below use one illustrative example. Acme Co. sells furniture online, and a coding agent opened pull request #1042 for the request "Let customers edit their delivery address during checkout."

Which intent and scope checks come first?

These checks need only the task and the file list.

Does each changed file belong to the task?

Agents often edit files the task never named. Treat each such change as an out-of-scope edit until someone gives a reason. The Open Worldwide Application Security Project (OWASP) secure coding with AI cheat sheet calls missing these edits "the most common review failure mode in agentic coding."

List the files with git diff --name-only --no-renames origin/main...HEAD before reading them. Ask for a split when the list is too long to check, because pull request size sets how closely a reviewer reads. At Acme, the list includes src/lib/session.ts, which the request did not name.

Does the change do all of what the task asked?

The pull request description is text the agent generated, so compare the change with the request instead. Write the request as checks, e.g. "the cart keeps its items after an address change." Then find each check in the diff and the tests. A missing requirement draws no comment unless someone looks for it. At Acme, the summary says "Tests pass," but no test still checks the cart's item count.

Can a person explain each change?

The same cheat sheet states that "AI-generated code must have a human owner." The owner, the developer who accepts the change, answers review questions. The review waits until the owner can explain each changed file. At Acme, the owner has to say why the session code changed.

Which test and dependency checks come next?

An agent can pass a suite by editing the tests, so read these files before the application code.

Was any test or check weakened?

Start with deleted tests and edited assertions. A weakened assertion can pass on broken code, e.g. toBeDefined() in place of toHaveLength(2). That edit is test tampering.

Checks outside the tests can be weakened too. A disable comment or an edit to the linting configuration turns a rule off, and a changed continuous integration (CI) file can drop a job. An edited agent instruction file changes the agent's later runs too. At Acme, cart.test.ts now expects only that the items are defined.

Do the added tests check the result?

Tests that an agent writes from the same prompt as its code often check that the code ran, not what it returned. For each added test, ask which change to the code would make it fail. Assertion coverage measures how much behavior the tests check, not how many lines they run.

A batch of similar tests is a sign of test bloat, which adds review time without catching more bugs. At Acme, the added address test would still pass with an empty cart.

Are the added packages real and needed?

A model can generate an import for a package that does not exist, one of the AI hallucinations in code. An attacker can publish a package under that name. Look up each added name in its registry before anything installs it, and run a dependency audit, e.g. npm audit.

Changes to the lock file should match the added packages. Install scripts in package.json run when someone installs the project, so an edit to them needs a reason. At Acme, no package was added, so this check passes.

Which runtime and failure checks close the list?

Reading shows only part of what these checks ask.

Is there a record of the change running?

An agent's "done" is a claim. Evidence that code works is a record of a check that ran on a stated version of the code, with what the software did. The record should name the version after the agent's last edit, because a run before that edit describes older code.

A green CI check is evidence about the tests, which may be the ones weakened above. In this illustrative example, Acme's pull request has no record, and one run of the checkout shows the failure. The address saved, but the cart emptied.

Are any errors swallowed?

Generated code often catches an error and carries on. Read each added catch block and each value returned on failure, e.g. an empty list. The code should stop the action or report the error.

A login check that lets the request through after an error is a fail-open check. Swallowed errors and fail-open checks are common bugs in AI-generated code, and a demo rarely triggers them. At Acme, the diff adds no catch block, so nothing is flagged.

Do access checks and secrets stay on the server?

A prompt rarely says who may use a feature, so generated code can check access only in the page. Look for a permission check on the server for each route that reads or changes a customer's data. Look for keys in source files and in code that ships to the browser. These checks are a small part of AI-generated code security.

At Acme, the address route checks on the server that the customer owns order A-1042, so this check passes.

What does this list leave out?

Applied to pull request #1042, the list flagged the session file, the weakened cart test, the address test that ignores the cart, and the missing run record. A finished list still leaves these out:

  • Design and readability. Whether the code fits the system still needs a person who knows the codebase.
  • A full security review. OWASP's secure code review cheat sheet has its own checklists by area, e.g. authorization.
  • Performance and accessibility. A diff shows little about load or screen readers.
  • An independent reader. An AI code review tool can read the diff for several items. The session that wrote the change reviews from the same context, so it tends to repeat its own misses.
  • Proof that the software works. A checked box records that someone looked. No findings is not the same as complete coverage.

The order of the list is not a measured ranking. Checks that a tool already runs, e.g. formatting, belong in the linter and CI, which keeps the list short.

How do you check the parts of the list that reading cannot?

Run the branch after the agent's last edit. Use the requested feature and the workflows next to it, e.g. the cart. Then make a dependency fail and sign in as a second customer, which covers the error and access items.

RunStory's private alpha tests your software in isolated sandboxes. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. The alpha starts with CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Should the checklist differ for code a person wrote?

A checklist for code a person wrote can keep most of the same items, next to a classic list's design and readability checks. Those items weigh less when the author has read and run the change.

Can a coding agent run the checklist on its own pull request?

A coding agent can run the reading items on its own pull request, e.g. listing the changed files. The agent cannot be the named owner, and in the session that wrote the change, its review starts from the context that produced the change. A person and a run cover the rest.

How long should a review checklist be?

A review checklist should be short enough to fit in a pull request template. The list above keeps 9 items by leaving checks that a tool already runs, e.g. formatting, to the linter and CI.

Where should a team keep its review checklist?

A team can keep its review checklist in the repository as a pull request template. GitHub fills a pull request opened in the browser with the template's text, but not one opened through the API.