# How to review a pull request written by a coding agent

Reviewing an AI-generated pull request means checking its task, scope, tests, and evidence of a real run, not only reading the diff.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   List what to check in an agent-written pull request
-   Explain why the diff alone is not enough
-   Apply a review routine that asks for evidence

## Related content

-   [What is code review?](https://specstory.com/learning/code-review/code-review)
-   [What is AI code review?](https://specstory.com/learning/code-review/ai-code-review)
-   [What is the difference between static and dynamic analysis?](https://specstory.com/learning/code-review/static-vs-dynamic-analysis)

## Key points

-   Compare the pull request with the task it came from before reading the code.
-   Read the test changes first, because a weakened test can make a broken change pass.
-   Ask for evidence that the software ran, and send failures back as reproduction steps.

## How do you review a pull request written by a coding agent?

To review an AI-generated [pull request](https://specstory.com/learning/glossary#pull-request), compare the change with the task the [coding agent](https://specstory.com/learning/ai-coding/coding-agent) was given, and check which files it touched. Read the test changes before the code. Then ask for evidence that the software ran, and send any failure back to the agent as [reproduction steps](https://specstory.com/learning/debugging/reproduction-steps).

In [a 2026 study](https://arxiv.org/abs/2601.15195) of 33,596 pull requests opened by coding agents, failing [continuous integration](https://specstory.com/learning/ci-cd/continuous-integration) (CI) or tests was the most common rejection reason at the code level. It covered 17.6% of the rejected pull requests the authors analyzed.

In [code review](https://specstory.com/learning/code-review/code-review) of a person's change, the reviewer can usually assume that the author read and ran the code. An agent's pull request supports neither assumption. A diff shows what changed in the files, not what the software does when someone uses it.

## What do you need before you start?

A reviewer needs four things before opening the diff:

-   **The task.** The issue or prompt states what the change should do, and saved [coding session history](https://specstory.com/learning/ai-coding/ai-coding-session-history) shows the instructions the agent received.
-   **A way to run the branch.** A local checkout or a preview environment lets someone start the software on the head commit.
-   **The CI results.** The reviewer reads which jobs ran and which were skipped, not only the green check mark.
-   **An owner for the change.** The developer who gave the agent its task answers review questions about each changed file.

## How do you review an agent's pull request step by step?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme gives a coding agent the request "Let customers edit their delivery address during checkout." The agent opens pull request #1042 with the summary "Address editing added. Tests pass."

### 1\. Restate the task as checks

Write down what the request requires. For Acme, the address changes, the cart keeps its 2 items, and the total stays $240.00. The request says nothing about sessions, so a session change needs a reason.

### 2\. Check the scope of the diff

Check out the branch, then list the changed files before reading any of them:

```bash
gh pr checkout 1042
git diff --name-status origin/main...HEAD
```

The second command prints one line per file, with `M` for modified and `A` for added:

```text
M	src/checkout/address.ts
M	src/lib/session.ts
A	tests/checkout/address.test.ts
M	tests/checkout/cart.test.ts
```

The file `src/lib/session.ts` is an [out-of-scope edit](https://specstory.com/learning/glossary#out-of-scope-edit) until the developer explains it. Shared code widens the [blast radius](https://specstory.com/learning/glossary#blast-radius), because every page that reads the session depends on it. Edits to `package.json` or the CI configuration need a reason too, because they change what installs and runs. A developer who cannot explain a changed file has not reviewed it yet.

### 3\. Read the test changes first

The pull request edits an existing cart test:

```diff
 test("cart keeps its items after an address change", async () => {
   const order = await startCheckout({ items: 2 });
   await order.setAddress("12 Elm Street");
-  expect(order.items).toHaveLength(2);
+  expect(order.items).toBeDefined();
 });
```

The edited assertion also passes for an empty cart. Weakening an existing assertion is a form of [test tampering](https://specstory.com/learning/verification/test-tampering). The added address test checks only the address, a common pattern in [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass) that pass on broken code.

### 4\. Trace the riskiest path in the code

Read the code where the scope check pointed. In `src/lib/session.ts`, the change regenerates the session after the address is stored. The regenerated session starts empty, and the cart lived in the old one. The diff never shows that the cart is stored in the session, so a reviewer finds this only by opening code outside it. Session code also needs the session management checks in the Open Worldwide Application Security Project (OWASP) [secure code review](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Code_Review_Cheat_Sheet.html) cheat sheet.

### 5\. Ask for evidence of a real run

A passing suite is evidence about the tests. A run is evidence about the software. Ask the developer for a record of a run on the head commit, with the steps and the result of each check from step 1.

Pull request #1042 has no such record. The reviewer runs `npm run dev` on the branch, adds 2 items, and changes the address. The address saved, but the cart emptied. The green CI check was a [false pass](https://specstory.com/learning/verification/false-pass).

### 6\. Send the failure back as reproduction steps

Post the failure as steps the agent can replay:

```text
Commit: 3f9c2e1
Steps: add 2 items to the cart, start checkout, change the
  delivery address to "12 Elm Street", continue.
Expected: the cart still holds 2 items and the total is $240.00.
Actual: The address saved, but the cart emptied.
Also: restore the item count assertion in cart.test.ts.
```

The agent then changes the session code and restores the assertion. This example is simplified. A real pull request often needs more than one round of fixes.

## How do you handle large agent pull requests?

GitHub's [Octoverse 2025](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) counted an average of 43.2 million pull requests merged each month, up 23% on the year before.

[Salesforce engineering](https://engineering.salesforce.com/scaling-code-reviews-adapting-to-a-surge-in-ai-generated-code/) reported that as AI raised code volume by about 30%, review time on its largest pull requests began to level off or even fall. The team read this as reviewers disengaging, not getting faster. A diff that long can end in a [rubber-stamp review](https://specstory.com/learning/glossary#rubber-stamp-review).

Three practices keep a large agent pull request reviewable:

-   **Ask for a split before review.** One pull request per concern lets a reviewer read and run each part. A team can set a [pull request size](https://specstory.com/learning/code-review/pull-request-size) above which it asks for a split instead of a review.
-   **Ask for a file map.** The description should say which part of the task each changed file serves. A file with no entry is a scope question.
-   **Triage overnight pull requests by evidence.** A background coding agent can open several pull requests while nobody watches. Send back the ones with failing CI unread, and ask the rest for a run record.

On GitHub, a branch protection rule can require an approving review from someone with write access before a merge, so overnight pull requests usually wait for a person. By default, administrators can bypass the rule.

## What are common mistakes?

These mistakes let a broken agent pull request through review:

-   **Trusting the summary.** The description is text the agent generated about its own work. A claim of "done" is not a record of a run.
-   **Reading only the diff.** The diff shows the changed lines, not the code that depends on them. Acme's cart lived in the session, and nothing in the diff showed that.
-   **Treating a green check as a run.** A check can pass on weakened tests. On GitHub, a required status check that reports `skipped` still meets the requirement.
-   **Asking the same agent to review its work.** The agent reviews with the context that produced the change, so it tends to repeat its misses. [AI code review](https://specstory.com/learning/code-review/ai-code-review) from a separate tool helps, but a second reading is still not a run.
-   **Asking for passing tests.** A request to make the tests pass lets the agent edit the tests. Ask for a fix to the code, and keep the assertion that failed.

## How do you check that it worked?

A review worked when the fixed pull request passes three checks on its head commit:

-   **Every changed file has a reason.** The developer has explained `src/lib/session.ts`, or the change is gone.
-   **The tests only gain checks.** The restored assertion fails on the commit before the fix and passes on the fix.
-   **A run record matches the head commit.** The replay shows 2 items in the cart after the address change. A record from an earlier commit does not count.

These checks show that this failure is gone, not that the rest of the change works. To [verify a bug fix](https://specstory.com/learning/debugging/bug-fix-verification), a team also reruns the nearby workflows, e.g. login.

## How does RunStory help with reviewing agent pull requests?

A reviewer can ask for evidence of a run, but someone has to produce it. RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. Your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps.

A reviewer can then ask for that evidence in the pull request, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Does AI-generated code change how teams review?

AI-generated code changes where review effort goes first. The author may not have read or run the change. Reviewers start with the task, the scope, and the test edits, and ask for evidence that the software ran.

### Should there be a line limit per pull request?

A line limit per pull request works best as a trigger for a split before review, not as a rule that rejects work. Behind it, each pull request should hold one concern that a reviewer can read and run.

### Can code merge without a human reviewer?

Code can merge without a human reviewer when the repository allows it. On GitHub, a branch protection rule can require an approving review before a merge, so passing checks alone cannot merge an agent's pull request. By default, administrators can still bypass the rule.

### What if the author does not understand the AI code they submitted?

An author who does not understand the AI code they submitted cannot answer review questions, so the review should pause. The reviewer can ask the author to explain each changed file and show a run first.

---

Source: [How to review AI-generated pull requests | SpecStory](https://specstory.com/learning/code-review/reviewing-agent-pull-requests)
