What is AI code review?
AI code review is a review practice that uses a large language model (LLM) to read a code change and comment on likely problems. It is one kind of automated code review, which also covers linters. The comments usually appear on the pull request, next to the changed lines, and a person reads them before deciding whether to merge.
AI code review is one form of code review. An AI reviewer that only reads the change performs static analysis. The International Software Testing Qualifications Board (ISTQB) defines static analysis as the automated evaluation of a component or system without executing it.
Teams use AI review to get comments before a person is free to read the change. In a 2024 industrial study, developers resolved 73.8% of an AI reviewer's comments, but pull requests took longer to close (from about 6 to about 8 hours on average). Acting on comments takes time, so an AI reviewer can add a step as well as a check.
How does AI code review work?
A typical AI review runs when a pull request opens or changes, in five steps:
- The review tool collects the diff (the lines that changed) and the pull request description.
- The tool adds context from the repository, e.g. files that call a changed function.
- The tool sends the change, the context, and the team's rules to a language model.
- The model returns findings, and the tool drops or ranks them by severity.
- The tool posts the remaining findings as comments on the changed lines.
The author, a person or a coding agent, edits the change, and the review can run again.
Reviewers differ in what they read and whether they run anything:
- Changed lines only. The model reads the diff alone, so it can miss a break in code outside the diff.
- Repository context. The tool searches the repository, so the model can also read callers and tests.
- Agentic review. A coding agent explores the repository and can run commands, e.g. the project's tests. That adds dynamic analysis, but only for the behavior those tests check.
The reviewer can be a hosted tool or a coding agent asked to review. A review in the session that wrote the change is AI self-verification, which starts from the same context, so it can miss the same things. A separate session starts without that context, and a different model can miss different things.
What is an example of AI code review?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent opens a pull request that adds this function:
async function saveAddress(order, address) {
await order.setAddress(address);
window.location.reload();
}
An AI reviewer reads the diff and posts three comments:
address.js:2 Possible bug: a blank address is saved without a check.
address.js:2 Possible race: two saves can overlap.
address.js:1 Style: rename order to currentOrder.
The first comment is correct, and the agent adds the check. The second is a false positive, because the page disables the save button while a save runs, so the developer dismisses it. The agent applies the rename, the next review pass posts no further comments, and a human reviewer at Acme approves the pull request.
After the merge, a customer with 2 items in the cart edits the address. The address saved, but the cart emptied. In this version of the example, the page keeps the cart in memory until payment, and the reload cleared it. The reload is in the diff, but the cart code is not, so no comment connected the two. One end-to-end test that edits the address with items in the cart would have shown the failure before the merge.
This example is simplified. A real pull request would draw more comments, and the team would need rules for which comments to act on.
What can AI code review catch, and what does it miss?
AI can review code, and it does best on problems visible in the text of the change. It often catches these:
- Local bugs. The reviewer reads the changed lines, so it can flag a bug inside them, e.g. a missing null check.
- Known risky patterns. AI reviewers often flag a database query built from strings of user input.
- Drift from the description. The reviewer can flag changes that the pull request text does not mention.
It often misses these:
- Failures at runtime. Some bugs appear only when the software runs with real data and configuration, e.g. a hydration error in a web app.
- Code outside its context. A problem in code the reviewer did not read stays hidden, as the cart code did in the example.
- Design fit. Whether the change fits the system's design depends on plans and history that the diff does not record.
- Missing work. The model comments only on lines it receives, so a requirement the change left out draws no comment.
One review's comments are a sample of possible findings, not a full list. The same review run twice can return different comments, and each false positive costs time to dismiss.
Results also depend on the setup around the model. A small September 2026 study compared a review model in a harness that runs several review passes with the same model given a single prompt. The harness found about 1.6 times as many verified bugs, at about 10 times the tokens.
Checking a change by running the software is often called runtime verification. Reading a change and running code find different bugs, so one does not replace the other.
Should an AI be allowed to approve a pull request?
An AI reviewer can comment on a pull request, but approval is safer with a person who can answer for the change. On GitHub, a branch protection rule can require a set number of approvals from reviewers with write access before a merge. GitHub also has a public preview setting, off by default, that lets its own AI reviewer's approval count toward that number.
A model's approval records that one reading found no blocking problem. It does not show that the software works or name a person who can explain the change. Text in the pull request can also steer the model, a form of prompt injection, e.g. a code comment that asks the reviewer to approve. A coding agent can say "done" before anyone has run the change, and an AI approval adds a second claim of the same kind.
A human approval can be empty too. A rubber-stamp review approves a change without checking it closely, and the risk grows when agents open more pull requests than people can read. A team that allows AI approval can limit it to file paths with little risk, e.g. documentation, and require a run of the software before merge.
How is AI code review different from human review?
An AI reviewer can read each changed line on each push, but its comments vary between runs, and it works only from the context it is given. A human reviewer brings knowledge of the system and the team's plans, and the review passes that knowledge between people. A common setup runs the AI pass first for local bugs and leaves design fit and the merge decision to a person.
How does RunStory help with AI code review?
In the checkout example above, the reviewer read the reload but never ran the checkout, so nothing connected it to the cart. This is an illustrative example, not a live run.
RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. Your coding agent receives the actions RunStory took and evidence of the unexpected result. RunStory is in private alpha for CLIs and web apps. Your team keeps the final release decision.
FAQs
Is AI code review good enough?
AI code review is good enough as an extra pass that catches local bugs before a person reviews the change. It is not enough as the only check, because a review that reads the change can miss failures that appear only when the software runs.
Can AI code review tools replace human reviewers?
AI code review tools cannot replace human reviewers, but they can take over part of the reading of each changed line. A person still judges design fit, passes knowledge of the system between people, and answers for the decision to merge.
Should a team use a dedicated review tool or ask the coding agent?
A dedicated review tool and a coding agent asked to review both send the change to a language model, so either can work. The setup around the model shapes the results. Whichever a team picks, a review in a separate session starts without the author's context.
Does AI code review send code outside the team?
AI code review that runs as a hosted tool sends the diff and the context it gathers to a model provider, outside the team. Check the tool's data retention and training terms before turning it on for a private repository.