What are precision and recall in AI code review?
Precision and recall are two measures of how well a code reviewer finds real problems. Precision is the share of the reviewer's comments that point to a real problem. Recall is the share of the real problems that the reviewer comments on. Raising one usually lowers the other. Flagging every changed line tends to raise recall and lower precision.
In statistics, precision is also called positive predictive value, and recall is also called sensitivity. Any code review can be scored this way, and benchmarks of AI code review tools often report both.
Each measure counts one kind of error. Precision falls with each false positive, a comment about a problem that does not exist. Recall falls with each false negative, a real problem that draws no comment. The International Software Testing Qualifications Board (ISTQB) defines a false positive as a test result that reports a defect that does not exist. Many review tools use the term more loosely, for any unhelpful comment, including a correct one that nobody acted on.
A false positive costs the time to read and dismiss a comment. When a coding agent edits the code to answer each comment, a false positive can become a change that nobody needed. A false negative can let a bug merge.
How do review benchmarks score tools?
A review benchmark scores each tool against a fixed list of known problems, called the ground truth, in five steps:
- The benchmark's authors collect pull requests with known problems, e.g. bugs that a later commit fixed.
- The benchmark runs each review tool on each pull request and saves its comments.
- A person or a model matches each comment to one known problem, or to none.
- The benchmark counts each matched comment as a true positive, each other comment as a false positive, and each missed problem as a false negative.
- The benchmark computes precision and recall for each tool, and often a combined score, e.g. F1.
The standard formulas for the three scores are these:
precision = true positives / (true positives + false positives)
recall = true positives / (true positives + false negatives)
F1 = 2 * precision * recall / (precision + recall)
F1 is the harmonic mean of precision and recall, an average that sits closer to the lower of the two than a plain average does. The F-beta score is a variant that weights recall against precision with a number called beta. A beta above 1 favors recall, and a beta below 1 favors precision.
Step 3 is often done by a large language model (LLM) that compares each comment with each known problem, a setup called LLM-as-a-judge. A wrong match by the judge changes both scores.
Other checks can be scored with the same measures. A mutation score is a recall measure for tests, the share of injected bugs that a test suite detects.
What is an example of precision and recall?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Two AI reviewers comment on the agent's pull request. The team later finds four real problems in the change:
- A blank address saves without a check.
- A failed save shows no error message.
- The address field accepts any length.
- The address saved, but the cart emptied.
Reviewer A posted 2 comments, and both named real problems. Reviewer B posted 10 comments, and 3 of them named real problems. The reviewers score as follows:
| Measure | Reviewer A | Reviewer B |
|---|---|---|
| Comments posted | 2 | 10 |
| Precision | 2 of 2, or 1.0 | 3 of 10, or 0.3 |
| Recall | 2 of 4, or 0.5 | 3 of 4, or 0.75 |
| F1 score | About 0.67 | About 0.43 |
Reviewer B found one more real problem, but someone had to read and dismiss 7 wrong comments to get it. Reviewer A suits a team whose coding agent acts on each comment. Reviewer B suits an extra pass over risky code, e.g. payment logic, where a missed bug costs more than a wrong comment.
Neither reviewer flagged the emptied cart, a bug that shows only when the checkout runs with items in the cart.
This example is simplified. A careful benchmark uses many pull requests and repeats each review, because the same reviewer can post different comments on the same change.
Why do vendor benchmarks disagree?
Published review benchmarks often rank the same tools in different orders. Many are run by a company that sells one of the tools being ranked. The scores also differ because each benchmark makes its own choices at each step:
- Different bug sets. Each benchmark picks its own pull requests and kinds of bug. A tool can score well on one set and poorly on another.
- Different matching rules. A loose judge counts a vague comment near a bug as a match, and a strict judge requires the comment to name the bug. Looser matching raises both scores.
- Different ground truth. Some benchmarks list bugs that people confirmed by hand. Others count a comment as correct when developers later changed the code it pointed to. That measures which comments people acted on, so a correct comment that nobody acted on counts as a false positive.
- Unscored false positives. A benchmark that counts only the known bugs each tool caught reports recall and leaves out precision.
- Incomplete bug lists. A real bug that is missing from the list counts as a false positive when a tool flags it. A benchmark with a short list can rank the most thorough tool lower.
- One run per tool. Review output varies between runs and between settings, e.g. an effort level. A single run is one sample of what each tool can post.
Two scores are comparable only when both benchmarks use the same pull requests and the same scoring rules.
What does a benchmark score leave out?
A score describes one reviewer on one list of known problems, and it leaves out four things:
- Severity. A missed injection flaw and a missed typo each count as one false negative, so the score weighs them the same.
- Whether the change runs. A score shows what a reviewer found by reading, not whether the change works. Reading code and running code find different bugs. A tool that reads runtime context, e.g. production logs, still has not run the change.
- Cost. The score does not count what each review costs to run, or the time people spend dismissing false positives.
- The team's own code. A benchmark's pull requests come from other projects, which may use other languages and rules.
A clean review is not evidence that code works. High precision means that a reviewer's comments are usually right. It says nothing about the problems that drew no comment.
A team can score a reviewer on its own pull requests. A person labels each comment on a sample of recent changes as a real problem or not, and precision is the share labeled real. Recall needs a list of the real problems in those changes, e.g. bugs traced back to them later, and no team can be sure that list is complete.
How are precision and recall different from accuracy?
Accuracy is the share of all decisions that are correct, including true negatives, the clean places that drew no comment. Precision and recall leave true negatives out. In code review, clean lines far outnumber buggy ones. On a change of 200 lines with 2 buggy lines, a reviewer that posts nothing is right about 198 lines. Its accuracy is 198 of 200, or 0.99, and its recall is zero.
A review tool that quotes an accuracy figure should say what it counted as a true negative.
FAQs
What is an F1 score?
An F1 score is a single number that combines precision and recall as their harmonic mean. The F1 score stays low when either measure is low, so a reviewer cannot score well by pushing one measure up while the other stays low. The F-beta score weights one measure more heavily.
Should a code review tool favor precision or recall?
A code review tool should favor precision when people or a coding agent act on each comment, because each wrong comment then costs time or an unneeded change. A tool can favor recall on an extra pass over risky code, where a missed bug costs more than a wrong comment.
How do you measure a reviewer's precision on your own pull requests?
To measure a reviewer's precision on your own pull requests, have a person mark each comment on a sample of recent changes as a real problem or not. Precision is the share marked real. Recall is harder to measure, because no team can be sure its list of real problems is complete.
Is an ignored review comment a false positive?
An ignored review comment is not always a false positive. Under the ISTQB definition, a false positive reports a problem that does not exist, so a correct comment that nobody acted on is a true positive. Some review tools and benchmarks still count it as a false positive.