Skip to content

Why are AI code reviews noisy?

AI code reviews get noisy when a reviewer posts many false positives and minor comments, so developers learn to skim them and miss the few real problems.

Last updated , 8 min read

Why are AI code reviews noisy?

AI code reviews are often noisy because the reviewer flags code that could be wrong without checking whether it is wrong. Some flags are false positives, and they often arrive in the same form as real bug reports and style comments. Developers learn to skim them, and a real finding can get dismissed with the rest.

AI code review is a form of code review in which a language model reads a pull request and comments on the changed lines. A false positive is a check result that reports a problem that does not exist. Review noise is any comment that costs reading time without pointing to a change worth making.

In a 2024 industrial study, developers resolved 73.8% of an AI reviewer's comments, but pull requests took longer to close (from about 6 to about 8 hours on average). The authors list "faulty reviews, unnecessary corrections, and irrelevant comments" among the tool's drawbacks. Resolving a comment shows that someone acted on it, not that it named a real problem.

A coding agent told to address the review can edit the code for each comment, including comments about problems that do not exist.

What causes noise in AI code review?

When many comments are wrong or minor, people skim and dismiss them. This habit is alert fatigue, the loss of attention that follows when warnings arrive so often, many of them false alarms, that people start to ignore them. Approval fatigue builds the same way from a coding agent's permission prompts, which people end up approving without reading.

How review noise hides a real finding Reviewer posts many comments Many are wrong or minor People skim and dismiss Real finding dismissed too Next review: read with less attention
The real finding reaches the reader in the same form as the noise, so it gets the same quick dismissal.

Four causes feed this loop.

Why does a reviewer flag code it has not checked?

A language model generates review comments from patterns in the text of the change, e.g. a value used without a null check. The pattern shows that a bug is possible. Unless the reviewer runs the code, it cannot confirm that the bug occurs, so the reader has to settle each suspicion. Tools that read runtime context still have not run the change.

Why does missing context produce wrong comments?

A reviewer that receives only the changed lines cannot check the code that makes them safe, e.g. a check in the calling function. It flags a problem that another file already handles. A tool that searches the repository avoids some of these comments, but no reviewer can read a decision that was never written down.

Why do minor and serious comments look the same?

Many reviewers post each finding as a line comment in the same format. A naming nit and a bug that loses data look alike, so the reader has to open each one to find the bug. Style comments often repeat what linting already checks.

Why do comments change from run to run?

A language model usually samples its output with some randomness, so two runs on the same diff can return different comments. Two tools can also differ in model, context, and instructions. A check whose output changes on the same code has the effect of a flaky test, which people learn to ignore.

What does a noisy review look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent opens a pull request that saves the address and then calls session.reload(). An AI reviewer posts 9 comments on the changed lines:

address.ts:12  Nit: rename addr to address.
address.ts:14  Possible bug: address may be undefined here.
address.ts:14  Consider wrapping setAddress in try/catch.
address.ts:18  Nit: prefer const over let.
address.ts:21  Possible race: two saves can overlap.
address.ts:22  Reloading the session may drop unsaved cart items.
address.ts:25  Consider extracting this block into a helper.
address.ts:30  Nit: missing doc comment.
address.ts:31  Nit: trailing whitespace.

The review goes this way:

  1. The developer skims the list. Four comments are nits, two are optional, and two flag problems that cannot happen, e.g. an undefined address, which the calling function rejects.
  2. The developer asks the agent to address the review. It renames the variable, adds a check for an undefined address, and wraps the save in a try block.
  3. On line 22, the agent replies that the reload refreshes the session. The developer resolves it with the others.
  4. A human reviewer approves the pull request, which is now 14 lines longer.
  5. After the merge, a customer with 2 items in the cart edits the address. The address saved, but the cart emptied.

The reviewer named the bug on line 22, but the comment sat among 8 others of equal weight. The agent's check for an undefined address guards against a case that cannot occur.

This example is simplified. A real review would span more files, and a checkout run with a full cart would show which comment was real.

How can teams reduce review noise?

Teams usually combine several of these changes:

  • Move rule checks to a linter. Style checks fit a linter, whose rules each have a name and a setting that turns them off. The AI reviewer's instructions can then exclude them.
  • Write the review instructions down. Many review tools read an instructions file in the repository. The file can say what to skip and why, e.g. generated files, and changes to it get reviewed with the code.
  • Rank findings by severity. Severity labels let people read the serious findings first. Some tools can also group minor findings in one summary comment.
  • Filter drafts before they post. A second model call can grade each draft finding and drop the weak ones. That is a form of LLM-as-a-judge, and the judge has biases of its own.
  • Measure the share of real findings. The reviewer's precision on a sample of the team's own pull requests is the share of its comments that name real problems.
  • Check findings by running the code. A reviewer that runs the tests or the app before it posts can drop suspicions that the run does not confirm, and attach evidence to the rest.

A linter and an instructions file cut minor comments, and severity labels let people read them last. A judge model and a run of the code cut wrong comments, while a stricter prompt cuts the count and checks nothing. A quieter reviewer can also drop real findings, and its precision can rise while it drops them. No findings is not the same as complete coverage.

A person who reviews a pull request from a coding agent can read the AI comments last, after the task and the evidence of a run.

How is noise different from a false positive?

Every false positive is noise, but not all noise is a false positive. The International Software Testing Qualifications Board (ISTQB) defines a false positive as a test result that reports a defect that does not exist. Noise also includes correct comments that nobody needed, e.g. a nit on a line that works.

A reviewer that posts many correct but minor comments can score high precision and still be noisy, because precision counts wrong comments, not unneeded ones. Review tools often report both under the name false positive, so check which one a tool's figure counts.

How do you report fewer, stronger findings?

Post a finding only after a check shows the problem, and attach what the check did. A comment that a reload "may drop unsaved cart items" leaves the checking to the reader. A finding with reproduction steps and the observed result lets a person or a coding agent repeat the failure and confirm it.

RunStory's private alpha tests your software in isolated sandboxes. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. The alpha starts with CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Why do two AI reviewers disagree about one pull request?

Two AI reviewers often disagree about one pull request because they can use different models, receive different context, and follow different instructions. Each also usually samples its output with some randomness, so one reviewer run twice can return different comments.

What is alert fatigue in code review?

Alert fatigue in code review is the habit of skimming or dismissing review comments because many of them have been wrong or minor. A real finding then gets the same quick answer as the noise around it, and a bug can merge next to a comment that named it.

Does a stricter prompt make an AI reviewer quieter?

A stricter prompt usually makes an AI reviewer quieter, with fewer comments on each pull request. The prompt adds no check, so the remaining comments can still be wrong. A quieter reviewer can also drop real findings, so a team can compare the number of real findings before and after, not only their share.

How do you turn off a review rule that keeps misfiring?

To turn off a review rule that keeps misfiring, first check whether a linter can own the check. A linter rule has a name and a setting that turns it off. For an AI reviewer, add a line to the repository's review instructions that names the pattern to skip and why.