Skip to content

Why does pull request size matter?

Pull request size affects review quality, because reviewers catch less as diffs grow, so large changes get skimmed and more defects reach the main branch.

Last updated , 9 min read

Why does pull request size matter?

The size of a pull request sets how closely a reviewer can read it. A short diff, the set of changed lines, usually gets a careful reading. A long one often gets skimmed, so defects in it are more likely to reach the main branch. Small pull requests also wait less for review and are simpler to roll back.

Google's engineering practices guide says small changes are reviewed more thoroughly and are less likely to introduce bugs. It puts the right size at one change that does one thing, and calls 100 lines usually reasonable and 1,000 lines usually too large, leaving the call to the reviewer.

In code review, size usually means lines added plus lines deleted. The number of files counts too, because a change spread across many files takes longer to read than the same lines in one file.

Reviewers who skim a long diff often give a rubber-stamp review, an approval that no longer shows that anyone checked the change closely. Later, when git bisect points to the change that introduced a bug, a small pull request leaves fewer lines to search.

What causes large pull requests?

Large pull requests usually grow from one of three causes.

Why does one task grow into several changes?

A task that sounds small can touch several layers of the code. Authors then add nearby but separate work, e.g. a refactoring, which changes how code is structured without changing what it does. Each addition makes sense alone, but together they fill one diff with several concerns.

Why does a slow review queue lead to bigger pull requests?

When each review means a long wait, authors batch their work so that they wait once. A review bottleneck then produces larger pull requests, which take longer to review and lengthen the queue. Meanwhile the main branch moves, so a long-running branch collects merge conflicts, and each fix adds work before the review can end.

Why do mechanical edits inflate a diff?

Mechanical edits change many lines without changing behavior. A rename can touch dozens of files, and a package manager can rewrite hundreds of lines of a lock file for one update. These lines read fast, but they can bury the few lines that change what the software does.

What does an oversized pull request look like?

Here is an illustrative example. Acme Co. sells furniture online, and a developer there asks a coding agent to "Let customers edit their delivery address during checkout."

The agent changes the address form and the session code, and adds a test. It also runs npm update and renames a helper from getSession to loadSession in 27 files, reformatting each one. Git summarizes the pull request this way:

$ git diff --stat=72 --stat-count=6 origin/main...HEAD
 package-lock.json         | 412 ++++++++++++++++++++++----------------
 src/account/orders.tsx    |  25 ++-
 src/account/profile.tsx   |  25 ++-
 src/account/settings.tsx  |  25 ++-
 src/cart/cart-page.tsx    |  26 ++-
 src/cart/cart-summary.tsx |  26 ++-
 ...
 31 files changed, 842 insertions(+), 377 deletions(-)

The request needed 3 of the 31 files. The address form and the session code are files 14 and 17. The review goes this way:

  1. The reviewer skips the lock file and reads the first renamed files closely.
  2. After 10 files of the same rename, the reviewer scrolls faster.
  3. The session change, which clears the session after saving the address, gets a glance.
  4. The reviewer approves the pull request.
  5. After the merge, a customer with 2 items in the cart changes the address. The address saved, but the cart emptied.

Split by concern, the same work becomes 3 pull requests, and the address change is 3 files that a reviewer can read line by line. This example is simplified. A real review would also check the order total after the address change.

What changes when a coding agent writes the code?

When a coding agent writes the code, a large diff costs the author little and the reviewer as much as before. An agent can also run commands while it works, e.g. a package update, and the files they rewrite land in the same diff as the request.

Salesforce engineering reported that as AI raised code volume by about 30%, review time on its largest pull requests began to level off or even fall. It read this as reviewers disengaging, not getting faster.

Edits that the task never asked for are one way agents break working features, and a large diff hides them. A skimmed change that merges without a run adds to verification debt. When a team reviews a pull request from an agent, it can compare the scope of the diff with the task before reading the code. The team can also give the agent one concern per task and name the files it may change.

How can teams keep pull requests small?

These practices keep pull requests small:

  • One concern per pull request. Each pull request makes one kind of change. Google's guide advises keeping refactorings apart from feature changes and bug fixes.
  • Tests in the same pull request. Each pull request carries the tests for its change, and the software builds and passes its tests after each merge.
  • Unfinished work behind a flag. A feature flag, a setting that turns code on or off without a deploy, keeps a feature switched off until its last part merges.
  • Mechanical changes from a tool. A rename that a tool produced gets its own pull request, and the description names the command. A reviewer can rerun it, check that the tests pass with their assertions unchanged, and read a sample of files.
  • Generated files marked as generated. On GitHub, a linguist-generated line in .gitattributes hides a generated file from diffs by default. GitHub already hides some lock files, e.g. package-lock.json. The reviewer checks what produced the file, e.g. the command that ran.

A change of one line can still break checkout, because its blast radius covers every caller of the changed code. Splitting also costs a review and a merge per part, and parts split too finely can hide the design that connects them.

How are stacked pull requests different from one large one?

Stacked pull requests are a chain of small pull requests in which each one builds on the one before it. Each uses the previous branch as its base, so its diff shows only its own change. One large pull request puts the same work in one diff with one review.

One large pull request compared with a stack One large pull request main Packages, rename, and address change 31 files in one diff, one review Stacked pull requests main Part 1 Update packages 1 file Part 2 Rename helper 27 files Part 3 Edit address 3 files Each arrow points to the base branch of a pull request
Both end with the same code on main. In the stack, reviewers read and merge each part on its own, starting with part 1.

Acme's change can go up as a stack of 3 parts. The address change calls the renamed loadSession, so part 3 builds on part 2. The package update depends on neither, so it could also go up alone against main. The first two parts start this way:

git switch -c update-packages main
# commit part 1, then open its pull request against main
git push -u origin update-packages
gh pr create --base main --fill
git switch -c rename-helper
# commit part 2, then open its pull request against part 1
git push -u origin rename-helper
gh pr create --base update-packages --fill

If part 1's branch is deleted after it merges, GitHub changes the base of part 2 to main. When part 1 gets more commits during review, git rebase --update-refs update-packages on the top branch replays the parts above it and moves each of their branches. The author then pushes each moved branch with git push --force-with-lease.

A stack lets review start before the whole feature exists, and each part can be reverted on its own, at the cost of extra rebasing. If the team squashes or rebases each pull request on merge, part 2 needs a rebase onto main, or its diff can show part 1's changes again.

How do you check a change that is too big to read?

A few changes are hard to split, e.g. a framework upgrade that changes a call used in many files. Get the reviewer's agreement before sending one, and write extra tests, as Google's guide advises. Then read the riskiest files closely, e.g. shared session code, and run the workflows the change touches.

Running the software checks what a change does, not how its diff reads. RunStory is in private alpha for CLIs and web apps, and the alpha tests your software in isolated sandboxes. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. Your team keeps the final release decision.

Join the RunStory alpha →

FAQs

How do you split a large pull request?

A large pull request splits best along its concerns. Mechanical edits, e.g. a rename, go into their own pull request, and dependent parts form a stack. Each part should build and pass its tests on its own.

How do you review large refactors?

A large refactor should change structure, not behavior, so the reviewer first checks that the existing tests pass with their assertions unchanged. If a tool made the change, the reviewer can rerun its command and read a sample of the files.

What counts as the size of a pull request?

The size of a pull request usually means its lines added plus lines deleted, read together with the number of files it touches. Lines a reviewer skips, e.g. a lock file, weigh less than lines that change behavior.

Should generated files count toward pull request size?

Generated files, e.g. a lock file, usually count apart from the code a reviewer reads, because the reviewer checks what produced them. On GitHub, a repository can mark such files as generated so that diffs hide them by default.