Why is CI becoming the bottleneck for AI-written code?
Continuous integration (CI) is becoming the bottleneck because coding agents push changes faster than a team's CI runners can build and test them. Each push usually starts a pipeline run, so the extra runs wait in a queue and results arrive late. Teams under that pressure often cut checks to keep changes moving.
Anthropic reported that its CI job volume grew 25-fold in six months as Claude came to write about 80% of its code. The growth forced a redesign of how it chooses which tests to run.
In a CI/CD pipeline, each job in a run holds a runner, the machine that runs it, for as long as its checks take. Scaling CI/CD for AI coding agents means getting results back sooner without dropping the checks that catch real failures.
What causes the CI bottleneck?
CI load is the total runner time that a team's runs need. It equals the number of runs multiplied by the average runner time of one run, and coding agents raise both.
Why do agents start more CI runs?
One developer can run several agents at once, each in its own Git worktree, and each opens a pull request. A background coding agent works with no person watching, so nobody paces its pushes. Agents also push in small steps, so each fix attempt can start another run.
GitHub's Octoverse 2025 counted an average of 43.2 million pull requests merged each month, up 23% on the year before. More pull requests, from people and agents, mean more pushes, and each push can start a run.
Why do test suites grow with agent output?
Agents often add tests with each change. Each added test runs in each later run, so runs get longer even when the number of changes stays flat. Linear reported that its test suites nearly quadrupled in 2026 as coding agents sped up shipping, and it had to rework CI to keep feedback fast.
Why does slow CI feedback cost more with an agent?
CI feedback time is the time from a push to a result the author can act on, including the time a run waits in the queue. A coding agent's turn usually ends when its output says "done." A later failure has to be brought back to an agent, often by someone pasting the log into another session. The fix then starts another run, which waits in the same queue.
Why do reruns add load?
A flaky test passes and fails on the same code. A job that a flaky test fails is often rerun, and the rerun uses about as much runner time as the first attempt. More runs give a flaky test more chances to fail, so reruns grow with agent output.
What does an overloaded CI look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme runs 4 coding agents at once, and one task is "Let customers edit their delivery address during checkout." Acme's CI workflow runs the whole suite on 2 events:
on:
push:
pull_request:
jobs:
test:
runs-on: [self-hosted, linux]
steps:
- uses: actions/checkout@v7
- run: npm ci
- run: npm test
- run: npm run test:e2e
One morning goes through these steps:
- Each agent opens a pull request and pushes about 5 times. Each push to a branch with an open pull request starts 2 runs, one for each event.
- Acme has 4 runners, and a full run takes 20 minutes. By late morning, 18 runs are waiting in the queue.
- The checkout agent's output says "done," and its turn ends. Its last run starts 70 minutes later.
- The end-to-end test fails. The address saved, but the cart emptied. Nobody reads the log until a reviewer opens the pull request after lunch.
- A flaky test fails 2 other runs, and a developer reruns both, which adds 40 runner minutes.
- To clear the queue, the team moves the end-to-end tests to a nightly run.
When the queue stays long, teams often cut slow or unreliable checks first, and these can catch what unit tests miss:
- End-to-end tests. They move to a nightly run or stop being required, so a change that breaks checkout can merge.
- The version matrix. Runs on several runtime versions shrink to a run on one.
- Flaky tests. They get skipped or retried until they pass, which can hide a real failure.
- The full suite. Test selection replaces it, with no full run to catch what selection missed.
Each cut check adds to verification debt, the gap between the code a team produces and the code anyone has checked by running or testing it. This example is simplified. A real queue would also hold runs from people and scheduled jobs.
How can teams scale CI for coding agents?
These approaches cut runner time or waiting time, and some trade away coverage:
- Cancel runs that a later push replaced. In GitHub Actions, a
concurrencygroup per workflow and branch, withcancel-in-progress: true, cancels the run in progress when a newer run for that branch is queued. Limitingpushruns to the main branch also removes the example's duplicate runs. - Run fast checks in the agent's loop. Shift-left testing moves checks earlier, e.g. the unit tests for the changed package before the agent pushes. CI still reruns them, because an agent can skip a local check.
- Run the affected tests. Test impact analysis maps which tests exercise which code, so a change runs only the tests it could affect. It can miss a test whose link to the code is not recorded.
- Put the fastest signal first. A smoke test runs first and checks that the build starts and its core functions work, so a broken build fails early. A suite that follows the test pyramid keeps most tests fast, and its few end-to-end tests try fewer of the software's workflows.
- Move slow suites to a merge queue. A merge queue tests each approved pull request with the main branch and the changes ahead of it. It adds runs of its own, so it saves runner time only when slow suites run there instead of on each push. A failure then arrives after approval, often after the agent's turn has ended.
- Fix flaky tests before adding runners. Each fixed flaky test removes the reruns it caused.
- Add capacity. More runners cut the wait, and test sharding splits one run across several machines. On GitHub Actions, private repositories pay by the minute for standard GitHub-hosted runners beyond an included quota. Runners that a team hosts itself cost the machines they run on.
A common pattern runs fast checks and the affected tests on each push. The full suite runs before a merge or on a schedule. A faster pipeline can also move the wait to the next queue, the review bottleneck, where people read the pull requests that pass.
How is CI load different from slow tests?
Slow tests make each run long, while CI load grows with the number of runs even when each test is fast. With slow tests, a run takes long after it starts. When load exceeds the runners' capacity, a run waits long before it starts.
The two compound, because a slow suite makes each extra run cost more runner time. Faster tests shorten each run and also cut the load. Fewer runs and more runners cut the wait, but not the time each run takes.
How does RunStory help with CI for coding agents?
When a busy pipeline cuts its end-to-end tests, fewer of the software's workflows get tried. Coding agents outpace testing, even with AI review. RunStory runs your software against the change in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures back to your coding agent while you keep working. It is in private alpha for CLIs and web apps.
FAQs
What is CI feedback time?
CI feedback time is how long an author waits between a push and a result they can act on. Queue time counts as well as run time, so feedback slows when agents fill the queue, even if no test gets slower.
Should every agent commit run the full suite?
Running the full suite on every agent commit gives the most coverage, but each push then costs a full run, and agents push often. A common split runs fast checks and the affected tests on each push, and keeps the full suite for merges or a schedule.
Should agents wait for CI before starting the next task?
An agent can start the next task while CI runs, as long as a failure still reaches a person or another agent session. Running fast checks in the agent's loop before the push leaves fewer failures for CI to return late.
Can CI costs grow with coding agents?
CI costs can grow with coding agents, because agents start more runs and their added tests make each run longer. On GitHub Actions, standard GitHub-hosted minutes for private repositories are billed beyond a plan's quota, and runners a team hosts itself cost the machines they need.