Skip to content

What is CI failure triage?

CI failure triage is deciding why a CI run failed, whether from a bug in the change, a flaky test, the environment, or the pipeline, and who should act.

Last updated , 9 min read

What is CI failure triage?

CI failure triage is the work of finding out why a continuous integration (CI) run failed and deciding who should act on it. It is also called build failure triage. The usual causes are a bug in the change, a flaky test, a fault in the CI environment, and a broken pipeline. Each cause has a different owner.

Triage starts from a red check on a pull request or on the main branch of a CI/CD pipeline. The change's author usually triages a failed pull request, and on the main branch a rotating build sheriff often does. Triage ends with a likely cause, an owner, and a next step, e.g. a revert.

At Google in 2016, about 84% of the times a test went from pass to fail involved a flaky test. A failure sent to the wrong owner waits longer, and that wait adds to change lead time, one of the DORA metrics.

How does CI failure triage work?

The person on triage takes each failed job through these steps:

  1. The person finds the first error in the log, because later errors often follow from it.
  2. The person checks where the job failed. A failure during setup points to the environment or the pipeline. A compile error in the changed code points to the change.
  3. The person reruns the job on the same commit. In GitHub Actions, gh run rerun RUN_ID --failed reruns only the failed jobs, with that commit's retry settings. A changed result, or a pass on retry, points to a flaky test or an intermittent bug.
  4. The person checks recent runs on the main branch. The same failure there usually means that an earlier change broke the test.
  5. The person writes a triage note with the commit, the first error, the rerun result, the likely cause, the owner, and the next step.
How triage sorts a failed CI job Failed job read the first error Failed during setup? yes Environment or pipeline fault runner or workflow owner no Result changes on rerun? yes Flaky test or intermittent bug more reruns find the owner no Main branch fails too? yes Already broken on main whoever triages main no Bug in the change author of this change
Each question can send a failure to an owner outside the change. A result that changes on rerun needs more reruns on the change and on main before it has an owner.

Each cause leaves different signs:

  • A bug in the change. The same compile error or wrong value comes back on each rerun, and the main branch passes. The change's author fixes it, or sends it back to the coding agent that wrote the code.
  • A flaky test. The result changes between runs of one commit, and the test has failed on unrelated changes before. An intermittent real bug gives the same pattern, so more reruns come first.
  • An environment fault. The runner or a service that the job calls fails, e.g. the runner's disk fills up. A rerun on another runner often clears it. An exit code of 137 often means that the Linux kernel ended the process to free memory.
  • A broken pipeline. The workflow file or a script it runs is wrong, e.g. a step that needs a secret, which a run from a fork does not receive.

A test that the change added can pass locally but fail in CI because it depends on a machine setting. Its failure sorts as a bug in the change, and the test needs the fix.

What is an example of CI failure triage?

Here is an illustrative example. Acme Co. sells furniture online. A coding agent opens a pull request for "Let customers edit their delivery address during checkout." Three of the run's five jobs fail, and a developer at Acme triages them:

  1. The build job stopped during npm ci with ENOSPC: no space left on device. It passes on a rerun on another runner.
  2. The e2e job's confirmation page test timed out while waiting for the order number. It passes on a rerun, and it failed twice on the main branch in the past week.
  3. The unit job's cart test fails with Expected: 2 and Received: 0 on each rerun and passes on the main branch. The address saved, but the cart emptied.
  4. The developer sends that failure to the coding agent with the first error and the failing command, npx jest src/checkout/address.test.js, which serve as reproduction steps.

The developer posts this triage note on the pull request:

Triage for commit 3f2a91c
build  environment fault  runner team     free disk space on runners
e2e    flaky test         checkout owner  track outside this change
unit   bug in the change  coding agent    fix the cart code

The agent's fix turns the unit job green on the next run. This example is simplified. A real triage would also check whether the change touched the confirmation page before calling that test flaky.

What changes when a coding agent writes the code?

Coding agents add red runs to sort, one reason CI for coding agents becomes a bottleneck. An agent told to make the checks pass can edit code or tests for a red check, whatever the cause. When the cause is a flaky test or a runner fault, that edit changes code that was not broken. The agent's summary can also label its own failure as flaky without a rerun, and that label is a claim for triage to check.

An agent that sorts its own failures often works from the log alone. A 2026 study found that for 58% of the flaky end-to-end tests it studied, the test code and CI logs were not enough to find the cause. Finding those causes would need more evidence from running the tests.

A practical adjustment is to triage before the agent acts. Send the agent only the failures classified as a bug in its change, as a bug report with the first error and a failing command. Route the rest to their owners.

What is self-healing CI, and what can it hide?

Self-healing CI is a pipeline setup in which an automated step, usually a coding agent, reads a failed run, writes a fix, and reruns the checks. Some setups have a person review each fix before it merges.

Because the step fixes what the log shows, it can hide what triage would find:

  • A real bug behind a flaky result. A fix that adds test retries or a longer wait can turn an intermittent bug green.
  • A weakened test. A fix that changes the expected value or skips the test turns the check green while the software is still wrong.
  • A fault outside the code. A longer timeout for a slow runner leaves the runner's fault in place.
  • A failure from another change. A fix for a test that broke on the main branch adds an unrelated edit to the pull request.

A self-healing step is safer when it acts only on failures triaged as a bug in the change. It should stop after a set number of attempts and leave test edits to a person's review.

How is triage different from root cause analysis?

Triage decides what kind of failure a red run is and who acts on it, usually from the log and the run history. Root cause analysis comes after, when the owner that triage chose reproduces the failure and traces it to the condition that caused it. Triage can end at "flaky test, owned by the checkout team," while root cause analysis asks why the test is flaky.

A triage class is a likely cause, not a confirmed one. A wrong class sends the analysis to the wrong owner, e.g. a race condition in the software filed as a flaky test.

How do you turn a red build into a reproducible finding?

A red build becomes a reproducible finding when one command, run on the same code, makes the failure happen again. Keep that command with the first error, so the owner or a coding agent can check a fix.

A CI log often shows where a run stopped but not what the software did before it. RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

Who should triage a failed CI run?

A failed CI run on a pull request is usually triaged first by the author of the change. A failure on the main branch often goes to a person on a rotating duty, sometimes called the build sheriff.

How do you tell an infrastructure failure from a test failure?

An infrastructure failure usually stops a job during setup, and no test reports a wrong value. A test failure names a test and a wrong value or a timeout. A rerun on another runner often clears an infrastructure failure.

Should a coding agent fix CI failures on its own?

A coding agent can fix a CI failure on its own when triage has classified it as a bug in the agent's change. A flaky test, a runner fault, or a broken pipeline belongs to another owner, and an agent told to make it pass can weaken or skip a test instead.

What should a triage note record?

A triage note records the commit, the first error, the rerun result, the likely cause, the owner, and the next step for each failed job. The rerun result shows whether the failure repeats. The owner and next step keep the failure from waiting.