# What is root cause analysis (RCA)?

Root cause analysis is working back from a failure to the underlying reason it happened, so a fix removes the cause instead of hiding the symptom.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Define root cause analysis
-   Explain common techniques, e.g. the 5 whys
-   Identify the signs of a symptom fix

## Related content

-   [What is debugging?](https://specstory.com/learning/debugging/debugging)
-   [What is git bisect?](https://specstory.com/learning/debugging/git-bisect)
-   [What are reproduction steps in a bug report?](https://specstory.com/learning/debugging/reproduction-steps)
-   [What is a flaky test?](https://specstory.com/learning/debugging/flaky-tests)

## What is root cause analysis (RCA)?

Root cause analysis (RCA) is a method that traces a failure past its symptoms to the underlying cause, so the fix removes that cause. It is also called causal analysis. A symptom is what someone observes, e.g. a crash. The root cause is the condition that, once removed, stops that failure from happening again.

The International Software Testing Qualifications Board (ISTQB) [defines root cause analysis](https://glossary.istqb.org/en_US/term/root-cause-analysis) as a process that identifies the root cause of a defect. Its glossary calls a root cause a source of a defect whose removal makes that type of defect rarer or stops it.

RCA is the part of [debugging](https://specstory.com/learning/debugging/debugging) that finds the cause a fix should remove. The people closest to the failure usually run it, e.g. the developer who reproduced it, working with the code's owners. A fix made without it can remove the symptom and leave the cause, so the failure can return in another form.

## How do you find a root cause?

Finding a root cause usually follows these steps:

1.  The team reproduces the failure from its [reproduction steps](https://specstory.com/learning/debugging/reproduction-steps).
2.  The team lists what changed between the last good state and the failure, e.g. the code merged in between.
3.  The team narrows down where the first wrong value appears, e.g. with log lines.
4.  The team asks why that value is wrong, then asks why about each answer, until an answer names a cause whose removal would stop the failure.
5.  The team tests each candidate cause. Removing it should stop the failure, and restoring it should bring the failure back.
6.  The team fixes the cause and records the contributing factors, which are conditions that helped the failure happen or made it worse without causing it alone.

Three common techniques give this search a structure:

-   **The 5 whys.** The team asks why the failure happened, then why about each answer, often five times. The technique follows one chain of causes.
-   **Fishbone diagram.** Also called an Ishikawa diagram, it sorts possible causes into categories, e.g. the environment, drawn as branches toward the failure.
-   **Change analysis.** The team compares the last working state with the failing one and checks each difference. [Git bisect](https://specstory.com/learning/debugging/git-bisect) narrows the commits in between to one with a binary search.

Failure mode and effects analysis (FMEA) works in the other direction. It lists how a design or process could fail before any failure happens.

## What is an example of root cause analysis?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) to "Let customers edit their delivery address during checkout." After the release, the order summary page crashes. The log shows this [stack trace](https://specstory.com/learning/glossary#stack-trace):

```text
TypeError: Cannot read properties of undefined (reading 'length')
    at renderSummary (/srv/acme/src/checkout/summary.js:18:31)
```

The quickest fix adds a default at line 18, `checkout.items ?? []`. The crash stops, and the page shows an empty order. The address saved, but the cart emptied.

The developer removes the default and asks why three times:

1.  The page crashed because `checkout.items` was undefined.
2.  The field was undefined because the checkout record lost its items when the address changed.
3.  The record lost its items because the address handler saves `{ address: newAddress }` over the whole record.

Answer 3 is the root cause, because the wrong state starts there. The developer then asks why the defect reached customers. The only test for the change never checked the cart, so that test is a contributing factor. The fix copies the record and replaces only the address.

A test changes the address on a cart with 2 items and checks that 2 items remain. It fails on the old code, passes on the fix, and stays as a [regression test](https://specstory.com/learning/testing/regression-testing). This example is simplified. A real analysis would also check the other code that writes the checkout record.

## What changes when a coding agent writes the code?

When a coding agent wrote the change, the analysis can start from the agent's own account of the cause. That account is generated text, not evidence, and it can be as wrong as a [false completion claim](https://specstory.com/learning/verification/coding-agent-done-claims). A person still owns the analysis and tests the account as a candidate cause, by removing that cause and putting it back.

An agent's edit can also [break working features](https://specstory.com/learning/debugging/agents-break-working-features) outside its task, so change analysis over the whole diff can find the cause sooner than reading only the failing code. The agent's [session history](https://specstory.com/learning/ai-coding/ai-coding-session-history) records the request behind that code, so the team can check what a fix at the cause must keep doing.

Stopping at "the agent made a mistake" ends the analysis too early, the same way "human error" does. The practical adjustment is to ask one more why, about the check that let the mistake through. The team then adds that check, e.g. a [code review](https://specstory.com/learning/code-review/code-review) rule that asks each change for a test of what it must keep working.

## What is a blameless postmortem?

A blameless postmortem is an incident review that works out how a failure happened and how to prevent a repeat, without blaming the people involved. [Google's SRE book](https://sre.google/sre-book/postmortem-culture/) describes a postmortem as a written record of an incident, its impact, how it was resolved, its root causes, and the actions to prevent a repeat.

The same chapter says a blameless postmortem must "focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior." Blame hides information, because people who expect it leave out what they did and saw. The chapter moves the question from who to blame to why a person or team had incomplete or incorrect information.

## What are the limits of root cause analysis?

Root cause analysis has four common limits:

-   **Many failures have more than one cause.** A failure often needs several conditions at once. Naming one root cause can drop the others from the fix, so postmortems often list them as contributing factors.
-   **The stopping point is a judgment.** The count of five is a rule of thumb, and different people can follow the same failure to different causes.
-   **A cause without a reproduction is hard to confirm.** For a [flaky test](https://specstory.com/learning/debugging/flaky-tests) or an intermittent failure, a pass after removing a cause can happen by chance.
-   **Removing one cause covers one failure.** It does not show that the software has no other defects.

## How is a root-cause fix different from a symptom fix?

A symptom fix changes the code where a failure shows, e.g. the line that threw an error. A fix to the root cause changes the code or the process where the wrong state starts. Both can make the failing check pass. A symptom fix can still be a reasonable first step to stop harm while the analysis continues.

Diagram: Where a symptom fix and a fix at the cause land

The symptom fix stops the crash, and the cart still empties. The fix at the cause keeps the items, so neither the crash nor the empty cart returns.

A change is often a symptom fix when it shows these signs:

-   The change sits at the line in the stack trace, not where the bad value was created.
-   The change catches or hides an error, e.g. with an empty `catch` block.
-   The change edits the failing test instead of the code.
-   Nobody can say why the wrong value appeared.

## How do you confirm that the root cause is gone?

Remove any symptom fix first, so a hidden failure cannot pass as a fixed one. Then rerun the whole workflow that failed, not only the step that crashed, and check its result, e.g. the items left in the cart. The same steps [verify a bug fix](https://specstory.com/learning/debugging/bug-fix-verification), and keeping them as a test makes the suite fail if this failure returns.

When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### What are the 5 whys?

The 5 whys are a root cause analysis technique that asks why a failure happened and then asks why about each answer. Teams stop when an answer names a cause whose removal would stop the failure, often after five rounds. The technique follows one chain of causes, so it can miss others.

### Is there always one root cause?

A failure does not always have a single root cause. A failure often needs several conditions at once, and each one can be worth fixing. Postmortems often name the main cause and list the others as contributing factors.

### Who should run a root cause analysis?

A root cause analysis is usually run by the people closest to the failure, e.g. the developer who reproduced it, working with the code's owners. When a coding agent wrote the change, the agent's account of the cause is one candidate to test, and a person still owns the conclusion.

### What is a contributing factor?

A contributing factor is a condition that helped a failure happen or made it worse, without causing it alone. In the Acme example, the test that never checked the cart was a contributing factor, because it let the defect reach customers.

---

Source: [What is root cause analysis? | RCA for software | SpecStory](https://specstory.com/learning/debugging/root-cause-analysis)
