# What is code coverage?

Code coverage is the share of a program's code that runs during its tests, which shows what was executed but not whether the tests checked the results.

Last updated September 29, 2026, 11 min read

## Learning objectives

After reading this article you will be able to:

-   Define code coverage and its common measures
-   Explain why covered code is not checked code
-   Compare coverage with test coverage and mutation score

## Related content

-   [What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)
-   [What is the difference between mutation score and code coverage?](https://specstory.com/learning/test-quality/mutation-score-vs-code-coverage)
-   [What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)
-   [What is property-based testing?](https://specstory.com/learning/test-quality/property-based-testing)

## Key points

-   Code coverage reports which parts of a program's code ran while its tests ran.
-   A line counts as covered when it runs, even if no test checks its result.
-   Use coverage to find code that no test runs, and other checks to judge test strength.

## What is code coverage?

Code coverage is a test measure that reports the percentage of a program's code that ran while its tests ran. A coverage tool counts code items, e.g. lines, and divides the number that ran by the number that exist. A high number shows that code was executed. It does not show that any test checked the result.

The International Software Testing Qualifications Board (ISTQB) [defines coverage](https://glossary.istqb.org/en_US/term/coverage) as the percentage of chosen coverage items that a [test suite](https://specstory.com/learning/glossary#test-suite) exercises. Code coverage applies that idea to the structure of the code itself. Tools usually report it for each file and as one total for the project. Teams most often collect it from [unit testing](https://specstory.com/learning/testing/unit-testing), where tests call functions directly.

No single percentage suits every codebase. Many teams use 80% coverage as a target by [rule of thumb](https://www.atlassian.com/continuous-delivery/software-testing/code-coverage). Where a bug costs more, e.g. in checkout code, teams usually expect more coverage and stronger checks. The list of uncovered lines is often more useful than the total, because it names code that no test has run.

## How is code coverage measured?

A coverage tool, e.g. coverage.py for Python, measures coverage during a normal test run in five steps:

1.  The tool instruments the program. It adds counters to the code, or it asks the language runtime to report each line that runs.
2.  The test runner runs the test suite against the instrumented program.
3.  The tool records each line and each branch outcome that ran.
4.  The tool reads the source code and lists the lines and branch outcomes that could have run.
5.  The report divides what ran by what could have run and lists the lines that never ran.

Diagram: How a coverage tool measures a test run

Coverage and the pass or fail result come from the same run. The coverage report counts execution only and never reads the assertions.

The coverage count does not depend on what the [assertions](https://specstory.com/learning/glossary#assertion) check. A test that calls a function and checks nothing adds as much coverage as a test that checks the result. The pass or fail result and the coverage report come from the same run, but they measure different things.

Open-source coverage tools exist for most widely used languages:

-   coverage.py for Python
-   Istanbul for JavaScript
-   JaCoCo for Java
-   `go test -cover` for Go

Teams often turn the total into a [quality gate](https://specstory.com/learning/ci-cd/quality-gate). With coverage.py, `coverage report --fail-under=80` exits with code 2 when the total is below 80, so the build step fails. Some teams apply the threshold only to the lines a change touches, so gaps in older code do not stop the change.

## What are the types of code coverage?

Tools count different items, and each type answers a narrower or wider question about what ran.

### Function coverage

Function coverage counts the functions that the tests called at least once. It is the coarsest of these measures, because one call covers a function however few of its lines ran.

### Statement and line coverage

Statement coverage counts the executable statements that ran, and line coverage counts the source lines that ran. Most reports lead with one of the two. They differ mainly when a line holds more than one statement or a statement spans several lines.

### Branch coverage

Branch coverage is a code coverage measure that counts which outcomes of each decision the tests have run, e.g. the false side of an `if`. An `if` with no `else` still has two outcomes, so every line can run while one outcome never does. Full branch coverage implies full statement coverage, but not the reverse. Some tools measure branches only when asked, e.g. coverage.py with `--branch`.

### Condition coverage

Condition coverage checks that each boolean part of a decision has been both true and false. In `if paid and in_stock:`, both `paid` and `in_stock` need a true run and a false run. Modified condition/decision coverage (MC/DC) is stricter and shows that each condition can change the outcome on its own. It is used where a failure is dangerous, e.g. in flight control software.

### Path coverage

Path coverage counts the distinct routes from a function's entry to its exit that the tests have run. The number of paths grows fast, because each independent `if` doubles it and a loop can make it unbounded. Few teams measure it in full, and most coverage reports leave it out.

## What is an example of code coverage?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The agent writes a function and one test:

```python
# checkout.py
from models import Order

def change_address(order, new_address):
    if not new_address.strip():
        raise ValueError("delivery address is empty")
    return Order(id=order.id, address=new_address)

# test_checkout.py
from checkout import Order, change_address

def test_change_address():
    order = Order(id="A-1042", items=["sofa", "lamp"])
    updated = change_address(order, "12 Elm Street")
    assert updated.address == "12 Elm Street"
```

The test passes. The developer measures coverage for `checkout.py` with branches turned on:

```text
$ coverage run --branch --include=checkout.py -m pytest
$ coverage report
Name          Stmts   Miss Branch BrPart  Cover
-----------------------------------------------
checkout.py       5      1      2      1    71%
-----------------------------------------------
TOTAL             5      1      2      1    71%
```

Four of the 5 statements ran. The `raise` line did not, and the `if` took only its false side, so 1 of its 2 branch outcomes is missing. The Cover column combines statements and branches, so it counts 5 of the 7 items.

The developer asks the agent for full coverage. The agent adds a second test that passes a blank address and expects a `ValueError`. The report now shows full coverage, with 0 missed statements and 0 partial branches, and both tests pass.

Then a customer reports the failure. The address saved, but the cart emptied. `change_address` returns a different `Order` with no items, and neither test checks the items.

The first test needs one added line, `assert updated.items == ["sofa", "lamp"]`. The test now fails with `AssertionError: assert [] == ['sofa', 'lamp']`. The coverage report does not change, because the added line runs no extra code in `checkout.py`. This example is simplified. A real project would measure coverage across many files and would also need a check on the order total.

## What changes when a coding agent writes the code?

A coding agent can raise a coverage number faster than a reviewer can read the tests behind it. When a prompt names a coverage target, the agent's output can reach it with tests that call functions and assert little. That is one reason [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass) pass on broken code.

The coverage report also shows the code that an agent's tests never ran. [A 2025 benchmark](https://arxiv.org/abs/2508.00408) of 3,909 Python functions from real code was built to avoid leakage from training data. On that benchmark, unit tests written by large language models (LLMs) averaged 45% statement coverage and killed 40% of injected mutants. A mutant is a copy of the code with one small deliberate bug, and a test kills it by failing. Much of the code those tests were written for never ran.

A practical adjustment is to ask the agent for a failing test first, as in [test-driven development](https://specstory.com/learning/testing/test-driven-development). A test that failed before the change and passes after it has shown that it checks something. A reviewer can then read the uncovered lines in the files the agent changed, which is where testing [AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing) can start.

## Why can 100% coverage hide broken code?

Full coverage means each counted item ran at least once. [Researchers](https://arxiv.org/abs/2506.02954) have found test suites with 100% code coverage that caught only 4% of deliberately injected bugs, which is why coverage alone is a weak signal. The author of [a 2011 post](https://testing.googleblog.com/2011/02/this-code-is-crap.html) on the Google Testing Blog writes that "you can have great code coverage and lousy tests."

Coverage has four limits:

-   **Coverage counts execution, not checking.** The Acme tests covered the line that emptied the cart, and no assertion checked the items.
-   **Coverage cannot measure missing code.** When the code never handles a case, there is no line for the report to mark as missed. The report measures the code that exists, not the behavior the request asked for.
-   **Coverage does not record the inputs.** A line that ran once with one value counts as covered, but it can still fail with another value, e.g. an empty cart.
-   **Full coverage costs the most at the end.** The last uncovered lines are usually error handlers and rare branches. Tests written only to reach them often check little.

A test that passes on broken code is a [false pass](https://specstory.com/learning/verification/false-pass), and a full coverage report gives no sign of one. No findings is not the same as complete coverage, and complete code coverage is not the same as checked behavior.

## How is code coverage different from test coverage?

Code coverage measures the code, and a tool reports which parts of it ran during the tests. Test coverage often means a measure of the requirements or risks, and counts how many have at least one planned test. It often comes from a traceability matrix, which maps each requirement to its tests, not from a coverage tool. The terms overlap, because the ISTQB glossary lists "test coverage" as a synonym for coverage in general, and some tools use it to mean code coverage.

Both differ from mutation score, which counts checking instead of execution or planning:

| Measure | What it counts | What a high number shows |
| --- | --- | --- |
| Code coverage | Code items that ran during the tests | The code was executed |
| Test coverage | Requirements or risks with at least one planned test | Each listed item has a test |
| Mutation score | Injected bugs that made at least one test fail | The tests check results |

## What does test quality cover?

Each page in this topic answers one part of how to tell whether tests catch real bugs:

-   [Test oracles](https://specstory.com/learning/test-quality/test-oracle) supply the expected result that a test compares with the actual one.
-   [Mutation testing](https://specstory.com/learning/test-quality/mutation-testing) measures whether a test suite detects small deliberate bugs injected into the code.
-   [Mutation score](https://specstory.com/learning/test-quality/mutation-score-vs-code-coverage) and code coverage answer different questions, since one counts caught bugs and the other counts executed code.
-   [Property-based testing](https://specstory.com/learning/test-quality/property-based-testing) checks rules that must hold for every valid input against many generated inputs.
-   [Golden file testing](https://specstory.com/learning/test-quality/golden-file-testing) compares a program's output with an approved copy and fails when the output changes.
-   [Adversarial testing](https://specstory.com/learning/test-quality/adversarial-testing) tries to break software with unexpected or hostile input.
-   [Tautological tests](https://specstory.com/learning/test-quality/tautological-test) cannot fail, because their expected value comes from the same source as the result.
-   [Catching tests](https://specstory.com/learning/test-quality/catching-tests) are generated for one code change and fail if that change added a bug.
-   [Fuzzing](https://specstory.com/learning/test-quality/fuzzing) runs a program on many random or malformed inputs to find crashes and hangs.
-   [Metamorphic testing](https://specstory.com/learning/test-quality/metamorphic-testing) checks that the outputs of related inputs relate by a known rule.
-   [Checking agent-written tests](https://specstory.com/learning/test-quality/checking-agent-written-tests) means breaking the code on purpose and confirming the tests fail.

## What does coverage leave unchecked in a running app?

A coverage report lists the code that ran during the tests. It cannot show whether a customer can finish a checkout in the built app. To find out, run the built app through the workflow that a change touched, e.g. checkout, and check the cart at the end.

RunStory runs the software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures to the coding agent and verifies the fix. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### What is a reasonable code coverage percentage?

A reasonable code coverage percentage is a target that a team sets by rule of thumb, not a standard that fits every codebase. The list of uncovered lines often tells a team more than the total, because it names the code that no test has run.

### Is 100% code coverage worth chasing?

Complete code coverage is rarely worth chasing as a goal in itself. The last uncovered lines cost the most to reach, and tests written only to reach them often check little. A fully covered function can still return a wrong result.

### How do teams choose the right amount of coverage?

Teams choose the right amount of coverage by risk. Code where a bug costs more, e.g. checkout code, usually gets a higher target and stronger checks. Some teams also apply the target only to the lines a change touches.

### Does path coverage find all bugs?

Path coverage does not find all bugs. A test can run every path and still check the wrong result, and no path exists for behavior the code never implements. Full path coverage is also rarely reachable, because each independent decision doubles the number of paths and a loop can make it unbounded.

---

Source: [What is code coverage? | Code vs. test coverage | SpecStory](https://specstory.com/learning/test-quality/code-coverage)
