Skip to content

What is the difference between mutation score and code coverage?

Mutation score is the share of small injected bugs that a test suite catches, while code coverage is the share of code that runs during its tests.

Last updated , 8 min read

What is the difference between mutation score and code coverage?

Code coverage measures how much code runs during the tests, while mutation score measures what share of small injected bugs the tests catch. Coverage shows what the tests reached, and the score from mutation testing, also called mutation analysis, shows what they checked. A test suite can run every line and still miss most of the injected bugs.

PIT, a mutation testing tool for Java, reports the same measure as "mutation coverage."

Many teams use 80% coverage as a target by rule of thumb, and a suite can meet that target with tests that check no result.

Researchers have found test suites with 100% code coverage that caught only 4% of deliberately injected bugs, which is why coverage alone is a weak signal.

What is code coverage?

Code coverage is a test measure that reports the share of a program's code that runs during its tests. A coverage tool records which parts of the code execute while the tests run, e.g. lines, and divides them by the total.

Coverage counts a line as covered when the line runs. No assertion has to check what it produced. A unit test that calls a function and checks nothing still marks each line it runs as covered.

Coverage is most useful as a negative signal. An uncovered line has no test that runs it, while a covered line may or may not be checked.

What is a mutation score?

A mutation score is the share of valid mutants that a test suite detects. A mutant is a copy of the code with one small deliberate change, e.g. >= replaced with >, and it survives when every test passes. Stryker defines the score as the percentage of valid mutants that were detected, and it counts a mutant in code that no test runs as undetected.

Some tools also report a score over covered code only. PIT calls it test strength, and Stryker calls it the mutation score based on covered code. This score separates how well the tests check code from how much code they run.

Tools count these numbers in different ways, so a score is comparable only with scores from the same tool and settings. Equivalent mutants behave exactly the same as the original code, so they can make a perfect score unreachable.

When should a team use mutation score or code coverage?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent writes this function and 2 tests for an order that holds 2 items:

export function changeAddress(order, address) {
  if (address.trim() === "") {
    throw new Error("Enter a delivery address");
  }
  return { ...order, address, items: [...order.items] };
}

test("saves the new delivery address", () => {
  const updated = changeAddress(order, "12 Elm Street");
  expect(updated.address).toBe("12 Elm Street");
});

test("rejects an empty address", () => {
  expect(() => changeAddress(order, "")).toThrow();
});

The team runs both measures on the change:

  1. The coverage report shows that the tests ran every line and both branches of changeAddress.
  2. A mutation tool, e.g. Stryker, makes 10 mutants of the function and runs the tests against each one.
  3. The tests kill 7 mutants, and 3 survive.
  4. One survivor replaces [...order.items] with [], and both tests still pass on it. The address saved, but the cart emptied.
  5. The developer adds expect(updated.items).toHaveLength(2) to the first test. The tests now kill 8 of the 10 mutants, and coverage does not change.
  6. The last 2 survivors show that no test checks the error message or an address made only of spaces.

Coverage found no gap here, because the tests already ran every line. The mutation run found the missing check on the cart. A team uses coverage to find code that no test runs, and mutation score to find results that no test checks.

This example is simplified. A real project has many functions, and its mutation run takes far longer than its coverage run.

What changes when a coding agent writes the code?

A coding agent can meet a coverage target with tests that call the changed code and check nothing. Those tests kill only mutants that make the code crash or hang, so the mutation score stays low. That gap is a sign of AI-written tests that pass on broken code.

A 2025 benchmark of 3,909 real Python functions, chosen to avoid leakage from training data, reports both numbers. Unit tests written by large language models averaged 45% statement coverage and killed 40% of injected mutants. Both numbers were low.

Either number can rise without better checks. Many coverage tools leave out lines with an exclusion comment, e.g. # pragma: no cover in coverage.py, and Stryker leaves out mutants that a comment disables. An agent can also copy the code's current output into its assertions. Those tests kill mutants and still pass on the bugs the code already has.

A team checking agent-written tests can run both tools in its continuous integration and delivery (CI/CD) pipeline, with thresholds in files the agent does not edit. A diff check can flag added exclusion comments for a reviewer. A surviving mutant usually names a missing check, and its test needs an expected value from the request.

Can mutation score and code coverage be used together?

Mutation score and code coverage work together, and many mutation tools use coverage data. By default, Stryker and PIT first record which tests reach each part of the code, then run only those tests against each mutant. A mutant in code that no test runs counts as undetected, so low coverage also holds the mutation score down.

A common setup splits the work by cost:

  • Coverage on every change. Coverage costs one instrumented test run, so it can run on each change and flag code that no test reaches.
  • Mutation testing on changed code. A mutation run repeats the tests for each mutant, so it is often limited to the lines a change touches, e.g. with the --in-diff option of cargo mutants. A full run can happen on a schedule.
  • A gate on the score. Either number can act as a quality gate that fails the build. Stryker exits with an error when the score falls below its break threshold, and PIT has separate build thresholds for line coverage and mutation score.
  • A threshold that rises. A team can set the threshold near the current score and raise it as the tests improve. A stricter gate fails on any surviving mutant in changed code.

Neither measure shows that the software does what the request asked. Both judge the tests against the code as written, so a behavior that nobody wrote has no code to mutate and leaves no survivor. A test that expects the same wrong value the code returns still kills mutants, because a change to that value makes it fail. Its test oracle is wrong, so the suite can report a false pass while it scores well on both measures.

How do mutation score and code coverage compare?

The two measures differ on these points:

PointCode coverageMutation score
What it countsCode that runs during the testsInjected bugs that make a test fail
Question it answersDid any test reach this code?Would any test catch a change here?
CostOne test run with instrumentationOne test run per mutant, often only the covering tests
Can be inflated byTests with no assertions, exclusion commentsDisabling comments, expected values copied from the code
Blind spotResults that no test checksBehavior that nobody wrote
Example toolscoverage.py, Istanbul, JaCoCoStryker, PIT, mutmut

FAQs

Can a build fail on surviving mutants instead of coverage?

A build can fail on surviving mutants when the mutation tool has a score threshold, e.g. Stryker's break threshold. That gate also reflects coverage, because a mutant in code that no test runs counts as undetected. A team can keep the coverage gate as well, because it costs less to run.

What is a good mutation score?

A good mutation score is one that rises over time on code where a missed bug is costly. No single target fits every project, because tools count the score differently and equivalent mutants can make a perfect score unreachable.

Why doesn't 100% coverage mean the tests work?

Full coverage means every line ran during the tests, not that any test checked what those lines produced. A test with no assertion counts toward coverage as much as a careful test. Mutation score shows the difference, because a mutant survives when no test checks the changed behavior.

Is mutation score worth the extra CI time?

Mutation score is worth the extra CI time where a missed bug is costly or a coding agent wrote the tests. A surviving mutant names a missing check that coverage cannot show.