What is mutation testing?
Mutation testing is a technique that measures a test suite by planting small deliberate bugs in the code and checking whether the tests fail. Each changed copy of the program is a mutant. A mutant that makes at least one test fail is killed. A mutant that passes every test survives. The technique is also called mutation analysis.
In their 2011 survey, Yue Jia and Mark Harman report that researchers had studied mutation testing, a technique built on seeded faults, for over three decades. The survey traces it to early work by Lipton, DeMillo, and Sayward. Fault seeding is a related technique that inserts known bugs into code to see how many of them a test suite or review process detects. The International Software Testing Qualifications Board (ISTQB) glossary also uses fault seeding to estimate how many real defects remain, which mutation testing does not do.
Mutation testing answers a question that code coverage cannot. Coverage reports which lines ran during the tests. Mutation testing reports whether the tests would fail if those lines were wrong.
How does mutation testing work?
A mutation testing tool checks a suite in six steps:
- The tool runs the test suite, usually the unit tests, on the unchanged code and stops if any test fails.
- Many tools also record which tests run which parts of the code.
- The tool makes mutants by applying mutation operators, the rules for small changes that Stryker and PIT call mutators.
- The tool runs the tests on each mutant, either the whole suite or only the tests that cover the changed code.
- The tool marks the mutant killed when at least one test fails, and survived when every test passes.
- The tool reports the mutation score and lists each surviving mutant with its file, line, and change.
Each mutation operator makes one kind of small change. Common kinds include:
- Boundary change. The operator moves the edge of a comparison, e.g.
>=becomes>. - Negation. The operator reverses a condition, e.g.
==becomes!=. - Arithmetic swap. The operator replaces one arithmetic operator with another, e.g.
+becomes-. - Removed statement. The operator deletes a statement, e.g. the call that saves a record.
- Changed return value. The operator makes a function return a different constant, e.g.
falseinstead oftrue.
Open-source mutation testing tools exist for most widely used languages. PIT works with Java and the JVM, and it mutates compiled bytecode instead of source files. Stryker has versions for JavaScript and TypeScript, C#, and Scala. For Python, mutmut runs each mutant against pytest tests, and Infection covers PHP.
What is an example of mutation testing?
Here is an illustrative example. Acme Co. sells furniture online. Its checkout charges $15.00 for delivery, and delivery is free for orders of $200.00 or more. A coding agent wrote the fee function and two pytest tests:
def delivery_fee(subtotal):
if subtotal >= 200.00:
return 0.00
return 15.00
def test_free_delivery():
assert delivery_fee(240.00) == 0.00
def test_paid_delivery():
assert delivery_fee(150.00) == 15.00
Both tests pass, and together they run every line of the function. A developer at Acme runs mutmut on the fee function. The mutants that change a return value make a test fail and are killed. Two mutants move the threshold and survive. The first one changes the comparison:
- if subtotal >= 200.00:
+ if subtotal > 200.00:
Both tests still pass on this mutant, because $240.00 is above the threshold and $150.00 is below it. The second survivor changes 200.00 to 201.0, and both tests pass on it for the same reason. No test checks a subtotal of exactly $200.00. At that value the original code charges $0.00 and both mutants charge $15.00. The developer adds a test at the boundary:
def test_free_delivery_at_threshold():
assert delivery_fee(200.00) == 0.00
The added test fails on both mutants, so the tool now reports them as killed. The function did not change, and only the tests got stronger. This example is simplified. A real project would run the tool on the whole checkout and review each survivor before writing a test.
What is a mutation score?
A mutation score is the percentage of valid mutants that a test suite detects. Stryker's documentation builds it from these counts:
- Detected. Detected mutants are the killed mutants plus those that made the tests time out, e.g. through an infinite loop.
- Undetected. Undetected mutants are the survivors plus the mutants that no test runs, which Stryker reports as "no coverage."
- Valid. Valid mutants are the detected and undetected mutants together, so mutants that fail to compile or crash the test run are left out.
- Score. The mutation score is the number of detected mutants divided by the number of valid mutants, times 100.
Tools differ in which mutants they count, so two tools can give the same suite different scores. Stryker also reports a score based on covered code, which leaves out the mutants that no test runs.
A high score means the suite's test oracles, the expected results that tests compare against, are strict enough to catch small changes. A surviving mutant is usually a controlled false pass, because the tests report success on code that was broken on purpose. The list of survivors is often more useful than the score, since each one names a line and a change that no test checks.
What changes when a coding agent writes the code?
When a coding agent writes a function and its tests in one task, the tests often check little about the result, e.g. that the function returns a number. They pass and raise coverage, yet mutants survive, because many wrong versions of the code pass the same tests. This is one reason AI-written tests can pass on broken code.
Language models can also write the mutants. Meta's ACH system generated mutants aimed at privacy bugs and then wrote tests to catch them. Engineers accepted 73% of those tests. A catching test is a related idea, generated for one code change to fail if that change introduced a bug.
A score can be gamed too. Stryker and mutmut let a comment turn off mutation on a line, e.g. # pragma: no mutate in mutmut, and a disabled mutant does not count against the score. An agent can also loosen a failing assertion, a pattern called test tampering, which can show up as more surviving mutants. When checking agent-written tests, teams can run a mutation tool on the files an agent changed, read each survivor, and review any added suppression comment in the diff.
What are the limits of mutation testing?
Mutation testing has four limits:
- Cost. Each mutant needs its own test run, so a full analysis takes far longer than the normal suite. It pays off most where a missed bug is expensive, and least where coverage is low, since untested mutants only repeat the coverage report.
- Equivalent mutants. Some mutants change the code without changing what it does, so no test can kill them. These equivalent mutants can make a perfect score unreachable.
- Small faults. Mutants are small single changes, while many real bugs span several lines or are missing code. The method assumes the coupling effect, the idea that tests that catch small faults also catch larger ones. A study of 851 real Java bugs found that mutants written by language models resembled real bugs more closely than rule-based mutants, but more of them failed to compile.
- Tests, not the product. A high score shows that tests detect small changes, not that the code does what the user asked for.
Teams can run the full analysis on a schedule. In the continuous integration (CI) checks on each pull request, a tool can mutate only the changed code, e.g. cargo-mutants for Rust with --in-diff.
How is mutation testing different from code coverage?
Code coverage counts the lines or branches that run during the tests, while mutation testing checks whether the tests fail when those lines are wrong. Coverage needs one test run, and mutation testing needs a run for each mutant. The two overlap, because no test can kill a mutant on a line that no test runs. High coverage with a low mutation score is a common sign of tests that run code without checking it.
FAQs
Is mutation testing useful in practice?
Mutation testing is useful in practice on code where a missed bug is expensive and on tests that a coding agent wrote. Each surviving mutant names a change that no test checks.
What are mutation operators?
Mutation operators are the rules that a mutation testing tool uses to make mutants, and each one makes one kind of small change, e.g. moving the edge of a comparison.
Is mutation testing too slow for CI?
Mutation testing is slower than a normal test run. Teams keep it practical in CI by mutating only the code a pull request changes and running the full analysis on a schedule.
What mutation testing tools exist?
Mutation testing tools exist for most widely used languages, and a team picks the one that fits its language and test runner, e.g. PIT for Java and the JVM.