What is an equivalent mutant?
An equivalent mutant is a changed copy of the code that behaves the same as the original for every input. In mutation testing, a tool makes each mutant by applying one small deliberate change to the code. Because the change never alters a result, no test can fail on the mutant and kill it.
A surviving mutant, or survivor, is any mutant that passes every test in the test suite. Stryker's documentation defines the mutation score as the percentage of valid mutants that the tests detect. A Stryker report does not separate equivalent mutants from other undetected mutants, so each one counts against the score until someone marks it.
PIT, a mutation testing tool for Java, calls them equivalent mutations and uses the term more widely. Its documentation also counts mutants that change behavior outside the scope of testing, e.g. a removed call that writes a log line. Equivalent mutants are one reason a mutation score can stay below a perfect score on code with complete code coverage.
Why can't an equivalent mutant be killed?
A test kills a mutant only when some input meets four conditions in order:
- The test runs the changed line.
- The change gives the program a different state, e.g. a different value in a variable.
- The different state reaches an output, e.g. the value a function returns.
- An assertion in the test checks that output and fails.
Research on mutation testing calls the first three conditions reachability, infection, and propagation. An equivalent mutant can meet the first condition, but no input makes it meet both the second and the third. Either the change never alters the state, or the altered state never reaches an output. Changing i < n to i != n in a loop that counts up from zero to a list's length is the first case, because both stop the loop at the same step.
More inputs and stricter assertions cannot kill an equivalent mutant. Sorting survivors into killable and equivalent is called the equivalent mutant problem. It has no general solution, because deciding whether two programs behave the same for every input is undecidable. So some survivors need a person's judgment.
What is an example of an equivalent mutant?
Here is an illustrative example. Acme Co. sells furniture online. Its cart limits the quantity of each item to the number in stock. A coding agent wrote the function in cart.js and two Jest tests in cart.test.js:
export function limitQuantity(quantity, stock) {
if (quantity > stock) {
return stock;
}
return quantity;
}
// cart.test.js
import { limitQuantity } from "./cart.js";
test("keeps a quantity within stock", () => {
expect(limitQuantity(2, 5)).toBe(2);
});
test("caps a quantity above stock", () => {
expect(limitQuantity(8, 5)).toBe(5);
});
A developer at Acme runs Stryker on cart.js. The tests kill every mutant but one:
- if (quantity > stock) {
+ if (quantity >= stock) {
PIT names this kind of mutant a changed conditional boundary, and a test at the edge value often kills it. The developer adds one, expect(limitQuantity(5, 5)).toBe(5). The added test passes on the original and on the mutant.
At 5 and 5, the original returns quantity and the mutant returns stock, and both are 5. Below the stock level both versions return quantity, and above it both return stock. The cart sends whole numbers, and no pair of whole numbers separates them, so the mutant is equivalent.
The developer keeps the boundary test, because it states the rule at the edge. Then the developer rewrites the function body as return Math.min(quantity, stock);. For whole numbers the behavior is the same, and the comparison that produced the mutant is gone. Stryker can now make a Math.max mutant for that line, and the first test kills it.
This example is simplified. A real cart would also need tests for a quantity of zero and for stock that runs out during checkout.
What changes when a coding agent writes the code?
A coding agent told to reach a perfect mutation score, or to kill each survivor, cannot finish while an equivalent mutant remains. It can still make the report look finished. A mutator is a tool's rule for one kind of small change. An added disable comment turns off a whole mutator on a line, so it also hides the killable mutants there. A test that checks the text of a log line kills a mutant that removes the log call, and it breaks on harmless edits to the message.
Language models can also label mutants as equivalent. Meta's ACH system generated mutants aimed at privacy bugs and then wrote tests to catch them. Engineers accepted 73% of those tests. Before writing any tests, ACH used a language model to filter out the mutants that the model labeled equivalent.
A model's label and an agent's disable comment are both claims that someone has to check. A practical adjustment is to prefer a rewrite that removes the mutant over a disable comment. When a comment stays, it gives a reason, and a reviewer confirms that the tests killed the other mutants on that line before the comment hid them. The same review of survivors is part of checking agent-written tests.
How do tools and teams handle equivalent mutants?
Three common methods detect or avoid equivalent mutants, and each one catches only some of them:
- Skip likely patterns. PIT does not mutate lines that call common logging frameworks, unless a team turns that feature off.
- Compare compiled code. One research method compiles the original and each mutant with optimizations on. When the two compiled programs are identical, the mutant is equivalent.
- Analyze the code. Static analysis examines code without running it. It can show that a change never reaches an output, e.g. a change to a value that is overwritten before it is read.
A reviewer handles the rest by looking for an input on which the two versions return different results. If one exists, the survivor needs a test. If none exists, the team marks the mutant or changes the code. Neither Stryker nor PIT has an equivalent status, so a team uses one of these options:
- Disable comment. Stryker takes a comment, e.g.
// Stryker disable next-line EqualityOperatorwith a reason after a colon. It reports the mutant as ignored and leaves it out of the score. - Pragma comment. mutmut skips a line that ends with
# pragma: no mutate. - Configuration. PIT settings, e.g.
excludedMethods, exclude whole methods or classes from mutation. - Rewrite. A refactoring that keeps the behavior can remove the code that produced the mutant, as it did at Acme.
A comment covers a whole mutator on a line. At Acme, disabling EqualityOperator on the if line would also have removed the <= mutant, which the first test kills. PIT's documentation warns that a build may contain equivalent mutations, so a threshold needs careful thought. A quality gate on the score needs a threshold below a perfect score, or a rule that each survivor in changed code gets a test or a reason.
How is an equivalent mutant different from a surviving mutant?
A surviving mutant is a result that a tool reports after a run, when every test passed with the change in place. An equivalent mutant is a property of the change itself, and it holds whatever tests exist. Each equivalent mutant that the tests run survives, but a survivor can also be killable. A mutant on a line that no test runs gets a separate status, no coverage, in Stryker and PIT.
A killable survivor points to a gap in the tests. Either no test uses an input that separates the two versions, or the test oracle does not check the result that changed. More unit tests with sharper assertions close that gap. Mutation and property-based testing work together here, because generated inputs can find a value that separates a killable mutant.
FAQs
Why did a mutant survive?
A mutant survives when every test passes with its change in place. Either the tests miss an input or an assertion that would separate it from the original, or the mutant is equivalent and no test can fail on it. A reviewer looks for a separating input before calling a mutant equivalent.
Should equivalent mutants count against the mutation score?
Equivalent mutants should not count against the mutation score once a reviewer has confirmed them, because no test can kill them. In Stryker, a disable comment with a written reason takes a confirmed mutant out of the score. An unconfirmed survivor keeps counting, because it may point to a missing test.
Why is a perfect mutation score so rare?
A perfect mutation score is rare because one unmarked equivalent mutant makes it unreachable, and no tool can sort every survivor. Reaching it means a person reviews each survivor and either kills it or marks it. A quality gate on the score therefore usually needs a threshold below a perfect score.
How do you mark a mutant as equivalent in a mutation tool?
Marking a mutant as equivalent usually means turning off mutation on its line, because neither Stryker nor PIT has an equivalent status. Stryker takes a disable comment with a mutator name and a reason, mutmut takes a pragma comment, and PIT excludes whole methods or classes. Each option also hides any killable mutants in the same scope.