What is incremental (diff-scoped) mutation testing?
Incremental mutation testing is mutation testing that tests only the mutants a change could affect, not every mutant in the codebase. Its diff-scoped form plants small deliberate bugs only in the lines a change added or edited. A change usually touches few lines, so the run tests far fewer mutants, the changed copies of the program.
In a 2021 paper, Goran Petrović and colleagues describe how Google runs mutation testing incrementally, on changes while they are under review. Its service makes mutants only in changed lines that tests cover. It shows a limited number of the surviving mutants to the author and reviewers as comments on the change. A reviewer can mark one for the author to fix.
Tools use the name for two designs, and a script can build a third. A diff filter makes mutants only in the changed code. A result cache, e.g. StrykerJS incremental mode, reuses the results of an earlier run and retests only the mutants whose code or tests changed.
The technique measures the tests of each pull request in a way that code coverage cannot. Coverage shows whether the changed lines ran during the tests. A surviving mutant shows a changed line whose result no test checks.
How does incremental mutation testing work?
In continuous integration (CI), a diff-scoped run checks a pull request in five steps:
- The CI job saves the diff between the pull request and its base branch.
- The mutation tool makes mutants only in the changed code. Some tools skip changed lines that no test runs, since those mutants would survive.
- Many tools then pick the tests that cover each mutant's line, from coverage they record in a first run. Test impact analysis uses the same kind of map, and a stale map can skip a test that would fail.
- The tool runs the tests on each mutant and marks it killed when a test fails, or survived when every test passes.
- The tool reports each survivor with its file and line, and can exit with an error, which makes the run a quality gate.
The technique comes in two designs and a scripted variant:
- Diff filter. The tool tests only mutants in the changed code. cargo-mutants for Rust takes a diff file with
--in-diff, and Infection for PHP mutates only touched lines with--git-diff-lines. Mull for C and C++ does the same with itsgitDiffRefsetting. Stryker.NET's--sincetests whole files that differ from a branch or commit. - Result cache. The tool reuses each saved result until the mutant's code or tests change, so its report covers the whole codebase. StrykerJS does this with
--incremental, and PIT uses history files, an experimental feature. - Line range. A script passes the changed lines to a tool's target option, e.g.
--mutate src/checkout.js:12-30in StrykerJS.
What is an example of incremental mutation testing?
Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, acme, is written in Rust. A developer at Acme asks a coding agent to add a --status option, so acme export --status paid exports only paid orders. The agent's pull request adds a filter function:
+pub fn keep_order(order: &Order, status: Option<&str>) -> bool {
+ match status {
+ Some(wanted) => order.status == wanted,
+ None => true,
+ }
+}
The agent's only test checks that a paid order passes the paid filter. The GitHub Actions job fetches the full history with fetch-depth: 0, so origin/main exists. Then it saves the diff and runs cargo-mutants, which calls a killed mutant caught and a survivor missed, and lists both with -v:
$ git diff origin/main.. > git.diff
$ cargo mutants --no-times -v --in-diff git.diff
Found 3 mutants to test
ok Unmutated baseline
MISSED src/export.rs:13:5: replace keep_order -> bool with true
caught src/export.rs:13:5: replace keep_order -> bool with false
caught src/export.rs:14:38: replace == with != in keep_order
3 mutants tested: 1 missed, 2 caught
All 3 mutants sit in the changed function, and none in csv_line, an older export function. The survivor makes keep_order return true, so the filter keeps each order. The test still passes, and acme export --status paid would export pending orders too.
cargo-mutants exits with code 2 on a missed mutant, so the check fails. It also adds a "Missed mutant" warning at line 13 of src/export.rs. The reviewer asks for a test that drops a pending order, and the agent adds one. The rerun reports 3 mutants tested: 3 caught.
This example is simplified. A real project would also run all its mutants on a schedule, which at Acme would find a survivor in csv_line.
What changes when a coding agent writes the code?
When agents open many pull requests, a full mutation run on each one adds to the load on CI for coding agents. The mutant count of a diff-scoped run grows with the change, not the codebase. It fits on each pull request, and an agent can run it before its output says "done."
An agent writes code and tests in one task, so a test can confirm that the code runs without checking what it returns. A survivor on a changed line names that gap, the one a reviewer looks for when checking agent-written tests by hand.
An agent can also make the run pass without better tests. A skip marker, e.g. #[mutants::skip] for cargo-mutants, removes a function's mutants. A change that only weakens a test touches no line that cargo-mutants mutates, so an --in-diff run makes no mutants and exits with success.
A practical adjustment is to read survivors and added skip markers when reviewing agent pull requests. Each survivor gets a test or a written reason, and a person confirms each survivor the agent calls an equivalent mutant. A change that only edits tests needs a mode that retests the mutants they cover, e.g. Stryker.NET's --since.
What does incremental mutation testing miss?
Incremental mutation testing trades reach for speed. It misses four kinds of problem:
- Effects outside the diff. A change can leave code in another file less tested, e.g. a caller that depends on a changed return value. That caller sits in the change's blast radius but gets no mutants, which is one way coding agents break working features.
- Older gaps. A diff filter reports only on the change, so weak tests for code that nobody touched stay hidden, e.g. the
csv_linesurvivor at Acme. - Changes a cache cannot see. StrykerJS cannot detect a change to a dependency or a configuration file, e.g. a package upgrade. Its cache can then reuse a result that no longer holds.
- Removed and missing code. A mutant changes code that exists. If a change deletes a check, or leaves out one the request needed, no line holds that check for the tool to mutate.
How is incremental mutation testing different from a full run?
A full run mutates the whole codebase and gives a mutation score for all of its tests. It finds weak tests anywhere, but it needs a test run for each mutant in the codebase. Incremental mutation testing tests only the mutants a change touches, so it fits on each pull request. Teams that practice trunk-based development merge small changes at least once a day, so each diff is short and each run makes few mutants.
The cargo-mutants documentation calls incremental tests helpful for faster feedback but "not a substitute for a full test run." Teams can keep a full run on a schedule, e.g. each night. They can also run one after a dependency upgrade, because an upgrade changes no source line for a diff filter to mutate. In StrykerJS, a full run with --incremental --force also refreshes the saved results.
FAQs
Which tools support incremental mutation testing?
Several open-source mutation tools can mutate only changed code or reuse earlier results. cargo-mutants for Rust reads a diff file, and Stryker.NET compares the code with a branch. Infection for PHP can mutate only touched lines, while StrykerJS and PIT save earlier results and retest only the mutants whose code or tests changed.
Who should look at mutants that survive in a pull request?
The author of the pull request should look at its surviving mutants first and either add a test or write down why none is needed. A reviewer then checks those reasons. When a coding agent wrote the change, a person confirms any survivor that the agent calls equivalent.
How often should a full mutation run still happen?
A full mutation run can happen on a schedule, e.g. each night, and after each dependency upgrade. It finds weak tests outside the changed lines and refreshes the saved results that an incremental run reuses.
Can test impact analysis pick the tests for each mutant?
Test impact analysis can pick the tests for each mutant, and many mutation tools already work this way. They record which tests cover each line and run only those tests against a mutant on that line. A map that is out of date can skip a test that would have killed the mutant.