What is the difference between mutation and property-based testing?
Mutation testing, also called mutation analysis, checks the tests, while property-based testing (PBT) checks the code. A mutation tool plants small bugs, called mutants, in the code and counts how many the existing tests catch. A property-based test states a rule, e.g. that changing the address keeps the cart, and checks it on many generated inputs.
A mutation run reports a mutation score and the surviving mutants, each a change that no test detected. A property test passes, or it fails with a small input that breaks the rule. One result shows where the tests are weak, and the other shows where the code breaks a rule.
The two work together. In a property-based test, the property is the test oracle, the rule that separates right results from wrong ones. Mutation testing measures whether a suite's oracles fail when the code changes, while code coverage shows only which lines ran.
What does mutation testing test?
Mutation testing tests the tests. Stryker's documentation describes the method as inserting bugs into the production code and running the tests against each mutant. A failing test kills the mutant. If every test passes, the mutant survives, which usually means the tests have a gap there.
The tool reuses the existing tests and their inputs. Most tools first check that the suite passes on the unchanged code. After that, the tool treats the current code as correct, so its report does not flag a bug that is already in the code. Some survivors are equivalent mutants, which behave exactly as the original code does, so no test can kill them.
Mutation testing also differs from fuzzing, although both run the software many times. A mutation tool changes the code, keeps the test inputs, and grades the tests. A fuzzer changes or generates the inputs, keeps the code, and usually checks the code for crashes and hangs. Mutation-based fuzzing is a kind of fuzzing, named for those changes to inputs, not a kind of mutation testing.
What does property-based testing test?
Property-based testing tests the code. The International Software Testing Qualifications Board (ISTQB) defines property-based testing as checking test results against specified relations between inputs and expected results. A framework, e.g. Hypothesis for Python, generates many inputs, runs the code on each, and checks the property each time. The Hypothesis website calls the generating and running part fuzzing, and the property is what turns that run into a check of the code.
A property often states an invariant, a condition that must hold for every valid input, e.g. an order total that is never negative. Generators also produce edge cases that people rarely write, e.g. an address of one character. Most property tests are unit tests that call one function many times.
When a property fails, the framework shrinks the input to a small case that still fails. The code or the property is wrong, and a person reads the case to find out which. A property run does not grade the other tests in the suite.
When should a team use mutation or property-based testing?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes Checkout.set_address and writes one pytest test that sets the address to "12 Elm Street" and checks it. The test passes.
The developer runs mutmut, a mutation tool. Of 27 mutants, 12 survive. One survivor drops the cart after the address is saved:
- self.items = save_session(self.items, address)
+ self.items = None
The survivor shows that no test checks the cart, not that the code is wrong. The developer writes a property from the request with Hypothesis:
from hypothesis import given, strategies as st
from acme.checkout import Checkout
products = st.sampled_from(["chair", "desk", "lamp"])
@given(items=st.lists(products, min_size=1), address=st.text(min_size=1))
def test_address_change_keeps_cart(items, address):
checkout = Checkout(items)
checkout.set_address(address)
assert checkout.items == items
The property fails on the unchanged code. Hypothesis shrinks the failing input to a cart with one chair and the address '\x80'. The save step rejects some characters and returns an empty cart.
The mutation run did not flag this bug, because it treated the current code as correct. After the fix, a mutation run with both tests kills the self.items = None mutant.
Mutation testing fits an existing suite that needs grading, and property-based testing fits code that must follow a rule across many inputs. A team that doubts its tests often starts with a mutation run, which needs no extra test code, then writes properties where the survivors point.
This example is simplified. A real project would review each survivor and write more properties.
What changes when a coding agent writes the code?
A coding agent that writes the code and its properties in one task often takes the properties from the code, not the request. A mutation run can grade those properties, but its score exposes only one of two common weaknesses.
A property that checks only that nothing crashes usually kills few mutants, so the score stays low. A property that copies the code's formula often kills many, because a changed line stops matching the copy. That property still passes when the formula itself is wrong, since the copy shares the mistake. The score cannot separate it from a property taken from the request. Both kinds are reasons AI-written tests can pass on broken code.
An agent can also make a failing property pass without a fix. It can narrow the generator, e.g. to st.text(alphabet=string.ascii_letters), or filter inputs with assume(). With plain letters as addresses, the Acme property passes on the broken save step, and the self.items = None mutant is still killed.
A practical adjustment is to review a property's generator as closely as its assertion. Reviewers check any diff that narrows a generator, adds assume(), or lowers max_examples, the number of valid inputs each property tries.
Can mutation and property-based testing be used together?
Mutation and property-based testing work together, because a mutation tool runs property tests like any other test. A property kills a mutant when a generated input breaks the rule. Properties can kill mutants that example tests miss, e.g. a changed comparison at the edge of a range, when the generator reaches that edge.
Three settings keep the combination stable and affordable:
- Repeatable inputs. A random generator can kill a mutant on one run and miss it on the next. Hypothesis's
derandomizesetting makes each run try the same inputs. Hypothesis turns it on by default in continuous integration (CI), so a local mutation run needs it set by hand. - Fewer inputs. Each mutant reruns the properties that reach the changed code, and by default Hypothesis stops after 100 valid inputs pass. A settings profile that only mutation runs load can lower
max_examples, and CI can mutate only the changed code. - No shrinking. A mutation tool needs only the failure. Leaving the shrink phase out of Hypothesis's
phasessetting saves time on each killed mutant.
Neither technique shows that the software does what the request asked. A behavior that nobody stated or built has no property to fail and no code to mutate.
How do mutation and property-based testing compare?
The two techniques differ on these points:
| Point | Mutation testing | Property-based testing |
|---|---|---|
| What it checks | The test suite | The code |
| What it changes | The code, one small bug per mutant | The inputs, many generated per test |
| What a person supplies | A test suite that passes | A property and a generator |
| Result | A mutation score and a list of survivors | A pass, or a small failing input |
| Bugs already in the code | Not flagged, because the current code counts as correct | Reported when a bug breaks the property |
| Cost | One run of the covering tests per mutant, usually the slower one | Many runs of one test |
| Example tools | Stryker, PIT, mutmut | Hypothesis, QuickCheck, jqwik |
FAQs
Is mutation testing the same as fuzzing?
Mutation testing is not the same as fuzzing. Mutation testing changes the code to grade the existing tests, while fuzzing changes or generates the inputs and usually checks the code for crashes and hangs. Mutation-based fuzzing is a kind of fuzzing, named for its changes to inputs, not a kind of mutation testing.
Can property-based tests kill mutants?
Property-based tests can kill mutants, because a mutation tool runs them like any other test. A property kills a mutant when a generated input breaks the rule. A random generator can change the result between runs, so a setting that repeats the same inputs keeps the score stable.
Which should a team adopt first?
A team that already has tests and doubts them often adopts mutation testing first, because a mutation run grades the existing suite without extra test code. A team with code that must follow a clear rule across many inputs can start with a property test instead.
Is mutation testing slower than property-based testing?
Mutation testing is usually slower, because it reruns the covering tests once for each mutant. A property-based test is one test that runs its rule on many inputs. A mutation run over property tests costs the most, since each mutant reruns the properties that reach its change.