Test quality and test oracles
How to tell whether tests catch real bugs, from test oracles and code coverage to mutation testing and property-based testing.
Start here
What is code coverage?
Code coverage is the share of a program's code that runs during its tests, which shows what was executed but not whether the tests checked the results.
Concepts
What is a test oracle?
A test oracle is the source of truth a test checks a result against, e.g. a specification that states the correct output for a given input.
What is mutation testing?
Mutation testing checks a test suite by making small deliberate bugs in the code, called mutants, and counting how many of them the tests catch.
What is property-based testing?
Property-based testing states rules that should hold for every input, then checks them against many generated inputs and shrinks any failure it finds.
What are catching tests (catching JiTTests)?
Catching tests are tests generated for a specific change and meant to fail if that change introduced a bug, rather than to guard the code for years.
What is a tautological test?
A tautological test cannot fail because its expected value comes from the same code, mock, or constant as its actual result, so it passes on wrong code.
What is adversarial testing in software?
Adversarial testing is deliberately trying to break software with unexpected, hostile, or careless input and actions, instead of confirming the happy path.
What is fuzzing?
Fuzzing is feeding a program large numbers of random or malformed inputs to find the crashes, hangs, and input bugs that ordinary tests miss.
What is golden file testing (golden master)?
Golden file testing compares a program's output with a saved, approved copy called the golden file, and fails when anything in the output changes.
What is metamorphic testing?
Metamorphic testing checks how outputs should change when inputs change in a known way, so code can be tested when the exact right answer is unknown.
How-to guides
How to check that AI-written tests catch bugs
Checking AI-generated tests means showing that they fail on broken code, by running them on the old code and adding small bugs on purpose.
Comparisons
What is the difference between mutation score and code coverage?
Mutation score is the share of small injected bugs that a test suite catches, while code coverage is the share of code that runs during its tests.
Terms in this topic
- Adversarial testing
- Adversarial testing is a testing approach that deliberately tries to break software with unexpected, hostile, or careless inputs and actions, instead of confirming the expected path.
- Branch coverage
- Branch coverage is a code coverage measure that counts which outcomes of each decision point, e.g. both sides of an if statement, the tests have run.
- Catching test
- A catching test is a test that is generated for one code change and designed to fail if that change introduced a bug, then usually thrown away after review.
- Characterization test
- A characterization test is a test that records what existing code does now, right or wrong, so later changes can be checked against that recorded behavior.
- Code coverage
- Code coverage is a test measure that reports the percentage of a program's code that runs during its tests, which shows what was executed but not what was checked.
- Coverage-guided fuzzing
- Coverage-guided fuzzing is a fuzzing method that keeps the inputs which reach new code paths and mutates them further, so the fuzzer explores a program efficiently.
- Fault seeding
- Fault seeding is a technique that deliberately inserts known bugs into code to see how many of them a test suite or review process detects.
- Fuzzing
- Fuzzing is a testing technique that feeds a program large numbers of random or malformed inputs to find crashes, hangs, and bugs in input handling.
- Golden file
- A golden file is a saved, approved copy of a program's output that a test compares new output against, failing when anything in the output changes.
- Implicit oracle
- An implicit oracle is a general failure signal that marks a result as wrong without any expected value, e.g. a crash.
- Metamorphic testing
- Metamorphic testing is a technique that checks how a program's output should change when its input changes in a known way, so code can be tested without an exact expected output.
- Mutant
- A mutant is a copy of a program with one small deliberate change, e.g. a flipped comparison, that mutation testing uses to check whether the tests fail.
- Mutation score
- Mutation score is a test-strength measure that reports the share of valid mutants a test suite detects, which shows how well the tests check behavior.
- Mutation testing
- Mutation testing is a technique that checks a test suite by injecting small deliberate bugs, called mutants, and counting how many of them the tests catch.
- Partial oracle
- A partial oracle is a test oracle that checks some properties of a result when the full expected result is unknown, e.g. that an order total is never negative.
- Property-based testing
- Property-based testing is a technique that states rules that should hold for every input, then checks them against many generated inputs and shrinks any failing input.
- Shrinking
- Shrinking is a step in property-based testing that reduces a failing generated input to the smallest input it can find that still fails, so the bug is easier to understand.
- Tautological test
- A tautological test is a test that cannot fail because its expected value comes from the same code, mock, or constant as the result it checks. It passes whether the code is right or wrong.
- Test oracle
- A test oracle is the source of truth that a test compares a result against, e.g. a specification that states the correct output for a given input.
- Test oracle problem
- The test oracle problem is the difficulty of deciding whether a test passed or failed for a given input and state, which is hardest when nobody knows the correct answer.