What is property-based testing?
Property-based testing (PBT) is a technique that checks a general rule, called a property, against many inputs that a tool generates. A property says what must hold for every valid input, e.g. that reversing a list twice gives back the original list. When a generated input breaks the property, the tool shrinks it to a smaller input that still fails.
The International Software Testing Qualifications Board defines property-based testing as checking results against specified relations between inputs and expected results, which is close to metamorphic testing. Most practice follows the style of QuickCheck, a Haskell library that Claessen and Hughes described in 2000. QuickCheck uses random inputs, and later versions added shrinking.
The property is the test's test oracle, the rule that separates right results from wrong ones. Most property tests are unit tests that call one function repeatedly. Code coverage counts the lines those calls ran, but only the property checks the results.
Property-based testing is related to fuzzing, which feeds a program many random or malformed inputs and usually checks only for crashes and hangs. The Hypothesis website splits the work by who does it. Generating and running inputs is fuzzing, which the framework does, and the property is the part a person writes.
How does property-based testing work?
A property-based test runs as a loop of six steps. A framework, e.g. Hypothesis for Python, runs steps 2 to 6:
- The developer writes a property and a generator that describes valid inputs, e.g. any list of product names.
- The framework generates an input from the generator. Many generators favor small inputs and edge cases, e.g. zero.
- The test runs the code with that input and checks the property with an assertion.
- The framework repeats steps 2 and 3 many times. By default, Hypothesis stops after 100 valid inputs pass.
- When an input fails, the framework shrinks it and runs the smaller versions.
- The framework reports the smallest failing input it found. Hypothesis also saves it and tries it first on the next run.
Shrinking is the step that reduces a failing input to the smallest input it can find that still fails. The framework tries simpler versions of the input, e.g. a shorter list, and keeps each one that still fails.
These open-source frameworks are well known:
- Hypothesis for Python
- QuickCheck for Haskell
- fast-check for JavaScript
- proptest for Rust
- jqwik for Java
- FsCheck for .NET
- rapid for Go
Some frameworks also generate sequences of actions and check the code after each one, often against a simple model, e.g. a plain list for the cart. This form is called stateful testing.
What is an example of property-based testing?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's own test sets the address to "12 Elm Street" and passes.
The developer adds one property from the request. Changing the address must not change what is in the cart. The test uses Hypothesis with pytest. Hypothesis calls its generators strategies:
from hypothesis import given, strategies as st
from acme.checkout import Checkout
products = st.sampled_from(["chair", "desk", "lamp"])
@given(items=st.lists(products, min_size=1), address=st.text(min_size=1))
def test_address_change_keeps_cart(items, address):
checkout = Checkout(items)
checkout.set_address(address)
assert checkout.items == items
Hypothesis generates carts and addresses, and an address with an accented letter fails. The address saved, but the cart emptied. Hypothesis shrinks the failing input and prints this report:
Failing test case: test_address_change_keeps_cart(
items=['chair'], # or any other generated value
address='\x80',
)
The shrunk address is one character, '\x80', the first character after the American Standard Code for Information Interchange (ASCII) range. The comment on items says that any cart fails the same way. Together they point to the cause. The save step writes the session as ASCII text and empties the cart when a character does not fit.
The developer adds the shrunk input to the test with Hypothesis's @example decorator, so it runs as a regression test after the fix. This example is simplified. A real checkout would need more properties, e.g. one for the order total.
What makes a good property?
A good property comes from the requirement and fails when the code is wrong. These patterns are common starting points:
- Round trip. Loading a saved value returns the original, e.g. a saved cart loads with the same items.
- Something stays true. An invariant holds for every input, e.g. an order total is never negative.
- Same result twice. Running an operation a second time changes nothing, e.g. a discount code applied twice gives the same total as once.
- Checkable result. Some results are hard to compute but simple to check, e.g. a sorted list is in order and holds the same items.
- Reference version. A simpler, slower version of the code gives the same answer, e.g. a plain loop that adds up the cart.
A property that repeats the implementation's own formula is weak. If the code and the property share one formula, they are wrong together, and the test passes. A stronger property checks a consequence that the request states, e.g. that changing the address keeps the cart.
What changes when a coding agent writes the code?
A property that a coding agent writes from its own code tends to restate that code, or to check only that nothing crashes. Both kinds pass on the code they came from, which is one reason AI-written tests can pass on broken code.
In Dan Luu's 2026 experiment, agents told to fuzz a Zstd decoder built useful structured inputs in only 10 of 160 runs. About half of those found real bugs. Structured inputs of that kind are what a property test's generators build. When told to use QuickCheck, the agents mostly wrote simple tests that checked little.
In Anthropic's research, an agent built on Claude Code wrote Hypothesis tests from code and its documentation. People reviewed each bug report before it reached a project. The authors report that deriving properties from code with subtle semantics remains difficult. When code makes an implicit assumption, they write, "only the library maintainers can decide what the correct property to test is."
A practical adjustment is to take properties from the request, not the code. A person who knows the intent writes them before the agent starts, and the agent writes the generators.
What are the limits of property-based testing?
A passing property test shows that the rule held for the inputs that ran. It does not prove the code correct. The technique has five limits:
- A property checks only what it states. A property that checks only for crashes passes when the result is wrong.
- Random search can miss rare inputs. A bug that needs one exact value, e.g. a specific postcode, may not appear in 100 tries.
- Invalid inputs waste runs. Fully random inputs often fail input validation and reach only the rejection path. Building valid, structured inputs takes work.
- Some behavior has no general rule. Whether a page layout looks right is hard to state as a property.
- Runs cost time. Each property runs many times, so it takes longer than one example.
A generator also covers only the inputs it can make. Adversarial testing by a person adds actions that the team's generators do not describe, e.g. clicking the checkout button twice. Mutation testing can measure whether the properties catch injected bugs.
How is property-based testing different from example-based testing?
Example-based testing checks chosen inputs against expected outputs that a person writes down. Property-based testing checks generated inputs against a rule. The two differ in these ways:
| Aspect | Example-based test | Property-based test |
|---|---|---|
| Inputs | A few, chosen by a person or an agent | Many, generated by a framework |
| Expected result | An exact value for each input | A rule that holds for every input |
| Failure report | The input the author chose | A small failing input, found by shrinking |
Chosen examples tend to follow the happy path. Generators also produce inputs that nobody thought to write, e.g. an address with an accented letter. Test-driven development usually starts from examples, and a team can add a property once the rule is clear.
FAQs
What should you use property testing for?
Property testing suits code whose correct behavior follows a rule across many inputs, e.g. a save step that must return what it stored. It is less useful where no general rule exists, e.g. whether a page layout looks right.
How do you come up with properties?
Properties usually come from the requirement, not from the code. The common patterns are round trips, conditions that stay true, repeated operations that change nothing, results that can be checked directly, and comparison with a simpler version.
Is it bad if properties mirror the implementation?
A property that mirrors the implementation is weak, because it shares the code's own formula and passes whether that formula is right or wrong. A comparison with a simpler, separate version is different, because the two versions do not share one formula.
How is property-based testing different from fuzzing?
Property-based testing checks a stated rule about each result. Fuzzing feeds a program large numbers of random or malformed inputs and usually checks only for crashes and hangs. Both generate inputs, and the property is the part a person adds.