# What is differential testing?

Differential testing gives the same inputs to two or more implementations, or two versions of one program, and treats any difference as a possible bug.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define differential and back-to-back testing
-   Explain how comparing versions finds regressions
-   Identify when two implementations share the same bug

## Related content

-   [What is code coverage?](https://specstory.com/learning/test-quality/code-coverage)
-   [What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)
-   [What is metamorphic testing?](https://specstory.com/learning/test-quality/metamorphic-testing)
-   [What is golden file testing (golden master)?](https://specstory.com/learning/test-quality/golden-file-testing)

## What is differential testing?

Differential testing is a technique that runs the same inputs through two or more programs that should agree and flags any output that differs. The programs can be separate implementations of one specification, e.g. two JSON parsers. They can also be an old and a new version of one program. It is also called back-to-back testing.

William McKeeman [described differential testing](https://en.wikipedia.org/wiki/Differential_testing) in a 1998 paper on testing C compilers with generated programs. The International Software Testing Qualifications Board (ISTQB) [defines back-to-back testing](https://glossary.istqb.org/en_US/term/back-to-back-testing) as testing that uses a pseudo-oracle, an independently built variant of the software that gets the same inputs. Wikipedia treats the two names as synonyms, but by the ISTQB wording an earlier version of the same program does not count as an independent variant.

One program serves as the [test oracle](https://specstory.com/learning/test-quality/test-oracle) for the other, so nobody writes expected values by hand. A team can then check many generated inputs, and [code coverage](https://specstory.com/learning/test-quality/code-coverage) shows how much of the code those inputs reached. [Golden file testing](https://specstory.com/learning/test-quality/golden-file-testing) compares output with a saved copy instead of a second program. [Metamorphic testing](https://specstory.com/learning/test-quality/metamorphic-testing) runs one program on related inputs and checks a rule that links the outputs.

## How does differential testing work?

A differential test repeats five steps for each input:

1.  A generator creates an input, e.g. a random order file. When [fuzzing](https://specstory.com/learning/test-quality/fuzzing) supplies random or malformed inputs, the method is often called differential fuzzing.
2.  The test runner gives the same input to each program.
3.  The test runner records each program's output and whether the program crashed or hung.
4.  The test runner removes values that can differ without a bug, e.g. timestamps, and compares the outputs.
5.  If the outputs match, the runner moves to the next input. If they differ, or one program crashed or hung, it saves the input, often shrunk to a shorter one that still shows the problem. A person then decides which program is wrong.

Diagram: How a differential test compares two programs

Each program's output is the check for the other's. A difference marks an input to investigate, and matching outputs can still both be wrong.

Program A, called the reference, usually comes from one of these sources:

-   **An earlier version.** The code from before a change, e.g. a [refactoring](https://specstory.com/learning/code-review/refactoring), should match the changed code wherever the change was not meant to alter behavior.
-   **An independent implementation.** A separate program built to the same specification, e.g. a second C compiler, should give the same answer wherever the specification fixes one.
-   **A simple model.** A short, slow version written only for testing checks the fast version that ships, e.g. a plain loop over a list compared with an indexed search.

Comparing versions finds a [regression](https://specstory.com/learning/testing/regression-testing) without any expected values, because the old version's output stands in for them. When the versions are far apart in history, [git bisect](https://specstory.com/learning/debugging/git-bisect) can find the change that caused a difference.

## What is an example of differential testing?

Here is an illustrative example. Acme Co. sells furniture online, and its `acme` tool exports orders as CSV. A developer at Acme asks a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) to "Make the export stream orders instead of loading them into memory." The output should not change.

The agent rewrites the export code, and its tests pass. The rewritten code joins each row's fields with commas. The developer builds the main branch as `acme-old` and the agent's branch as `acme-new`. A short script gives both builds the same generated orders:

```bash
for seed in $(seq 1 200); do
  ./make-orders --seed "$seed" > orders.json
  ./acme-old export --input orders.json --format csv > old.csv
  ./acme-new export --input orders.json --format csv > new.csv
  cmp -s old.csv new.csv || echo "seed $seed: outputs differ"
done
```

The script prints one line, `seed 137: outputs differ`. The developer reruns that seed and compares the two files:

```diff
-A-1187,1,$240.00,"12 Elm Street, Apt 4"
+A-1187,1,$240.00,12 Elm Street, Apt 4
```

The generated order has a delivery address with a comma in it. The old export put that field in double quotes, as RFC 4180 describes for CSV. The agent's version left the quotes out, so the row now has an extra column. The agent's tests used only addresses without a comma.

Here the old version is right. The agent fixes the quoting, the script finds no differences, and the developer keeps the orders from seed 137 as a regression test. This example is simplified. A real comparison would also compare error messages, e.g. for an order file that is not valid JSON.

## What changes when a coding agent writes the code?

A coding agent makes a second implementation cheap, so differential testing costs less to set up. It does not make the two versions independent. When one agent writes both from the same prompt, both can carry the same misreading, and the comparison passes. This is one limit of [AI self-verification](https://specstory.com/learning/verification/ai-self-verification).

In [Dan Luu's 2026 experiment](https://danluu.com/agentic-testing/), agents that each wrote a Zstd decoder in Rust were told to use differential testing. None built two full implementations to compare, and the comparisons they wrote generally checked too little to be useful. Where a comparison could have caught a bug, the agents wrote the same logic twice and "encoded the same bug in both versions."

A practical adjustment is to use a reference that the agent did not write. When the output should stay the same, e.g. after a port to another language, the code before the agent's change is that reference. When two agents build the same feature as [parallel coding agents](https://specstory.com/learning/ai-coding/parallel-coding-agents), a difference between their outputs marks an input to check.

Keep the reference, and the list of values the comparison ignores, out of the agent's edits. Adding a value to that list hides a difference without fixing it.

## When do both implementations agree and still fail?

Matching outputs show that the programs agree on the inputs that ran, not that either output is correct. Both programs can return the same wrong result in these cases:

-   **Shared misreading.** Both follow the same wrong reading of the requirement, e.g. two tax functions that round each item's tax when the rule rounds only the total.
-   **Shared code.** Both call the same library, so a bug in that library appears in both outputs.
-   **Copied bug.** A later version that keeps an old bug matches the earlier one on each input that triggers it.
-   **Missed inputs.** The generator never produces an input that reaches the bug. A run with no differences says nothing about inputs the generator never made.

A difference has the opposite problem, because it does not say which program is wrong. The changed version may have fixed a [pre-existing bug](https://specstory.com/learning/debugging/new-bug-vs-regression), caused a regression, or changed an output in an allowed way, e.g. the order of rows from an unsorted query. Each difference needs a person or a written rule to judge it.

## How is differential testing different from N-version programming?

N-version programming is a method for fault tolerance that runs several independently written versions of a program side by side in operation. A voter compares their outputs on each input and usually accepts the majority result. Differential testing compares versions to find bugs, usually before release, and a developer fixes each bug instead of outvoting it.

Both methods rely on versions failing in different ways. In a [1986 study](https://doi.org/10.1109/TSE.1986.6312924), Knight and Leveson had many versions of one program written independently from the same specification. Tests on which more than one version failed were substantially more common than independent failures would predict. Asking several coding agents for the same function repeats that setup at low cost, with the same risk.

## FAQs

### Is differential testing the same as back-to-back testing?

Differential testing and back-to-back testing are often treated as two names for one method, which runs comparable programs on the same inputs and compares the outputs. The ISTQB definition asks for an independently built variant, while differential testing also covers an old and a new version of one program.

### What is differential fuzzing?

Differential fuzzing is differential testing that uses a fuzzer as its source of inputs. The fuzzer generates many random or malformed inputs, and each input goes to two or more programs. A difference in output, or a crash in one program, is saved as a possible bug.

### Can an old version of a program be the oracle?

An old version of a program can be the oracle for inputs whose output should not change, e.g. after a refactoring. A difference then points to a regression or to an intended change. The old version cannot show that its own output is correct, so an old bug that the new version copies passes.

### Can two coding agents' implementations be compared this way?

Two coding agents' implementations can be compared this way, and each difference marks an input worth checking. Agents given the same prompt can make the same mistake, so their agreement is weak evidence. A reference the agents did not write, e.g. the code before the change, gives a stronger check.

---

Source: [Differential testing | Back-to-back testing | SpecStory](https://specstory.com/learning/test-quality/differential-testing)
