Skip to content

What is a minimal reproducible example?

A minimal reproducible example is the smallest self-contained program, input, or set of steps that still shows a bug, so anyone can run it.

Last updated , 9 min read

What is a minimal reproducible example?

A minimal reproducible example (MRE) is the smallest complete code, input, and setup that still makes a bug appear for anyone who runs it. Every part it contains is needed for the failure, and nothing the failure needs is left out. It is also called a minimal, complete, and verifiable example (MCVE), a minimal working example (MWE), or a reprex.

Stack Overflow's guide sets three tests for an MRE. The code is minimal when it uses as little code as possible that still produces the same problem. It is complete when the question itself holds, as text, every part someone else needs. It is reproducible when the author has run it and watched the problem happen. The guide also asks for the expected behavior and the exact wording of the error message.

A minimal example is a step in debugging, usually taken after a failure has been reproduced and before its cause is known. The few lines that remain narrow the search for the cause. The example also gives the person or coding agent that fixes the bug a check that fails before the fix and passes after it.

How do you cut a failing case down?

A developer cuts a case down in a loop of small removals and reruns:

  1. The developer starts from reproduction steps that fail on every run and saves the exact failure, whether an error message or a wrong result.
  2. The developer copies the case to a scratch branch, so the cuts do not touch the real project.
  3. The developer removes one part of the case, e.g. one step.
  4. The developer reruns the case and compares the result with the saved failure.
  5. If the same failure appears, the cut stays. If it disappears or changes, the part goes back.
  6. The developer repeats steps 3 to 5 until no part can be removed without losing the failure.
  7. The developer runs the result from a clean checkout to confirm that it needs nothing else.
How a failing case is cut down Failing case failure saved Remove one part Rerun the case Same failure? yes no Keep the cut part not needed Put the part back failure needs it next cut Minimal example no part left to cut
A cut stays only when the rerun shows the same failure. A different failure counts as a lost failure, so the part goes back.

Stack Overflow's guide calls this method divide and conquer, and it also suggests restarting from scratch with only what the problem needs. The guide says that more code makes the problem harder to find. It also asks for readable code with descriptive names, rather than the shortest code possible.

When most of a large input is irrelevant, cutting it in half at each round needs fewer reruns than removing one piece at a time. Zeller and Hildebrandt's delta debugging algorithm automates that search and simplifies a failing test case to a minimal one that still fails. Property-based testing frameworks shrink the failing inputs they find, and fuzzing tools can minimize a crashing input, e.g. with libFuzzer's -minimize_crash=1 flag.

What is an example of a minimal reproducible example?

Here is an illustrative example. Acme Co. sells furniture online. A developer asked a coding agent to "Let customers edit their delivery address during checkout." A reviewer's browser steps now fail on every try, starting from a cart with an Oak Chair and a Pine Table. The address saved, but the cart emptied.

The developer cuts the case down in five rounds:

  1. The developer removes the Pine Table, and the cart still empties with 1 item.
  2. The developer removes the Oak Chair too. With an empty cart the failure stops showing, so the chair goes back.
  3. The developer replaces the browser steps with a script that calls setAddress, the function in the agent's diff. The item list comes back undefined, the same loss the browser showed as an empty cart.
  4. The developer copies the saved checkout record into the script, and the item list is still undefined.
  5. The developer runs the script from a fresh clone, where it also fails, so it needs nothing else.

The result needs only Node.js and the checkout code:

// repro.mjs: run from the repository root with `node repro.mjs`
import assert from "node:assert/strict";
import { setAddress } from "./src/checkout/address.js";

const checkout = { id: "A-1042", items: ["oak-chair"] };
const updated = setAddress(checkout, "12 Elm Street");

assert.equal(updated.address, "12 Elm Street");
assert.deepEqual(updated.items, ["oak-chair"]);

The address check passes, and Node.js stops at the item check:

AssertionError [ERR_ASSERTION]: Expected values to be strictly deep-equal:
+ actual - expected

+ undefined
- [
-   'oak-chair'
- ]

The script fails before a fix and exits with code 0 after one. This example is simplified. A real reduction often takes more rounds, e.g. when a cut changes the error instead of removing it.

What changes when a coding agent writes the code?

A coding agent can run the loop of cuts and reruns itself, because each round is one command and one result to compare. A minimal example also gives the agent that fixes the bug one short file and one command, instead of a whole app and a written report.

While cutting a case down, an agent can replace a real part with a stub it wrote, e.g. a fake database that returns a fixed cart. If the stub's cart has no item list, the reduced case shows the same failure for a reason the original never had. A fix aimed at that case can miss the real bug.

The adjustment is to check that a reduced case calls the real code from the diff and shows the same failure as the full reproduction. Any part the agent wrote should copy the values of the part it replaced, as the Acme script's checkout record does. A bug report for the agent can then carry both the minimal example and the full steps.

When does reducing hide the cause?

Each cut removes a condition, and the cause can sit in one that seemed unrelated. Reduction hides the cause in these situations:

  • A cut changes the failure. Removing a part can swap the original error for another one, e.g. an import error. The rerun has to show the same error or wrong value, not only fail. Callers in the stack trace change with each cut, so the developer compares the error and the line that raised it.
  • A cut hides the symptom. The empty cart in the Acme example stopped the failure from showing without fixing anything. A part the failure needs is not always the part that holds the cause.
  • Timing drops out. A race condition depends on the order of events that run at the same time. A smaller case runs faster and can change that order.
  • One clean run misleads. When the original fails only some of the time, a pass after a cut can be chance, as with any pass on retry.
  • A rebuilt case leaves out what nobody suspected. A case restarted from scratch holds only the parts the developer thought to add. If the bug needs another part, e.g. a cache, the rebuilt case passes.

A minimal example shows what a failure needs, not why it happens.

How is a minimal example different from reproduction steps?

Reproduction steps show that a failure happens in the full software, from the state a user was in. A minimal example keeps only the code, input, and setup that the failure needs, so it often runs without the rest of the app. Steps usually come first, and the minimal example is cut down from them or rebuilt from scratch. The steps stay useful after a fix, because only they cover the parts the reduction removed.

How do you check the full workflow after the small case passes?

A minimal example that passes after a fix shows that one small path through the code works. Bug fix verification reruns the original steps in the full software, e.g. in the browser. Regression testing then checks the behavior around the fix, and the minimal example can join that suite as a test.

RunStory's private alpha tests your software in isolated sandboxes. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It starts with CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

What is an MCVE?

An MCVE is a minimal, complete, and verifiable example, another name for a minimal reproducible example. Verifiable means the same as reproducible here, that is, the author ran the code and watched the problem happen.

Why do Stack Overflow questions need a minimal example?

Stack Overflow asks for a minimal example in debugging questions because the people answering need to reproduce the problem before they can help. The guide says that more code makes the problem harder to find. It also asks for the code as text in the question.

How small should a minimal example be?

A minimal example should be small enough that removing any remaining part loses the failure, and complete enough to run from a clean checkout. It should still be readable, with descriptive names, because the shortest possible code can be harder to follow.

Can a coding agent reduce a failing case?

A coding agent can reduce a failing case when it can rerun the case after each cut and compare the result with the saved failure. A person or a separate check should still confirm that the reduced case calls the real code and shows the original failure.