What are reproduction steps in a bug report?
Reproduction steps are the part of a bug report that lists, in order, the actions and inputs that make a failure happen again. They are also called steps to reproduce or repro steps. The list starts from a stated setup, so another developer or a coding agent can follow it and get the same failure.
A bug report usually holds a title, the environment, the steps, the observed and expected results, and evidence. Observed versus expected is the part of a bug report that states what the software did next to what it should have done. Many templates call the observed result the actual result. Both results are written as facts, without guesses about the cause. Evidence is what the reporter captured when the failure happened, e.g. a screenshot.
Reproduction is usually the first thing a developer does in debugging after a failure is reported. A developer who cannot make a failure happen again has to find its cause from the evidence, e.g. a log, and cannot check a fix directly.
How do reproduction steps work?
Reproduction steps move a failure from the person who saw it to the person or agent who fixes it. The handoff has five stages:
- The reporter records the setup, also called the preconditions, which is the version, the platform, and the data the steps start from.
- The reporter writes each action as one numbered step with its exact input.
- The reporter states the observed and expected results.
- A developer or a coding agent follows the steps on a matching setup.
- The follower compares the result with the report. A match confirms the bug.
Steps are reproducible when they leave nothing for the follower to guess. Five properties do most of that work:
- A stated starting state. The steps name the account and its data, e.g. a cart that holds 2 items. A team with test data management can name a known data set instead of describing one.
- One action per step. Each step holds a single action, so a follower is less likely to skip one.
- Exact inputs. Each step quotes the literal input, e.g. "12 Elm Street" instead of "enter an address."
- A checkable end. The last step ends at a result someone can check, e.g. an error message. A stack trace or an exit code makes that result exact.
- A stated rate. The report says how often the steps fail, e.g. 4 of 4 tries. For an intermittent failure, the rate tells the follower how many runs to try, because one pass does not rule the bug out.
What is an example of reproduction steps?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." After the change, a reviewer at Acme tries the feature and files a short report that says "Changing the delivery address breaks checkout."
The developer opens the checkout on a test account whose cart is empty, which the test build allows, and changes the delivery address. The address saves and nothing else changes, so the developer closes the report as "cannot reproduce." An empty cart has nothing to lose, so the failure cannot show. The report left out the one condition the bug needs, which is a cart with items in it.
The reviewer rewrites the report. It now has a setup, one action per step, the observed and expected results, and a rate:
Setup: shop.example.com, build 4.12.0, Chrome, test account with an empty cart
1. Add the Oak Chair to the cart.
2. Add the Pine Table to the cart.
3. Select Checkout.
4. On the delivery step, select Edit address.
5. Enter "12 Elm Street" in the Delivery address field.
6. Select Save.
Observed: The address saved, but the cart emptied.
Expected: The address changes to 12 Elm Street, and the cart still holds 2 items.
Rate: 4 of 4 tries
Before sending it, the reviewer follows the rewritten steps from an empty cart and gets the failure. With 2 items in the cart, the developer gets the failure on the first try. This example is simplified. A real report would also attach evidence, e.g. the browser console log from step 6.
What changes when a coding agent writes the code?
A coding agent that fixes the bug is the reader of the steps. The agent can follow only the steps its tools can perform. A step that says "select Edit address" needs a headless browser or another tool that acts on the page.
The agent also needs a starting state it can create itself, e.g. a test account that a script fills with 2 items. An agent that cannot run the steps can still report "done," but no run of the reported steps supports it. That report is a false completion claim.
Agents also write steps. An agent that finds a bug during exploratory testing has to record its actions, because nobody else watched the run.
The practical change is to hand the agent a reproduction script. A reproduction script is a runnable script that replays the steps that trigger a bug, so a person or an agent can repeat the failure without guessing. For Acme, it can be a Playwright test whose last two checks are the expected result. The addItemsToCart helper is hypothetical:
import { test, expect } from "@playwright/test";
import { addItemsToCart } from "./helpers/cart";
test("editing the delivery address keeps the cart", async ({ page }) => {
await addItemsToCart(page, ["oak-chair", "pine-table"]);
await page.goto("https://shop.example.com/checkout");
await page.getByRole("button", { name: "Edit address" }).click();
await page.getByLabel("Delivery address").fill("12 Elm Street");
await page.getByRole("button", { name: "Save" }).click();
await expect(page.getByText("12 Elm Street")).toBeVisible();
await expect(page.getByTestId("cart-item")).toHaveCount(2);
});
The script fails before the fix and passes after it. A person or a separate test runner reruns it to verify a bug fix.
What are the limits of written reproduction steps?
Written steps record what the reporter noticed, and a failure can depend on conditions that nobody noticed. The main limits are these:
- Hidden conditions stay hidden. A condition the reporter never saw can decide whether the bug appears, e.g. a feature flag that is on for one account.
- Timing does not fit in a list. A timing bug depends on the order in which concurrent operations run, which no written step controls. The same steps can fail once and then pass on retry on the same code.
- Prose leaves room to guess. A person fills gaps from experience, and an agent fills them with values its model generates or its tools' defaults. Two followers can run different actions from the same text.
- Evidence shows one run. A video, a screenshot, or a console log shows what happened once. It supports the steps but cannot replace them, because none of it can be run against the software again.
Steps that reproduce a failure show that it happens, not why it happens. Finding the cause is the rest of debugging.
How do reproduction steps differ from a minimal reproducible example?
Reproduction steps trigger a failure in the full software, starting from the state a user was in. A minimal reproducible example is the smallest complete code, input, and setup that still shows the same bug. Stack Overflow's guide asks for code that is minimal, complete, and reproducible, and asks the author to test it before posting. Steps usually come first. A developer then builds the minimal example from them, either by removing parts the failure does not need or by starting from scratch with only what it needs.
How does RunStory help with reproduction steps?
A bug report that leaves out a condition or an exact action can end as "cannot reproduce." RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
What does a bug report include?
A bug report includes a title, the environment, the reproduction steps, the observed and expected results, and evidence. The steps let a reader make the failure happen again. The results and the evidence let the reader check that the failure is the same one.
What are AI bug reproduction steps?
AI bug reproduction steps are reproduction steps that a coding agent writes or follows. An agent can follow them when its tools can perform each step, and a reproduction script leaves the least to guess. An agent that finds a bug has to record the actions it took.
What do you do when developers cannot reproduce your bug?
When developers cannot reproduce your bug, compare their setup with yours, including the version, the platform, and the data. Then follow your own steps again from the stated starting state, add any condition or input the steps left out, and state how often the failure happens.
Should a bug report include video, screenshots, and console logs?
A bug report should include video, screenshots, and console logs when they show the failure, because they are evidence of the observed result. They support the reproduction steps but do not replace them. A recording shows what happened once, and the steps let someone make it happen again.