Key points
- A coding agent works from its context and tools, so write down what a colleague would infer.
- Paste the evidence as text, and add a command that fails before the fix and passes after it.
- State what counts as done and what must not change, then run the steps yourself before sending.
How do you write a bug report a coding agent can fix?
To write a bug report a coding agent can fix, describe one failure with its observed and expected results, the version, and a starting state. Add exact steps, paste the evidence as text, and include a command that fails before the fix. Then say what counts as fixed, and run the steps yourself before sending the report.
A good bug report records one failure so that someone else can make it happen again and fix it. Its usual parts are a title, the environment, the reproduction steps, the observed and expected results, and evidence. Reproducing the failure is usually the first step of debugging, so a report that nobody can follow stalls the fix.
A coding agent reads the report as its task and can run commands. A command that fails before the fix gives the agent and the reviewer one shared check.
What do you need before you start?
Collect these before you write the report:
- The failure, run by you. Follow the steps once yourself, so the report describes a failure you saw and the conditions it needs.
- The version. Note the commit or build that failed, e.g. the output of
git rev-parse --short HEADon a clean working tree. - The raw output. Copy the error message, the stack trace of a crash, and a command's output and exit code as text.
- A starting state the agent can create. The agent needs a way to set up the data, e.g. a script that fills a test cart.
- A tool for each step. Each action needs a tool the agent can use, e.g. a browser it drives from code.
How do you write a bug report a coding agent can fix step by step?
Here is an illustrative example. Acme Co. sells furniture online. A developer asked a coding agent to "Let customers edit their delivery address during checkout." A reviewer tries the change on staging and writes a report for the agent.
1. Name the failure in the title
The title states the action and the result, e.g. "Changing the delivery address empties the cart." A title that says "Checkout is broken" names neither.
2. State the observed and expected results
The observed result is a fact, "The address saved, but the cart emptied." The expected result is a cart that still holds 2 items.
3. Record the version and the starting state
The report names commit 3f9c2ab on staging and Acme's seed script, which fills a test cart with 2 items. An agent that starts from an empty test cart cannot trigger the failure.
4. Write the steps and paste the evidence
The steps put one action on each line, with the literal input. For evidence, the reviewer copies the failing request from the browser's network panel as text:
PATCH /api/checkout/A-1042 200
Request: {"address": "12 Elm Street"}
Response: {"address": "12 Elm Street", "items": []}
The response lists no items, although the cart held 2 items. The agent can search the code for that route and compare its own run line by line. A screenshot of the empty cart shows the result, but not the request that caused it.
5. Add a command that fails
The reviewer turns the steps into a browser test, e.g. with Playwright, and adds the command that runs it with retries off:
npx playwright test tests/repro/address-keeps-cart.spec.ts --retries=0
On the current build, the command exits with code 1, and the test reports 0 items where it expects 2. A reporter who cannot write the test can have the agent write it. Before any fix, the reporter checks that the test fails on the reported symptom, 0 items where 2 are expected, and not on a selector or a timeout. When the browser is not needed, a minimal reproducible example is shorter.
6. State when the task is done
The report ends with the condition for "done." It also names the test file the agent must not edit, because an edited test can pass without a fix. The finished report follows:
Title: Changing the delivery address empties the cart
Environment: staging, commit 3f9c2ab, Chromium
Start: test account, cart seeded with 2 items (npm run seed:cart -- --items 2)
Steps: 1. Open /checkout.
2. Select Edit address.
3. Enter "12 Elm Street".
4. Select Save.
Observed: The address saved, but the cart emptied.
Expected: The address is 12 Elm Street, and the cart still holds 2 items.
Evidence: PATCH /api/checkout/A-1042 returned "items": [] (200)
Repro: npx playwright test tests/repro/address-keeps-cart.spec.ts --retries=0
Fails: 5 of 5 runs on commit 3f9c2ab (--repeat-each=5)
Done when: the repro command exits 0 and npx playwright test checkout passes
Do not edit: tests/repro/address-keeps-cart.spec.ts
Guess, not confirmed: the address handler replaces the checkout record
A team can save these fields as a template, e.g. a GitHub issue form in .github/ISSUE_TEMPLATE. This example is simplified. A real report would also attach the full server log.
What does an agent need that a person would infer?
A developer fills gaps in a report from what they know about the product. A coding agent has only its context, which includes the report, the project's instruction files, and what its tools return. Choosing that text is called context engineering. A fresh agent session also starts without the chat where the bug was discussed.
Where the report is silent, the agent works from defaults. An agent needs these written down:
- The starting state. A colleague assumes a cart with items in it. An agent needs a command that creates one.
- How to run the software. The start and seed commands can live in AGENTS.md, so reports do not repeat them.
- What counts as wrong. A person sees that an empty cart is a bug. An agent needs a value that a check can compare.
- The limits of the change. A colleague knows the payment code is out of scope. An agent needs the files it must not edit.
- The end of the task. An agent's loop ends when the model stops calling tools and says "done," so the report names the command that must pass first.
A colleague can also ask the reporter a question. An agent that runs without a person watching usually continues with what it has.
What are common mistakes?
These mistakes slow down or mislead a coding agent:
- Several failures in one report. A fix for one of them can end the agent's run while the others remain.
- A summary instead of a result. Stack Overflow's guide says that "It doesn't work" is not descriptive enough and asks for the exact wording of the error message.
- Evidence the agent cannot reach. A dashboard behind a login stays outside the agent's context. Paste the text, or give the agent a tool that reads it, e.g. a log server it calls over the Model Context Protocol.
- A whole log pasted in. A long log fills the agent's context. Paste the lines around the error, and save the full log in a file the agent can read.
- A guess written as a fact. A line that says "the cache is stale" can send the fix to the wrong code.
How do you check that it worked?
A report works when someone who has only the report can reproduce the failure. Stack Overflow's guide also suggests checking the example again in a fresh environment. For a report to an agent, check these statements:
- The reproduction command fails on a clean checkout of the named commit, with the result the report describes.
- A fresh agent session that receives only the report reproduces the failure before it changes any code.
- The "Done when" line names a command that can pass or fail, not a judgment.
After the fix, bug fix verification checks that the same command exits 0 and the file marked "Do not edit" is unchanged. Reports written in a hurry, e.g. during a bug bash, need these checks most.
How does RunStory help with bug reports for coding agents?
A report written from memory can leave out the one condition that a failure needs. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
How do you write a bug report developers can fix quickly?
A bug report developers can fix quickly covers one failure. It gives the version, the starting state, the steps, and what happened next to what should have happened. Running the steps yourself before sending the report catches a missing condition early.
What is a good bug report template?
A good bug report template has fields for the title, the version, the starting state, the steps, the observed and expected results, and the evidence. A template for a coding agent adds a command that fails, the condition that counts as done, and the test files the agent must not edit.
Which evidence should a bug report for an agent attach?
A bug report for an agent should attach the raw output of the failure as text, e.g. the stack trace, and the command that makes it fail. Text lets the agent search the code and compare its own run line by line.
How long should a bug report for an agent be?
A bug report for an agent should be as long as the facts of one failure need. Facts that do not change between bugs, e.g. the start command, belong in AGENTS.md. A long log can be cut to the lines around the error, with the full log in a file.