Skip to content

What is a definition of done for coding agents?

A definition of done for coding agents lists the evidence a change must produce, e.g. a passing run of the app, before the work counts as finished.

Last updated , 8 min read

What is a definition of done for coding agents?

A definition of done for coding agents is a shared checklist of the evidence a change needs before the team accepts it as finished. It turns "done" from a sentence the agent writes into checks that anyone can repeat, e.g. a test run that passes on the final code.

The term comes from Scrum. The 2020 Scrum Guide defines the definition of done as "a formal description of the state of the Increment when it meets the quality measures required for the product." In Scrum, work that does not meet it cannot be released. An organization can set one as a minimum standard for its Scrum Teams. Otherwise, the Scrum Team creates one, shared by every team on the product.

A coding agent's final message can say "done" when nothing has checked the final code. One developer audited 516 such claims from his own sessions and found no passing test run behind 65 to 69% of them, depending on how strictly they were counted.

An unchecked "done" is a false completion claim. A definition of done is one part of how teams test AI-generated code.

How does a definition of done work?

On a human team, developers check their own work against the list, and some items rest on judgment, e.g. "documentation is updated." A coding agent's report on its own work is generated text, not a record of what ran. So each item on an agent's list needs a check that runs apart from the agent and leaves a record. The process has six steps:

  1. The team writes the list and adds it to the agent's instructions, e.g. AGENTS.md, and to the pull request template, in files the agent cannot change.
  2. The agent edits the code for its task and runs the checks that the list names.
  3. The agent reports that the task is "done."
  4. An agent hook, e.g. a Claude Code Stop hook, or a continuous integration (CI) job runs the listed checks again on the final code.
  5. If a check fails, its output goes back to the agent, which edits the code and reports again.
  6. When the checks pass, a reviewer reads the recorded evidence, checks the judgment items, and accepts or rejects the change.
Where a definition of done is checked Agent edits the code Agent reports "done" Checks run on the final code pass Reviewer reads the evidence fail, and the output goes back to the agent
The check runs after the agent's report, not inside it. A failing check returns its output to the agent, and only a passing run reaches the reviewer.

A point in the pipeline where a change is checked against the list, e.g. before merge, is a quality gate. Some gates block the change, and others only report.

What is an example of a definition of done for a coding agent?

Here is an illustrative example. Acme Co. sells furniture online. Its definition of done for the web store has five items. An agent hook runs items 1 to 4 when the agent reports that it is finished, and a reviewer checks item 5:

1. npm test exits 0 after the agent's last edit.
2. The diff deletes no test and marks no test as skipped.
3. The store builds, starts, and loads its home page.
4. The checkout end-to-end test passes on a fresh copy of the store.
5. The pull request includes the hook's log of each command and its exit code.

A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes the checkout code, adds a unit test for the address field, and reports "Done. All tests pass."

When the hook runs items 1 to 4, the first three pass, including npm test. Item 4 fails:

checkout.e2e.ts: expected cart to have 2 items, received 0

The address saved, but the cart emptied. The agent's unit test checked the address and never checked the cart. The hook sends the failing line back to the agent, which edits the code that cleared the cart. On the next run, items 1 to 4 pass. A reviewer checks item 5 by reading the hook's log in the pull request, then merges the change.

This example is simplified. A real definition of done would also cover the team's other rules, e.g. a security review for payment code.

What evidence should a definition of done ask for?

Each item should name evidence that someone can read back after the agent stops.

In a 2026 study of 33,596 pull requests opened by coding agents, failing CI or tests was the most common code-level reason for rejection (17.6% of rejected PRs analyzed). A definition of done that asks for a passing run can find that failure before a reviewer does.

A definition of done for a coding agent can ask for these kinds of evidence:

  • A passing run on the final code. The test command ran after the agent's last edit and exited with code 0. A run before the last edit shows nothing about the code that ships.
  • Tests left intact. The diff deletes, skips, or weakens no existing test. Editing a test so that it passes is test tampering.
  • A running app. A smoke test shows that the build starts and its most basic functions work.
  • A workflow that runs end to end. An end-to-end test drives one real journey through the running system, e.g. checkout. It is one form of runtime verification.
  • A check for each acceptance criterion. Each of the task's acceptance criteria has a test or a recorded run. In spec-driven development, the criteria come from the written spec.
  • A rerun of the original failure. For a bug fix, the steps that failed before now pass, which is how teams verify a bug fix.
  • A record a reviewer can read. The hook or CI log lists the commands, their exit codes, and the relevant output, not a summary.

What are the limits of a definition of done?

A definition of done sets a minimum for each change, and it has these limits:

  • It checks only what it names. If no item runs the checkout, a broken checkout still meets the definition. Meeting the list is not the same as complete coverage.
  • An item can pass without checking anything. A test with no real assertion passes on broken code, which is a false pass. Items that name the behavior to check, not only the command to run, are harder to pass with an empty test.
  • Checks the agent can edit are weak. If the agent can change the checklist, the hook, or the existing tests, it can make items pass without fixing the code. A safer setup keeps the checklist and the hook in files the agent cannot change, and checks the diff for edits to existing tests.
  • Each item costs time. Every check adds to the wait before a change is accepted, and a long list can get skipped.
  • It does not judge the goal. A change can meet every item and still build the wrong feature.

How is a definition of done different from acceptance criteria?

A definition of done applies to every change on a product, while acceptance criteria belong to one task. For the checkout request, one criterion is that the cart keeps its items when the address changes. The definition of done asks for the checkout test to pass after any change, and a change is finished only when it meets both. The two overlap when a definition of done requires a passing check for each acceptance criterion.

How does RunStory help with a definition of done?

A definition of done is strongest when its evidence does not come from the agent that wrote the code. An agent saying "done" is a claim that needs to be verified. RunStory independently runs the software against your change and returns evidence to the coding agent. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Where do you keep a definition of done for a coding agent?

A definition of done for a coding agent belongs in the agent's instructions and the pull request template. Keep those files where the agent cannot change them, so the checks on its work stay outside its control.

Who writes the definition of done?

The team that owns the product writes the definition of done. An organization can instead set one as a standard that each team follows as a minimum. The coding agent follows the list but should not be able to edit it, because the agent would then control the checks on its own work.

Can an agent check its own definition of done?

An agent can run the checks in its definition of done, but its own report of the results is not evidence. The checks are stronger when an agent hook or a CI job runs them again on the final code and a reviewer reads the recorded output.

Is a definition of done the same as a quality gate?

A definition of done is not the same as a quality gate. The definition of done lists the evidence a change needs, and a quality gate is a checkpoint that enforces criteria at one stage, e.g. before merge. A team can use a gate to check its definition of done.