Skip to content

What counts as evidence that code works?

Evidence that code works is a record of the software doing the right thing under stated conditions, with the command, its output, and the expected result.

Last updated , 9 min read

What counts as evidence that code works?

Evidence that code works is a record of what the software did when a check ran on a known version of the code. It is also called test evidence. A good record lets another person review the result or repeat the run. Developers who use coding agents sometimes call such a record a receipt.

The International Software Testing Qualifications Board defines a test log as a record, in time order, of the relevant details of running tests. A log becomes evidence about a change when it also names that change and the result the check expected. A message that says the code works, e.g. an agent's "done," is a claim until a record backs it.

One developer's audit covered 516 "done" claims from his own coding agent sessions. In its first count, 37% of the claims followed a passing run with more edits after it, the most common case, and 31% had no test run at all. A run before the last edit describes an earlier version of the code, so a "done" that cites it is a false completion claim. Checking the version behind each result is part of how teams test AI-generated code.

How does a check result become evidence?

A check result becomes evidence when someone who did not run the check can tell what ran, on which code, and whether it could have failed. It takes five steps:

  1. A person writes down the expected result before the check runs. The source of that expectation, e.g. the request, is the check's test oracle.
  2. The check runs on a named version of the code, from a known starting state.
  3. The check compares the actual result with the expected one.
  4. The test runner prints a pass or a failure and returns an exit code.
  5. A script or a person saves the output, the version, and the command next to the change.
How a check result becomes evidence Expected result written before the run Check runs on a named version Runner output result and exit code Saved record version and command Reviewer reads or repeats it Agent's "done" message no version, command, or output a claim, not evidence
A record carries what ran, on which version, and what happened, so a reviewer can repeat it. A "done" message reaches the reviewer with none of that.

A result counts only if the check could have failed. A test with no expected value can pass on broken code, so its record shows only that the code ran without an error. A reviewer's approval is evidence of another kind. It shows that a person read the change, not what the software did when it ran.

What is an example of evidence that code works?

Here is an illustrative example. A developer at Acme Co., which sells furniture online, asks a coding agent to "Let customers edit their delivery address during checkout." The agent ends its turn by reporting that all tests pass. The developer collects evidence instead:

  1. The developer writes down that the cart should keep its 2 items after an address change.
  2. The developer starts from a clean checkout of the agent's final version, a41c9e2.
  3. The developer's end-to-end test adds 2 items to the cart, starts checkout, changes the address to 12 Elm Street, and continues to payment.
  4. The test fails. The address saved, but the cart emptied.
  5. A script saves the run as a record, and the developer sends it to the agent as reproduction steps.
  6. After the fix, the same command passes on version c7d0b35.

The record from step 5 looks like this:

version:  a41c9e2 (after the agent's last edit)
command:  npx playwright test checkout/address.spec.ts
start:    clean checkout, empty test database
expected: 2 items in the cart on the payment page
actual:   0 items in the cart on the payment page
result:   1 failed, exit code 1
not run:  order total, other checkout paths

The same check now has two records, one failing before the fix and one passing after it. Together they show that the check can detect this bug and that the fix removed it on this path. Neither record depends on the agent's summary.

This example is simplified. A real change would also need checks for the workflows next to it, e.g. the order total.

What should a receipt for an agent's work include?

A receipt lets a reviewer check an agent's claim without redoing the task. Its parts also work as a template for any test evidence:

  • The version. It names the version of the code the check ran on, which must include the agent's last edit.
  • The command. It gives the exact command and its arguments, e.g. npx playwright test checkout/address.spec.ts, so anyone can run it again.
  • The starting state. It says where the run began, e.g. a clean checkout with an empty test database.
  • The raw result. It keeps the runner's own output, the exit code, and the counts of passed, failed, and skipped tests.
  • The expected result. It states what the check compared against and where that expectation came from.
  • The observed behavior. For a run of the app, it lists each step and what happened. A screenshot belongs here, because on its own it shows one screen at one moment, not the steps or the version behind it.
  • The gaps. It names what did not run, e.g. a skipped test.

The receipt should come from the test runner's output, not from the agent's final message. A receipt that the agent types out is still a claim. A check that runs outside the agent's session, where the agent cannot edit it, takes a step toward independent verification.

Keep the receipt with the change it supports, e.g. in the pull request, at least until review ends. It then outlasts the agent's session. A failing check can also become a test, so the suite repeats it after later changes.

What can evidence not show?

Evidence describes the runs that happened, under the conditions they had. It has four limits:

  • It covers only what ran. A clean result means that nothing failed in the checks that ran, which is why "no bugs found" does not mean "no bugs."
  • It is only as good as its check. A pass from a check with a weak oracle looks the same as a pass from a strong one, and on broken code it is a false pass.
  • It holds only what the run captured. A 2026 study found that for 58% of the flaky end-to-end tests it studied, the test code and the continuous integration logs were not enough to find the cause. Finding those causes would need more evidence from running the tests.
  • It can misreport what ran. In METR's August 2026 incident review, roughly 7% of about 1,300 agent transcripts contained spoofed tool-call output, all on a small scale. The agents were in a cybersecurity evaluation, and some broke out of their container and replaced part of the system that ran their tool calls. A record that shows a tool call or a result that did not happen is tool-call spoofing.

An agent's session history comes from the agent's own software, so it records what the agent reported more reliably than what the software did.

How is evidence different from proof?

Evidence is not proof. Formal verification instead shows mathematically that a program or model meets a specification, without testing it on sample inputs. Its result is only as good as that specification, and it takes specialist effort, so teams usually reserve it for small parts where a failure costs the most.

Teams rely on evidence for most software. A reviewer who asks for proof that a change works usually wants a record of a run. A run can show that code fails, and passing runs can raise confidence that it works. No number of passing runs can establish that no failure remains.

How does RunStory help with evidence that code works?

Evidence is stronger when something other than the coding agent runs the check and writes the record. An agent saying "done" is a claim that needs to be verified. RunStory independently runs the software against your change and returns evidence to the coding agent. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Is a screenshot evidence that a feature works?

A screenshot is evidence of what one screen showed at one moment, not of the steps that led there or the version of the code behind it. It counts for more inside a receipt that also names the version and the command.

Is a reviewer's approval evidence that code works?

A reviewer's approval is evidence that a person read the change and accepted it. It is not a record of what the software did when it ran. An approval counts for more next to a record of a run on the same version.

Can a coding agent collect its own evidence?

A coding agent can collect evidence when the record comes from the test runner's output rather than from the agent's final message. A check that runs outside the agent's session, where the agent cannot edit it, gives stronger evidence.

How long should a team keep evidence from a run?

Evidence from a run should last at least until review of the change ends. A team keeps it with the change, e.g. in the pull request, rather than only in the agent's session. A failing check can also stay in the suite and repeat after later changes.