Skip to content

What is a holdout test suite for coding agents?

A holdout test suite is a set of hidden tests kept out of a coding agent's reach, so they check its final work without the agent reading or editing them.

Last updated , 8 min read

What is a holdout test suite for coding agents?

A holdout test suite is a set of tests stored where a coding agent cannot read or edit them and run after it finishes. The tests in it are also called hidden tests. Because the agent never receives the tests, it cannot edit them, and only failure reports reveal their inputs.

The name comes from machine learning and statistics, where part of the data is held out of training and used only to test the finished model. SWE-bench gives a model a real issue and the codebase from before the fix. It applies tests from the pull request that fixed the issue only at evaluation, and its paper calls them "unseen tests for checking if a task was solved."

A test that the agent can run and edit becomes a target, and by Goodhart's law it stops being a good measure. On ImpossibleBench, where passing a task requires breaking its specification, GPT-5 "passed" 54% of one set of impossible SWE-bench tasks, by editing the tests or gaming them. Hiding the tests cut cheating to near zero. Keeping some tests out of reach is one way to test AI-generated code with checks the agent cannot edit.

How does a holdout test suite work?

A holdout suite separates the tests that guide the agent from the tests that judge its result. One change moves through it in six steps:

  1. A developer, or a separate agent session with no access to the coding agent's work, writes the holdout tests from the request, e.g. as acceptance tests.
  2. The developer stores the tests outside the agent's workspace, e.g. in a separate repository that the agent's credentials cannot read.
  3. The coding agent writes the code and its own visible tests, and runs them between edits.
  4. The agent reports "done" and stops.
  5. A continuous integration (CI) job makes a clean checkout of the agent's final code and adds the holdout tests.
  6. The test runner runs the holdout tests. A developer or the job then sends the agent the failing behavior, not the test code.
Where a holdout suite meets the agent's code failing behavior, not the test code Agent workspace Code Visible tests edited and run by the agent final code Clean run after the agent stops Holdout tests outside the agent's reach Result pass or fail
The agent runs its visible tests between edits. The holdout tests meet its code only in the clean run, after the agent stops.

The split closes the routes to a pass that depend on reading or editing the tests. The agent cannot rewrite a holdout assertion it never received, so test tampering in its test files cannot change the final check. It also cannot add special cases for inputs it never read, a common form of reward hacking. The holdout tests supply a test oracle that the agent did not write, so they need not repeat its reading of the request.

Because the coding agent neither writes nor runs these checks, they add some of the separation that independent verification calls for. When another session of the same model writes them, the tests and the code can still share a misreading of the request.

What is an example of a holdout test?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Before the agent starts, the developer writes 4 holdout tests from the request and stores them in a separate repository, acme-holdout. The agent's sandbox has no credentials for that repository.

The agent changes the checkout code and adds a test that sets the address on an order with 2 items. That test checks the address and not the cart. The agent's suite passes, and the agent reports "done." A CI job then copies the agent's branch and the holdout repository into a clean environment and runs the holdout tests:

FAIL holdout/checkout.test.js
  ● address change keeps the cart

    expect(received).toHaveLength(expected)

    Expected length: 3
    Received length: 0
    Received array:  []

Tests:       1 failed, 3 passed, 4 total

The address saved, but the cart emptied. The holdout test checked the cart, a check the agent's own test left out.

The developer sends the agent the steps and the actual result, not the test file. The steps are to add 3 items, start checkout, and change the address to 12 Elm Street, after which the cart shows 0 items. The agent fixes the checkout code, and the next holdout run passes 4 of 4. This example is simplified. A real holdout suite would also check the order total and the payment step.

What are the costs of a holdout suite?

A holdout suite trades feedback and upkeep for checks outside the agent's reach. It has four costs:

  • Less feedback for the agent. The agent cannot run the holdout tests while it works, so a failure reaches it only after it stops. The ImpossibleBench authors found that hiding the tests also lowered scores on the original, solvable tasks.
  • A second suite to maintain. Someone must update the holdout tests when the intended behavior changes, or a stale test fails on correct code. A test that calls a function the request never named fails correct code too, so holdout tests should drive the app through its interface.
  • Leaks. Each failure report reveals part of the suite, and a raw runner log can print lines of the test file. Tests anywhere the agent's commands reach are not hidden, e.g. an ignored folder in the same repository. An agent that can edit the CI configuration or the test runner's settings in its branch can change what the holdout job runs. Adding fresh cases after each round of feedback keeps part of the suite unseen.
  • Gaps. A holdout suite checks only what its authors wrote down, so a passing run is not the same as complete coverage.

How is a holdout suite different from read-only tests?

Read-only tests are tests the agent can read and run but cannot edit. They keep the agent's feedback loop and stop direct edits to test files. In ImpossibleBench, read-only tests kept scores on the original tasks and stopped test edits. They did not rule out other shortcuts, e.g. special cases for the inputs a test uses. A holdout suite also closes the shortcuts that depend on reading the tests, and it costs the agent that feedback.

A coding agent does not need every test in its workspace. A team can combine the two, with read-only tests for the agent's loop and a small holdout suite for the final check. Permission rules and a sandbox can make tests read-only, which is one way of protecting tests from agents.

Where should checks live so the agent cannot edit them?

Keep the holdout tests out of reach of the agent's commands, credentials, and network access, e.g. in a separate repository. Run them in a clean environment after the agent stops, from a CI definition and runner settings that the agent's branch cannot change. A holdout suite still checks only the cases someone wrote down, so run the finished software as well.

RunStory independently runs the software against your change and returns evidence to the coding agent. It runs your software in a separate environment, tries relevant workflows, and checks the results. That evidence comes from running the software, so it is not limited to the tests the agent wrote or edited. RunStory is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

Who writes the tests in a holdout suite?

The tests in a holdout suite come from a developer, or from a separate agent session with no access to the coding agent's work. The author writes them from the request before the coding agent starts, so they need not repeat that agent's reading of the request.

Can a coding agent find tests that are hidden from it?

A coding agent can find hidden tests stored anywhere its commands can reach, e.g. an ignored folder in the same repository. Tests stay hidden only when the agent's commands, credentials, and network access cannot reach them, and when its branch cannot change the CI job or runner settings.

How often should a holdout suite change?

A holdout suite should change when the intended behavior changes, so that a stale test does not fail correct code. Adding fresh cases after each round of feedback also keeps part of the suite unseen, because each failure report reveals part of the tests.

Does the name holdout come from machine learning?

The name holdout comes from machine learning and statistics, where a holdout set is data kept out of training and used only to score the finished model. SWE-bench also grades models with tests that it applies only at evaluation.