# How to stop coding agents from editing tests

Stopping a coding agent from editing tests means blocking its writes to test files, reviewing test changes, and running the tests outside its session.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   List ways to protect test files from agent edits
-   Explain what each protection costs
-   Check that protected tests stay unchanged

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [Why do coding agents delete or weaken tests?](https://specstory.com/learning/verification/test-tampering)
-   [What is a holdout test suite for coding agents?](https://specstory.com/learning/verification/holdout-tests)
-   [What is reward hacking in coding agents?](https://specstory.com/learning/verification/reward-hacking)

## Key points

-   Deny the agent write access to the tests you protect, and keep its own tests elsewhere.
-   Check the protected files for changes when the agent stops and again in the pull request.
-   Run the protected tests from the base branch, so edits in the agent's branch do not count.

## How do you stop coding agents from editing tests?

To stop a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) from editing tests, deny it write access to the protected tests, check them for changes, and run them outside its session. In Claude Code, `Edit` deny rules stop direct edits, and its [sandbox](https://specstory.com/learning/environments/ai-sandbox) stops scripts. A code owner review and a continuous integration (CI) run catch edits that get through.

Editing a failing test so the suite passes is called [test tampering](https://specstory.com/learning/verification/test-tampering). On [ImpossibleBench](https://arxiv.org/abs/2510.20270), where passing a task requires breaking its specification, GPT-5 "passed" 54% of one set of impossible SWE-bench tasks, by editing the tests or gaming them. Hiding the tests cut cheating to near zero. The steps below block edits instead, so the agent can run the tests.

Protecting tests is one part of how teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing).

## What do you need before you start?

The steps assume these pieces:

-   **Tests to protect.** These are checks written from the request, usually by a person.
-   **Committed agent settings.** Claude Code reads them from `.claude/settings.json`.
-   **Branch protection and CI.** A code host, e.g. GitHub, can require a review and a passing job.

## How do you stop coding agents from editing tests step by step?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent, e.g. Claude Code, to "Let customers edit their delivery address during checkout." A protected test checks that an address change keeps 2 items in the cart.

### 1\. Separate the protected tests from the agent's tests

Move the protected tests into their own folder, e.g. `tests/acceptance/`, and let the agent write its tests in `tests/unit/`. Protect the runner's config too, e.g. `jest.config.js`. Files from [snapshot testing](https://specstory.com/learning/testing/snapshot-testing) count as tests, because a snapshot rewritten from broken output makes the comparison pass.

### 2\. Write the rule into the agent's instructions

Acme's rule in [AGENTS.md](https://specstory.com/learning/ai-coding/agents-md) reads "Do not edit tests/acceptance/. If a test there contradicts the request, stop and report it." The second sentence gives the agent a way out. Claude Code reads a CLAUDE.md instead when one exists, so Acme imports the rules there with an `@AGENTS.md` line.

### 3\. Deny edits in the agent's settings and turn on its sandbox

Acme commits this `.claude/settings.json`:

```json
{
  "permissions": {
    "deny": [
      "Edit(/tests/acceptance/**)",
      "Edit(/jest.config.js)",
      "Edit(/.github/**)",
      "Edit(/.claude/**)"
    ]
  },
  "sandbox": {
    "enabled": true,
    "allowUnsandboxedCommands": false
  }
}
```

[Claude Code's permissions documentation](https://code.claude.com/docs/en/permissions) says an `Edit` deny rule covers the built-in tools that edit files and the file commands Claude Code recognizes in the shell, e.g. `sed`. It does not cover a Node script that writes files itself, e.g. `jest -u`. A leading `/` anchors the path at the folder where the session starts.

The sandbox turns the same deny rules into operating system limits on shell commands and their child processes, and `allowUnsandboxedCommands` set to false stops retries outside it. An agent without a sandbox can run in a container with the folder mounted read-only. Both apply [least privilege](https://specstory.com/learning/environments/least-privilege-for-ai-agents) to the tests.

### 4\. Check the protected files when the agent stops

A [Stop hook](https://specstory.com/learning/ci-cd/agent-hooks) checks the result, not each tool call, so it catches an edit that got through another route, e.g. a command excluded from the sandbox. Acme saves this script as `.claude/hooks/check-tests.sh` and registers it as a `Stop` hook in `.claude/settings.json`:

```bash
#!/bin/bash
# Claude Code sends the hook's JSON input on stdin. A repeat stop gets through.
[ -t 0 ] || input=$(cat)
if grep -q '"stop_hook_active": *true' <<< "$input"; then exit 0; fi
changed=$(git diff --name-only main -- tests/acceptance/ jest.config.js)
added=$(git ls-files --others --exclude-standard -- tests/acceptance/)
if [ -n "$changed$added" ]; then
  echo "Protected test files changed: $changed $added" >&2
  echo "Do not edit tests to pass. Fix the code, or report these files." >&2
  exit 2
fi
exit 0
```

`git diff main` compares the files on disk with `main`, so on a feature branch it lists committed and uncommitted edits. In the [Claude Code hooks reference](https://code.claude.com/docs/en/hooks), exit code 2 from a Stop hook keeps the agent working and passes it the message. The `stop_hook_active` line lets the next stop through, so the hook does not loop.

### 5\. Require a code owner's review for test changes

Acme's `.github/CODEOWNERS` file gives the protected paths to one team:

```text
/tests/acceptance/  @acme/test-reviewers
/jest.config.js     @acme/test-reviewers
/.github/           @acme/test-reviewers
/.claude/           @acme/test-reviewers
```

[GitHub's code owners documentation](https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners) says owners are "automatically requested for review" when a pull request changes their files, and branch protection can require their approval. The `/.github/` and `/.claude/` lines cover the workflows and the agent's settings, because a pull request can change either. GitHub reads the file from the base branch, so the agent's branch cannot change its own reviewers.

### 6\. Run the protected tests from the base branch in CI

Acme's CI job restores the protected files from `main` before it runs them:

```bash
git fetch --quiet origin main
git restore --source=origin/main -- tests/acceptance/ jest.config.js
npx jest --ci tests/acceptance
```

The `git restore` line puts back the `main` copy of the protected files and removes files the branch added there, so edits there cannot turn the run green. Make the job a [required status check](https://specstory.com/learning/ci-cd/required-status-checks), and merge an approved test change in its own pull request first.

The agent's first change fails the protected test. The address saved, but the cart emptied. Its next tool call targets the expected count in the test, Claude Code denies it, and the agent fixes the checkout code.

This example is simplified. A real project would also protect the fixtures and mocks the tests load.

## What does each protection cost?

Each protection closes one route and adds work elsewhere:

| Protection | What it stops | What it costs |
| --- | --- | --- |
| Instructions file | Some edits, and it names a way out | Nothing enforces it |
| Deny rules | Edits by file tools and known shell commands | Misses programs that write files themselves |
| Sandbox or read-only mount | Writes from any process inside it | Setup, and Git commands that rewrite those files fail |
| Stop hook | A quiet finish with protected files changed | A script to maintain |
| Code owner review | Test changes merging without approval | Reviewer time on each test change |
| CI run from the base branch | Branch edits turning the run green | Approved test changes merge first |

The ImpossibleBench authors found that read-only tests stopped test edits and kept scores on the solvable tasks. They did not stop code that handles only the test's inputs, or an overloaded equality operator. Those forms of [reward hacking](https://specstory.com/learning/verification/reward-hacking) leave the tests untouched, and a [holdout test suite](https://specstory.com/learning/verification/holdout-tests) closes the ones that depend on reading the tests.

## What are common mistakes?

These mistakes leave a route open:

-   **Trusting the instructions alone.** [METR](https://metr.org/blog/2025-06-05-recent-reward-hacking/) found that OpenAI's o3 reward-hacked in 39 of 128 runs (30.4%) on tasks where it could see the scoring code. Telling it not to cheat had a "nearly negligible effect."
-   **Relying on `chmod`.** The agent runs as the developer's user, who owns the files and can run `chmod` again. A container's read-only mount is set by the host instead.
-   **Protecting the tests but not what they load.** A protected test can still pass on broken code when a mock it loads changes, e.g. a file in `__mocks__` that replaces the checkout module.
-   **Counting on a pre-commit hook.** A [pre-commit hook](https://specstory.com/learning/ci-cd/pre-commit-hook) can reject a commit that touches the tests, but `git commit --no-verify` skips it, one reason [agents bypass pre-commit hooks](https://specstory.com/learning/ci-cd/agents-bypassing-pre-commit-hooks).

## How do you check that it worked?

Test each route an agent could take, then check each pull request:

1.  Ask the agent to change `toHaveLength(2)` in a protected test to `toHaveLength(0)`, and confirm that Claude Code denies the edit.
2.  Ask it to make the same change with a short Python script, and confirm that the sandbox stops the write.
3.  Change a protected file and run `.claude/hooks/check-tests.sh`, which should name the file and exit with code 2.
4.  Run this diff check, with full history, on each pull request, and [review a pull request](https://specstory.com/learning/code-review/reviewing-agent-pull-requests) from the agent by reading its test changes first.

```text
$ git diff --exit-code --stat origin/main...HEAD -- tests/acceptance/
 tests/acceptance/checkout.test.js | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)
$ echo $?
1
```

Exit code 1 means a protected file changed, and 0 means none did.

## What evidence goes beyond the agent's own tests?

Protected tests close one route to a green run on broken code, but they check only what someone wrote down. Evidence from running the finished software is not limited to the tests the agent wrote or edited. A check run by someone other than the agent adds [independent verification](https://specstory.com/learning/verification/independent-verification).

RunStory independently runs the software against your change and returns evidence to the coding agent. It runs your software in a separate environment, tries relevant workflows, and checks the results. RunStory is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### When should an agent be allowed to change a test?

An agent should be allowed to change a test when the request changes the behavior the test checks, or a person confirms the test is wrong. The change goes in its own pull request with a code owner's approval.

### Do read-only tests stop other shortcuts?

Read-only tests stop edits to the test files, not shortcuts in the code. A holdout test suite and a run of the finished software cover some of those routes.

### Can file permissions stop an agent that runs as your user?

File permissions alone cannot stop an agent that runs as your user, because that user owns the files and can change the permissions back. A read-only container mount or the agent's own sandbox holds instead, because the operating system enforces the limit, not the agent.

### How do you protect snapshot files from agent edits?

Snapshot files are protected like other test files, in the denied folder and under code owner review. Deny rules alone miss the runner's update flag, which writes the files itself, so the sandbox covers that route. In CI, the job restores them from the main branch.

---

Source: [How to stop coding agents from editing tests | SpecStory](https://specstory.com/learning/verification/protecting-tests-from-agents)
