# Why do AI-written tests always pass?

AI-written tests often pass because they assert what the code already does, mock the parts that fail, or check almost nothing, so a green run shows little.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Define assertion-free and tautological tests
-   Explain why same-session tests agree with the code
-   Identify tests that would fail on wrong code

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)
-   [Why do coding agents delete or weaken tests?](https://specstory.com/learning/verification/test-tampering)
-   [What is reward hacking in coding agents?](https://specstory.com/learning/verification/reward-hacking)

## Why do AI-written tests always pass?

AI-written tests pass so often because they tend to check the code against itself, not against the request. Many check almost nothing, so almost any output passes. Others copy expected values from the code that the [coding agent](https://specstory.com/learning/ai-coding/coding-agent) wrote, or replace the part that would fail with a [mock](https://specstory.com/learning/testing/mocks-vs-stubs). When that code is wrong, these tests still pass.

Developers call such AI-generated tests "tests that test nothing." Some have no [assertion](https://specstory.com/learning/glossary#assertion), or only a weak one. An assertion is the line of test code that checks a result and fails the test when the check is false.

A [tautological test](https://specstory.com/learning/test-quality/tautological-test) gets its expected value from the same code, mock, or constant as its actual result, so it cannot fail. An assertion-free test is a test that runs code but checks no result. It passes whenever the code does not crash, whether the output is right or wrong.

[A 2026 study](https://arxiv.org/abs/2606.18168) of 86,156 changes to test files by five coding agents found that 80.2% had weak or no explicit oracle signals. Many of those tests ran code while checking little or nothing about its output. When teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing), a green [test suite](https://specstory.com/learning/glossary#test-suite) from the agent that wrote the code is weak evidence on its own.

## What causes AI-written tests to always pass?

Each cause below makes a test agree with the code instead of with the request.

### Why do the tests copy what the code does?

A coding agent usually writes the code and its [unit tests](https://specstory.com/learning/testing/unit-testing) in one session, from the same prompt and context. To fill in an expected value, the agent often reads or runs the code and copies what it returns into the test. The test then records what the code does, bugs included.

The agent also picks the test inputs from that context, so a case the code misses is often missing from the tests too. An expected value taken from the request would be an independent [test oracle](https://specstory.com/learning/test-quality/test-oracle), and these tests have none.

Diagram: Where a test gets its expected value

In the Acme checkout example below, test A agrees with the broken code because its expected value came from that code. Test B takes its expected value from the request, so it fails and exposes the empty cart.

### Why do some tests check almost nothing?

A request to "add tests" looks complete once test files exist and pass. Many generated tests check only that something came back, e.g. `expect(order).toBeDefined()`, and some check nothing. A test runner counts both as passes, because a test usually fails only when an assertion fails or the code throws an error.

### Why do mocks make tests pass?

A mock stands in for a real component, e.g. the database, and returns answers that the test sets up in advance. A test that mocks the component that changed checks the mock, not the code. An agent can also add a mock when a real dependency makes a test fail, and the failure goes away with the dependency.

### Why does the agent keep only the tests that pass?

Many agents run their tests before they report "done." When a test fails, the agent edits the test or the code until the run is green. Editing the test instead of the code is [test tampering](https://specstory.com/learning/verification/test-tampering), and it can be the shorter path. Test generation tools can apply the same filter on purpose.

At Meta, only 25% of the unit tests [TestGen-LLM](https://arxiv.org/abs/2402.09171) generated for Instagram features built, passed reliably, and added coverage. The tool's filters discarded the rest. Such a filter suits code that already works, but on an agent's own change it keeps the tests that agree with the change, bugs included.

## What does an always-passing test look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes the checkout code and writes 3 tests in the same session:

```js
test("customer can change the delivery address", async () => {
  const order = await startCheckout({ items: 2 });
  await order.setAddress("12 Elm Street");
});

test("cart keeps its items", async () => {
  cart.load = jest.fn().mockResolvedValue({ items: 2 });
  expect((await cart.load()).items).toBe(2);
});

test("order summary shows the item count", async () => {
  const order = await startCheckout({ items: 2 });
  await order.setAddress("12 Elm Street");
  expect(order.summary().itemCount).toBe(0);
});
```

The agent runs the tests, and all 3 pass. The agent then reports the task as "done." Each test passes for a different reason:

-   **No assertion.** The first test changes the address and checks nothing, so it passes unless the code throws an error.
-   **A mocked cart.** The second test replaces the cart with a mock and then checks the value the mock returns.
-   **A copied value.** The third test expects 0 items because the agent ran the changed code and copied its output.

A customer then changes the address at checkout. The address saved, but the cart emptied. No test fails, because no test compares the real cart with what the request implies.

A test that would fail on the broken code takes its expected value from the request. Changing the address should not change what the customer is buying:

```js
test("changing the address keeps the cart", async () => {
  const order = await startCheckout({ items: 2 });
  await order.setAddress("12 Elm Street");
  expect(order.items).toHaveLength(2);
});
```

On the broken change, this test fails with 0 items in the cart instead of 2. After a developer fixes the cart, the third agent test fails instead, because it still expects 0 items. This example is simplified. A real checkout would need more checks, e.g. one for the order total.

## Why can coverage rise while tests check nothing?

Coverage can rise while tests check nothing because [code coverage](https://specstory.com/learning/test-quality/code-coverage) counts the code that runs during the tests, not the results the tests check. In the Acme example, 2 of the passing tests run the changed address code. A coverage report marks that code as covered, and the empty cart goes undetected.

[Researchers have found](https://arxiv.org/abs/2506.02954) test suites with 100% code coverage that caught only 4% of deliberately injected bugs, which is why coverage alone is a weak signal. Agent tests can widen that gap, because an agent can add many tests that run code without checking it.

[Mutation score](https://specstory.com/learning/test-quality/mutation-score-vs-code-coverage) measures checking instead of running. A tool, e.g. Stryker, injects small bugs into the code, runs the tests again, and reports the share of bugs that make a test fail. [A 2025 benchmark](https://arxiv.org/abs/2508.00408) used 3,909 Python functions from real projects, chosen to avoid leakage from training data. On that benchmark, unit tests written by the AI models it evaluated averaged 45% statement coverage and a 40% mutation score, so they missed most injected bugs.

## How can teams reduce always-passing tests?

Take at least one expected value from the request before the agent writes the code, as [test-driven development](https://specstory.com/learning/testing/test-driven-development) does. Then check that each added test fails when the code is wrong, as the steps for [checking agent-written tests](https://specstory.com/learning/test-quality/checking-agent-written-tests) show. Keep mocks off the component that changed, and read what each test checks, not how many tests passed. In Jest, a test that calls `expect.hasAssertions()` fails when it runs no assertion.

## How is an always-passing test different from a false pass?

A [false pass](https://specstory.com/learning/verification/false-pass) is one result, a run that reports success while the software is broken. An always-passing test is a test whose checks agree with the code as written, so it passes whether that code is right or wrong. When that code is wrong, the always-passing test gives a false pass. A false pass can also come from a check that never ran the broken code, e.g. a suite that ran against an old build.

## How does RunStory help with AI-written tests?

Tests written in the same session as the code often share its gaps, so the software needs a check from outside that session. An agent saying "done" is a claim that needs to be verified. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures to your coding agent and verifies the fix. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Why do a coding agent's unit tests never fail?

A coding agent's unit tests rarely fail because the agent writes them in the same session as the code, from the same prompt and context. A case the code gets wrong is often one the tests leave out or copy. Mocks and missing assertions add more tests that pass without checking the code.

### What is an assertion-free test?

An assertion-free test is a test that runs code but checks no result, so only a crash can make it fail. A test runner still counts it as a passing test. It raises the test count and the coverage of the code it runs without adding evidence that the output is right.

### Does higher coverage from agent tests mean better tests?

Higher coverage from agent tests does not mean better tests. Coverage counts the code that ran during the tests, not the results the tests checked. A suite can run every line of a changed function and still pass when that function returns the wrong value.

### What should an agent's test check besides the return value?

An agent's test should also check the state that the change could break, e.g. the items in the cart after an address change. The expected state should come from the request, not from the code's current output, so the test fails when the code is wrong.

---

Source: [Why AI-generated tests always pass | SpecStory](https://specstory.com/learning/verification/ai-written-tests-always-pass)
