Skip to content

How to check that AI-written tests catch bugs

Checking AI-generated tests means showing that they fail on broken code, by running them on the old code and adding small bugs on purpose.

Last updated , 9 min read

Key points

  • A useful test fails when the code is wrong, so break the code and run the test.
  • Run added tests on the old code, and compare each expected value with the request.
  • Mutation testing finds code changes that no test catches, and running the software finds behavior no test covers.

How do you check that AI-written tests catch bugs?

To check whether AI-generated tests are useful, make the code wrong on purpose and confirm that the tests fail. Run added tests on the old code, break the changed lines by hand or with a mutation testing tool, and compare each expected value with the request. Then run the software to find what no test checks.

AI-written tests often pass because they check the code against itself, not the request. A high code coverage number does not settle the question, because coverage counts code that ran, not results a test checked.

A 2025 benchmark used 3,909 Python functions from real code, chosen to avoid leakage from training data. On that benchmark, unit tests written by language models averaged 45% statement coverage and killed 40% of injected mutants.

What do you need before you start?

The checks below need 4 things:

  • The request. The prompt or issue says what must change and what must stay the same.
  • The change and its base. The agent's committed branch and its starting commit let the tests run on both versions.
  • A mutation testing tool. A tool for the language, e.g. mutmut for Python, makes deliberate bugs in bulk.
  • A way to run the software. A local build or preview environment runs the changed workflow.

How do you check that AI-written tests catch bugs step by step?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds this function to src/acme/checkout.py:

def change_address(order, new_address):
    address = new_address.strip()
    if not address:
        raise ValueError("delivery address is empty")
    return Order(id=order.id, address=address)

The agent also writes 3 tests and reports that they pass.

1. List the test changes

Run git diff --name-status main...HEAD -- tests/ first. At Acme it prints one added file. Any edit to an existing test needs a reason, because weakening a test so it passes is test tampering. Reviewing agent pull requests can start the same way.

2. Read what each test asserts

The agent's test file holds 3 tests:

import pytest
from acme.checkout import Order, change_address

def test_change_address():
    order = Order(id="A-1042", items=["sofa", "lamp"])
    updated = change_address(order, "12 Elm Street")
    assert updated.address == "12 Elm Street"

def test_returns_an_order():
    updated = change_address(Order(id="A-1042"), " 12 Elm Street ")
    assert updated is not None

def test_blank_address_is_rejected():
    with pytest.raises(ValueError):
        change_address(Order(id="A-1042"), " ")

Ask where each expected value came from. The first test expects the address it sent, a real test oracle. Five patterns mark a test likely to pass on wrong code:

  • A weak or missing assertion. A check that a result is not None passes for almost any result.
  • A copied value. An expected value copied from the code's output matches it, bugs included, so the test is close to tautological.
  • A mocked change. A test that mocks the changed code checks the mock.
  • A missing check. No test covers state the request keeps the same, e.g. the cart.
  • Only a clean input. No test sends an empty, boundary, or invalid value.

3. Run the added tests on the old code

Commit the agent's work first, because git restore overwrites uncommitted edits. Then put the old source back, keep the added tests, and run them:

git restore --source="$(git merge-base main HEAD)" -- src/
pytest tests/test_checkout.py
git restore -- src/

The run stops with an ImportError, which shows only that the old code lacks the function. For a bug fix, the added test should fail there with the bug's symptom and pass after the fix, a fail-before, pass-after check. Some developers call this check "prove red." A test that passes on both versions does not check the fix, and the steps to verify a bug fix start from a failing run.

4. Break the changed code by hand

Change one line the way a plausible bug would, e.g. keep the spaces in the saved address:

-    return Order(id=order.id, address=address)
+    return Order(id=order.id, address=new_address)

The run still ends with 3 passed. The second test sends an address with spaces but never checks the saved value. Replacing its assertion with assert updated.address == "12 Elm Street" makes it fail on this bug.

5. Add the check the request implies

The request asks for an address change only, so the cart should keep its items. One line in the first test states that rule, and the test fails:

>       assert updated.items == ["sofa", "lamp"]
E       AssertionError: assert [] == ['sofa', 'lamp']

The address saved, but the cart emptied. The function returns an Order without the items.

6. Run the changed workflow

Unit tests call the function directly, so also run the built store. Add 2 items, change the address at checkout, and check the cart. An end-to-end test can repeat those steps after the fix. This example is simplified. A real change would also need a check on the order total.

How do you use mutation testing on agent tests?

A mutation testing tool automates step 4. It makes small changes to the code, called mutants, and reruns the tests on each. Use it on agent tests in 4 steps:

  1. Limit the run to the changed code, e.g. mutmut run "acme.checkout*" at Acme. Stryker takes a --mutate flag for the same job.
  2. Read the survivors, the mutants that no test caught, not the score.
  3. Add an assertion from the request for each survivor a customer could notice, and write down why the rest can stay.
  4. Check the diff for added comments that turn mutation off, e.g. # pragma: no mutate in mutmut, because they raise the score without better tests.

On the agent's original tests, mutmut made 9 mutants and 4 survived:

$ mutmut results
    acme.checkout.x_change_address__mutmut_3: survived
    acme.checkout.x_change_address__mutmut_4: survived
    acme.checkout.x_change_address__mutmut_5: survived
    acme.checkout.x_change_address__mutmut_6: survived
$ mutmut show acme.checkout.x_change_address__mutmut_6
# acme.checkout.x_change_address__mutmut_6: survived
--- src/acme/checkout.py
+++ src/acme/checkout.py
@@ -2,4 +2,4 @@
     address = new_address.strip()
     if not address:
         raise ValueError("delivery address is empty")
-    return Order(id=order.id, address=address)
+    return Order(id=None, address=address)

Three survivors change the error message, and the fourth sets the order id to None. Only the id survivor needs an assertion, because a customer would notice an order without its id. No mutant touched the missing items, because a tool can only change code that exists.

What changes when a coding agent writes the code?

A coding agent usually writes the code and its tests from the same context, then edits until the tests pass. A test that disagreed with the code can be changed before anyone sees it fail.

The agent can also write mutants for its own changed lines. A study of 851 real Java bugs found that mutants written by language models resembled real bugs more closely than rule-based mutants, but more of them failed to compile. A second agent in a fresh session can read the tests for missing checks, but it may share the first agent's blind spots, a limit of AI self-verification.

A practical adjustment is to ask the agent for one deliberate bug per added test, as a diff. A reviewer applies each diff and checks that the test fails.

What are common mistakes?

These mistakes make agent tests look stronger:

  • Counting coverage as the check. Researchers have found test suites with 100% code coverage that caught only 4% of deliberately injected bugs, which is why coverage alone is a weak signal.
  • Aiming tests at mutants. A test written to kill one mutant may not check what the request asked. Name each test after a behavior.
  • Accepting a reported red run. An agent's summary can say a test failed on the old code. Rerun it, because a summary is not a record of a run.

How do you check that it worked?

The routine worked when these statements are true:

  • Each added test for a bug fix fails on the code before the fix.
  • Each deliberate bug a customer could notice makes at least one test fail.
  • Each mutation survivor has an added assertion or a written reason.
  • A run of the changed workflow gives the result the request describes.

These checks show that the tests can fail on wrong code, not that the code is correct.

How do you get evidence the agent's tests did not produce?

Tests that fail on broken code can still leave a workflow unchecked. Outside evidence comes from running the software, not only its tests.

RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. It runs the software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

What should a reviewer look for in a test file?

A reviewer should first check whether the change edits an existing test, then where each expected value came from. Warning signs are weak assertions, values copied from the code's output, mocks of the changed code, no check on unchanged state, and only clean inputs.

Can one agent check another agent's tests?

One agent can check another agent's tests by reading them for missing checks or by writing mutants for the changed code. A second agent on the same model can share the first one's blind spots, so the tests still need to fail on broken code in a real run.

Would mutation testing give agents another metric to game?

Mutation testing gives an agent a score it can raise without better tests. An agent can add a comment that turns off mutation on a line, or write a test aimed at one mutant that ignores the request. Reading the survivors and the test diff shows both.

Should the same agent write the code and its tests?

The same agent can write the code and its tests. A reviewer should then apply deliberate bugs and rerun the tests, and each expected value should come from the request, not from the code's output.