Skip to content

What is AI test case generation?

AI test case generation uses a language model to write tests from code or specs, and some tools then drop tests that fail, flake, or add no coverage.

Last updated , 8 min read

What is AI test case generation?

AI test case generation is the use of a large language model (LLM) to write test cases from source code or a specification. It is also called AI test generation or LLM test generation. Some tools keep a generated test only if it builds, passes on repeated runs, and adds code coverage. A person reviews the kept tests.

In software development, a generated test case is usually test code, e.g. a unit test for one function. In quality assurance (QA) teams, it can also be a written case with steps and an expected result, generated from a requirement. A person checks each written case against its requirement.

Test generation is older than language models. Search-based tools, e.g. EvoSuite for Java, generate test code that runs more of the program. They add assertions that record its current behavior.

Filters can discard most of what a model writes. In Meta's evaluation of TestGen-LLM on Instagram code, only 25% of the test classes it tried to extend gained a test that built, passed reliably, and added coverage. For single tests, the paper's detailed results show a considerably lower success rate.

How do tools generate tests with a language model?

A tool that generates and filters tests from code can work through six steps:

  1. The tool picks a target, e.g. a function that no test covers.
  2. The tool builds a prompt from the target's code and its existing test file.
  3. The model returns candidate tests as source code.
  4. The tool adds each candidate to the test suite and runs the build and the tests.
  5. The tool drops each candidate that fails a filter. Some tools send a build error back to the model and ask for a corrected test.
  6. The tool offers the remaining tests to a person, e.g. in a pull request.

Generation can also start from a specification, e.g. a user story. The model then writes each expected result from the specification, so the test can fail when the code disagrees with it.

What is an example of a generated test?

Here is an illustrative example. Acme Co. sells furniture online. A coding agent has added a change_address function for the request "Let customers edit their delivery address during checkout." The function has no tests, so a developer at Acme runs a test generation tool on it:

  1. The tool sends change_address, the Order class, and tests/test_checkout.py to the model.
  2. The model returns 4 candidate tests.
  3. The first imports validate_address, which does not exist, so pytest reports an ImportError. The tool drops it.
  4. The second expects the cart to still hold sofa and lamp. It fails because the function empties the cart, so the tool drops a test that found the bug.
  5. The third passes but checks only the Order class, which tests/test_checkout.py already covers, so the tool drops it for adding no coverage.
  6. The fourth passes on 5 runs in a row and runs lines of change_address that no other test runs, so the tool keeps it.

The kept test reads as follows:

def test_change_address_updates_order():
    order = Order(id="A-1042", items=["sofa", "lamp"])
    updated = change_address(order, "12 Elm Street")
    assert updated.address == "12 Elm Street"
    assert updated.items == []

The model wrote the last assertion from the function's code. The address saved, but the cart emptied. The kept test treats the empty cart as correct, and it will fail when a developer fixes the bug. This example is simplified. A real tool would generate many more candidates and retry the ones that fail to build.

How are generated tests filtered?

Meta's TestGen-LLM paper describes 3 filters in order. The second makes two checks, so each generated test faces 4 checks:

  1. The test must build. In a Python project, the nearest check is that the test file imports.
  2. The test must pass once on the current code.
  3. The test must pass on each of 5 runs, to catch a flaky test.
  4. A coverage tool checks whether the test runs lines that the existing tests do not.

A test that fails any of these checks is dropped.

How a tool filters generated tests Model output candidate tests Builds? Passes? Stable? passes 5 runs Adds coverage? no no no no Dropped does not build, fails, flakes, or adds no coverage yes Kept test for review
A failing test is dropped even when the failure points to a real bug. The coverage check shows only that the kept test ran more code, not what it checked.

The second filter throws away information. A failing test can point to a real bug or to a wrong assertion. The authors write that the tool has no automated way to tell the two apart, so it discards any test that fails. The kept tests describe current behavior, which suits code that already works, e.g. code that is about to be refactored.

The filters check that a test runs and passes, not what it checks. Assertion coverage and mutation testing measure checking instead. A stricter pipeline can also require each test to fail on at least one mutant, a copy of the code with a small bug planted in it.

What changes when a coding agent writes the code?

A coding agent asked to add tests runs its own version of this pipeline. It often applies only the build and pass checks. It can also edit the code, not only the test, until the two agree. A generation tool treats the current code as correct, which is reasonable for code that has run in production. An agent's change has no such record, so AI-written tests that pass on it can record its bugs as expected behavior.

An edit that weakens the test instead of fixing the code is test tampering.

A coding agent often reads the existing tests first, and generated tests can copy their flaws. A 2026 study of four database systems found LLM-generated tests were slightly more likely to be flaky than existing tests. The flaky ones often assumed an ordering the system did not guarantee. Flakiness in the existing tests also carried over through the prompt.

A practical adjustment is to generate tests from a source the change did not write, e.g. the old code, as catching tests do. The steps for checking agent-written tests then confirm that each test fails on broken code.

What do generated tests usually miss?

Tests generated from code usually leave these gaps:

  • The request. A tool that reads only the code has no record of the request, so no filter compares an expected value with it.
  • Missing behavior. A tool writes tests for code that exists, not for a check the request implies and the code lacks, e.g. a limit on address length.
  • Strong assertions. A weak or tautological test can clear each filter, because the filters reward tests that run code and pass.
  • Connections between parts. Tools often test one function at a time, so a bug between the checkout page and the cart service can go unchecked.

Each kept test also adds run time and code to maintain, and the total can grow into test bloat. A green generated suite shows that the code still does what it did when the tests were written, not that it does what was asked.

How is test generation different from test evaluation?

Test generation writes tests. Test evaluation measures a test suite, e.g. with a mutation score, the share of injected bugs that make a test fail. TestGen-LLM uses another such measure, coverage, as its last filter, so a coverage number says little about the tests it kept. A fairer reading comes from a measure the tool did not target, e.g. a mutation score for tests filtered on coverage.

Neither step runs the built software end to end, as agentic testing does. Teams that test AI-generated code need that run too.

What evidence should sit beside generated tests?

Beside generated tests, keep at least one test whose expected value comes from the request, not the code, e.g. a cart that keeps its items after an address change. Then run the changed workflow in the built software.

RunStory runs your software, sends reproducible failures to your coding agent, and verifies the fix. It runs the software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

Can AI generate tests for code that has none?

AI test case generation can write tests for code that has none. Those tests expect what the code does, so a person should confirm each expected value before the tests guard later changes.

Should generated tests be reviewed like other code?

Generated tests need the same review as other code, plus a check of where each expected value came from. An expected value copied from the function's output passes whether the function is right or wrong, so the reviewer compares it with the request.

Which tests are worth generating first?

The tests worth generating first cover code that already works and has few tests, e.g. code that is about to be refactored. A change that nobody has checked needs tests written from the request first.

Why do generated tests fail to build or run?

Generated tests fail to build or run when the model writes code that does not match the project, e.g. an import of a function that does not exist. The build filter drops these tests, and some tools send the error back to the model for a corrected test.