# How many tests is too many?

Test bloat is the growth of a test suite beyond the behavior it checks, which adds run time, upkeep, and review work without catching more bugs.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Define test bloat
-   Explain its costs in CI time and review
-   Identify tests that can be merged or deleted safely

## Related content

-   [What is code coverage?](https://specstory.com/learning/test-quality/code-coverage)
-   [How to check that AI-written tests catch bugs](https://specstory.com/learning/test-quality/checking-agent-written-tests)
-   [What is AI test case generation?](https://specstory.com/learning/test-quality/ai-test-case-generation)
-   [What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

## How many tests is too many?

A test suite has too many tests when added tests cost run time, upkeep, and review work but catch no bug that the existing tests would miss. The condition is called test bloat, or test suite inflation. No fixed count marks it, because bloat depends on what each test checks, not on how many tests exist.

Each test carries costs for as long as it stays in the suite:

-   **Run time.** The test runs on each push, which lengthens each [continuous integration](https://specstory.com/learning/ci-cd/continuous-integration) (CI) run.
-   **Upkeep.** The test has to change when the code it checks changes.
-   **Review.** Each added test is more code for a reviewer to read.
-   **False failures.** Any test can turn into a [flaky test](https://specstory.com/learning/debugging/flaky-tests), which fails some runs with no bug behind it.

Bloat often builds up in [unit tests](https://specstory.com/learning/testing/unit-testing), the cheapest tests to write. High [code coverage](https://specstory.com/learning/test-quality/code-coverage) does not rule it out, because coverage ignores how many tests ran each line. Test bloat is older than [coding agents](https://specstory.com/learning/ai-coding/coding-agent), but agents make tests almost free to write.

## What causes test bloat?

Four causes recur, and most leave a test smell. The International Software Testing Qualifications Board [defines a test smell](https://glossary.istqb.org/en_US/term/test-smell) as an aspect of test materials that suggests potential problems, inefficiencies, or risks.

### Why are tests added more often than removed?

Adding a test needs only a passing run, while removing one needs evidence that it adds nothing. Few teams check whether each test still earns its run time, so the count seldom falls.

### Why do duplicate tests appear?

Duplicate tests call the same code with inputs that make no difference to it and assert the same result. They often start as a copy of the previous test. The same check can also sit at two levels of the [test pyramid](https://specstory.com/learning/testing/test-pyramid), a model whose usual advice puts each check at the lowest level that can catch its failure.

### Why do tests that check nothing stay?

A test that asserts almost nothing passes on right and wrong code alike, so nobody has a reason to look at it. Such tests often come from a coverage target. A [tautological test](https://specstory.com/learning/test-quality/tautological-test), whose expected value comes from the code under test, belongs in this group. So does a test of what the language already provides, e.g. that a `dataclass` stores its fields.

### Why do tests of implementation details cost more?

Some tests check how the code works instead of what it returns, e.g. the exact calls it makes to a mock. They fail when a developer restructures the code without changing its behavior, and fixing them catches no bug.

## What does test bloat look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme gives a coding agent the request "Let customers edit their delivery address during checkout." The prompt also asks for thorough tests, and these steps follow:

1.  The agent writes a `change_address` function and 14 tests.
2.  Nine tests send the same address and expect it back. They differ only in the order id.
3.  Three tests check only that the function returns a value.
4.  Two tests send an empty or blank address and expect a `ValueError`.
5.  The 14 tests pass, and a reviewer approves after reading a few.
6.  After the merge, a customer reports a failure. The address saved, but the cart emptied.
7.  A developer records coverage for each test. Two of them together run the same lines as all 14.

Three of the tests read as follows:

```python
def test_change_address_a1042():
    updated = change_address(Order(id="A-1042"), "12 Elm Street")
    assert updated.address == "12 Elm Street"

def test_change_address_a1043():
    updated = change_address(Order(id="A-1043"), "12 Elm Street")
    assert updated.address == "12 Elm Street"

def test_change_address_returns_order():
    assert change_address(Order(id="A-1042"), "12 Elm Street") is not None
```

Diagram: What the fourteen tests check

Most of the tests check the same behavior again, and the behavior that broke has no test. More tests of the first kind would make the suite larger without checking the cart.

The developer replaces the 9 address tests and the 3 that check nothing with one test of the rule the request implies:

```python
def test_change_address_keeps_cart():
    order = Order(id="A-1042", items=["sofa", "lamp"])
    updated = change_address(order, "12 Elm Street")
    assert updated.address == "12 Elm Street"
    assert updated.items == ["sofa", "lamp"]
```

The test fails with `AssertionError: assert [] == ['sofa', 'lamp']` until the agent fixes the function. The suite drops from 14 tests to 3 and now catches the bug. Both tests of an empty or blank address stay, because each can fail on a bug the other misses.

This example is simplified. A real suite would also check the order total.

## What changes when a coding agent writes the code?

A coding agent writes tests in the same loop as its code and usually keeps any test that passes. Some tools for [AI test case generation](https://specstory.com/learning/test-quality/ai-test-case-generation) also drop a passing test that adds no coverage, but an agent's loop has no such check.

A prompt that asks for thorough tests can then return many tests that differ only in their inputs. Those tests add review work, and a reviewer of [agent pull requests](https://specstory.com/learning/code-review/reviewing-agent-pull-requests) reads them before the code. A pull request padded this way is one form of [AI slop](https://specstory.com/learning/verification/ai-slop). Copies of one check share its blind spot, a common trait of [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass).

[Linear reported](https://linear.app/now/ci-bottleneck-reworked) that its test suites nearly quadrupled in 2026 as coding agents sped up shipping, and it had to rework CI to keep feedback fast. Growth alone is not bloat, because more code needs more tests. [CI for coding agents](https://specstory.com/learning/ci-cd/ci-for-coding-agents) covers the load on the queue.

A practical adjustment is to ask for one test per behavior the request names, and to add inputs to an existing test instead of copying it.

## How can teams prune a suite safely?

A test that has never failed is not bloat by that fact alone, because it may guard behavior nobody has broken yet. Pruning removes only tests whose checks other tests already make:

1.  List the slowest tests, e.g. with `pytest --durations=20`, and the flaky ones in the CI history.
2.  Record which test ran each line, e.g. with `dynamic_context = test_function` in coverage.py. Group the tests that run the same lines with the same assertion.
3.  Run a [mutation testing](https://specstory.com/learning/test-quality/mutation-testing) tool on that code. Many tools stop at the first failing test, so set the tool to report each one, e.g. with `disableBail` in Stryker for JavaScript.
4.  Mark a test as a candidate when each mutant it kills is also killed by a test you keep. Of exact duplicates, keep one.
5.  Merge candidates that differ only in inputs into one parametrized test, and delete those that check nothing.
6.  Remove tests in a pull request of their own, with mutation results from before and after.

Keep each test written for a fixed bug, and fix a flaky test rather than delete its check. Breaking the code on purpose, as in [checking agent-written tests](https://specstory.com/learning/test-quality/checking-agent-written-tests), shows whether a kept test can fail.

An unchanged mutation score shows only that the remaining tests catch the same injected bugs. [Test impact analysis](https://specstory.com/learning/ci-cd/test-impact-analysis) runs fewer tests on each change but removes none, so upkeep and review costs stay.

## How is test bloat different from low coverage?

Low coverage means some code never runs during the tests. Test bloat means many tests run and check the same code. Coverage rises with the first test that runs a line and then stays flat, so a bloated suite often shows a growing test count beside flat coverage.

The Acme suite ran each line of `change_address` and still missed the cart, because no test checked the items. A mutation score, the share of injected bugs that make a test fail, measures checking instead of running, but no mutant adds a behavior the code lacks.

## What should you check instead of adding more tests?

Check the behavior the request describes, e.g. that the cart keeps its items. One assertion on that result, or one run of the changed workflow, tells a team more than another test of code already tested.

A larger suite can still leave the changed workflow untried. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures to your coding agent and verifies the fix. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Should a team delete tests that never fail?

A test that has never failed should not be deleted for that reason alone, since nobody may have broken the behavior it guards yet. A team can break the covered code on purpose. A test that then fails still guards something, and one that no deliberate bug can fail is a candidate to remove.

### How do you find duplicate tests?

You find duplicate tests by comparing what each test runs and what it checks. Coverage recorded for each test groups the tests that run the same lines, and a mutation run shows which tests catch the same injected bugs. Tests in both groups with the same assertion are likely duplicates.

### Should generated tests stay in the suite?

Generated tests should stay in the suite when each one checks a behavior and fails when that behavior breaks. A generated test that repeats another test or checks almost nothing adds run time and review work. Merge it or remove it.

### What is a test smell?

A test smell is a sign in a test suite that the tests may have problems, waste effort, or carry risks. A duplicate test is one common example. A smell calls for a closer look, not for deleting the test on sight.

---

Source: [How many unit tests is too many? | Test bloat | SpecStory](https://specstory.com/learning/test-quality/test-bloat)
