# What is the test pyramid, and does it hold for AI-written code?

The test pyramid is a model for a test suite with many fast unit tests at the base, fewer integration tests, and a small number of end-to-end tests on top.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define the test pyramid and its layers
-   Identify how cheap agent-written tests change the shape
-   Explain alternatives to the test pyramid

## Related content

-   [What is software testing?](https://specstory.com/learning/testing/software-testing)
-   [What is the difference between unit, integration and end-to-end tests?](https://specstory.com/learning/testing/unit-vs-integration-vs-e2e-tests)
-   [What is unit testing?](https://specstory.com/learning/testing/unit-testing)
-   [What is end-to-end testing?](https://specstory.com/learning/testing/end-to-end-testing)

## What is the test pyramid, and does it hold for AI-written code?

The test pyramid, or testing pyramid, is a model of a test suite in which each higher level has fewer tests than the level below. [Unit tests](https://specstory.com/learning/testing/unit-testing) form the base, [integration tests](https://specstory.com/learning/testing/integration-testing) the middle, and a few [end-to-end tests](https://specstory.com/learning/testing/end-to-end-testing) the top. The model still fits AI-written code, but [coding agents](https://specstory.com/learning/ai-coding/coding-agent) make the base cheap and often leave the top thin.

The International Software Testing Qualifications Board (ISTQB) [defines the test pyramid](https://glossary.istqb.org/en_US/term/test-pyramid) as a graphical model with more testing at the lower levels than at the top. The pyramid describes the shape of a whole [software testing](https://specstory.com/learning/testing/software-testing) portfolio. The comparison of unit tests with [integration and end-to-end tests](https://specstory.com/learning/testing/unit-vs-integration-vs-e2e-tests) covers where teams draw the lines between levels. The model sets no fixed ratio between levels.

The shape sets how soon a team learns that a change broke something. Unit tests report soon after a save and point close to the cause. End-to-end tests report later, and a failure can sit anywhere in the stack.

A coding agent leaves those run times alone but cuts what tests cost to write. The base fills quickly, and the top, which needs the whole app and customer journeys, stays hard to fill.

## How does the test pyramid work?

Ham Vocke's [The Practical Test Pyramid](https://martinfowler.com/articles/practical-test-pyramid.html), published on martinfowler.com, credits the model to Mike Cohn's book Succeeding with Agile. Cohn's pyramid had unit tests at the base, service tests in the middle, and user interface tests at the top. Vocke keeps two of its rules, which are tests at several levels of detail and fewer tests at each level up.

The rules follow from the order in which a change meets the levels:

1.  The unit tests run first, often on each save. Each one calls a small piece of code in memory.
2.  On each push, [continuous integration (CI)](https://specstory.com/learning/ci-cd/continuous-integration) runs the unit tests again with the integration tests. These start real parts, e.g. a database, and catch mismatches between them.
3.  Before merge or release, the few end-to-end tests drive the whole app through its interface. They are the slowest level and the most likely to produce [flaky tests](https://specstory.com/learning/debugging/flaky-tests).
4.  After a deploy, a few [smoke tests](https://specstory.com/learning/testing/smoke-testing) check that the deployed app starts and its basic functions work. Many teams also run them on each new build, before slower tests. They sit near the top, because they run the whole software.
5.  When a test near the top catches a bug that the lower levels missed, the team adds a test lower down. Vocke advises pushing each check as far down as it can go, without repeating it higher up.

Diagram: The test pyramid and when each level runs

Each level up holds fewer tests, and each of those tests runs more of the real system. The slower levels run later, so most feedback comes from the base.

## What is an example of a test pyramid?

Here is an illustrative example. Acme Co. sells furniture online, and its checkout suite has this shape:

| Level | Tests | When they run |
| --- | --- | --- |
| Unit | 214 | On each save and each push |
| Integration | 31 | In CI on each push, with a real test database |
| End-to-end | 3 | Before each merge, in a browser |

A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change then moves up the pyramid:

1.  The agent edits `set_delivery_address` and adds 12 unit tests. All 226 unit tests pass in the agent's loop.
2.  On push, the 31 integration tests pass too. None of them saves an address on an order that holds items.
3.  Before merge, the checkout [journey test](https://specstory.com/learning/testing/user-journey-testing) adds 2 items to the cart, saves a delivery address, and finds 0 items on the review page.
4.  The developer adds an integration test that saves an address on an order with 2 items. It fails too, and points to the query that rewrites the order row.
5.  The agent changes the query, and both failing tests pass.

The 12 added unit tests widened the base but never checked the cart. The step 4 test now catches that bug on each push. This example is simplified. A real checkout suite would also check payment and stock levels.

## What changes when a coding agent writes the code?

The pyramid rests on cost. Unit tests are cheap to write and run, and end-to-end tests are not. A coding agent cuts the cost of writing tests, so one change can add many unit tests.

[Code coverage](https://specstory.com/learning/test-quality/code-coverage) rises with the base. Coverage counts the code that ran, though, not the results that were checked. Tests written from the same prompt as the code tend to share its gaps, which is one reason [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass) can pass on broken code.

Running an end-to-end test does not get cheaper. It still needs the whole app, and its journeys come from the request and the business, not from the code the agent wrote. So the pyramid can look healthier after each agent change while the top checks no more journeys than before.

Extra tests can still slow each push, a common strain in [CI for coding agents](https://specstory.com/learning/ci-cd/ci-for-coding-agents). [Test impact analysis](https://specstory.com/learning/ci-cd/test-impact-analysis) cuts each run to the tests a change could affect.

A practical adjustment is to read the pyramid per change. When a change adds only unit tests but alters a customer journey, a reviewer asks which higher test runs that journey.

## Is an inverted pyramid an anti-pattern?

An inverted pyramid has more tests at the top than at the base, e.g. more end-to-end tests than unit tests. Vocke warns against this shape, known as the test ice cream cone, because it is hard to maintain and slow to run. For most codebases it is an anti-pattern, because each change waits on the slowest tests, whose failures are the hardest to trace.

The right shape depends on where a program's bugs tend to sit. A thin web app over a database has little logic of its own, so its bugs tend to sit between parts. Tests that run several parts together then carry more of the evidence.

In microservices, many bugs sit at the boundaries between services. Some teams draw a honeycomb instead of a pyramid, with integration tests as the widest level. Vocke checks those boundaries with contract tests rather than with end-to-end tests that start several services. A contract test checks that a service and its callers still keep to the interface they agreed on.

The pyramid counts tests, not what they check, so a wide base of weak tests still looks balanced. It also does not say which journeys the top should hold. Vocke calls the pyramid overly simplistic, though still a good rule of thumb.

## How is the test pyramid different from the testing trophy?

The testing trophy is a shape that Kent C. Dodds proposed for JavaScript apps. In [a 2021 post](https://kentcdodds.com/blog/the-testing-trophy-and-testing-classifications), he shows its layers from the bottom as static checks, unit tests, integration tests, and end-to-end tests, with integration tests as the widest. Dodds ties the trophy to one principle, "The more your tests resemble the way your software is used, the more confidence they can give you."

Both shapes keep end-to-end tests few. They differ on the middle, and Dodds suggests that much of the difference between testing advice could come from different definitions of the terms.

## What should sit at the top of the pyramid?

The top should hold the few journeys that customers depend on, e.g. the checkout journey in the Acme example. Write them from the request, not from the code, and run them against the whole app before a change ships.

A coding agent's own tests usually stay near the base, where its session can run them. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### What is a balanced testing pyramid?

A balanced testing pyramid is a test suite in which each level up holds fewer tests, and each behavior is checked at the lowest level that can catch its failure. The levels do not repeat each other's checks, so most feedback comes quickly from the base.

### Is the test pyramid a rule or a guideline?

The test pyramid is a guideline, not a rule. The Practical Test Pyramid calls the model overly simplistic but still a good rule of thumb. Teams change the shape to fit where their bugs appear, e.g. with more integration tests when most failures sit between parts.

### Does the test pyramid apply to microservices?

The test pyramid applies to microservices, but more of the risk sits between services than inside them. Contract tests cover those boundaries by checking that a service and its callers still agree on the interface they share.

### Where do smoke tests fit in the pyramid?

Smoke tests fit near the top of the pyramid, because each one exercises the whole built or deployed software. A team keeps only a few, and runs them on a new build or right after a deploy to confirm that the software starts and its main functions work.

---

Source: [What is the test pyramid? | Testing pyramid | SpecStory](https://specstory.com/learning/testing/test-pyramid)
