What is the difference between unit, integration and end-to-end tests?
Unit, integration, and end-to-end tests differ in how much of the software each test runs at once. A unit test checks one small piece, e.g. a function, on its own. An integration test checks that two or more parts work together. An end-to-end test runs the whole system through a user journey.
All three are levels of software testing. Each level up runs more real code, so it can catch failures between parts that a lower level replaces. It is also slower and costs more to maintain, so it usually tries fewer inputs. Its failures point less precisely at the cause.
Teams draw the lines between the levels in different places. A test that runs a function with the real classes it calls is a unit test on some teams and an integration test on others.
What is a unit test?
A unit test checks the smallest testable part of a program, e.g. one function, apart from the rest of the system. It runs in the same process as the code, with no network or database, so it usually finishes in milliseconds.
When the code under test calls a slow or outside part, the test replaces that part with a test double, e.g. a stub that returns fixed answers.
Unit tests catch mistakes inside one piece, e.g. a wrong total when a discount applies. A failing unit test points close to the broken line. A unit test cannot show that the piece works with the parts it replaced.
What is an integration test?
An integration test checks that two or more parts of a system work together, e.g. a service and its database. It runs the real code at a boundary instead of replacing that code with a test double. The boundary is often a database or another service's API. The database is usually real, but a local fake may stand in for another service.
The Practical Test Pyramid separates narrow integration tests, which check one boundary, e.g. database code, from broad ones that run against live versions of other services.
Integration tests usually run slower than unit tests, because they start real dependencies and send data between processes. They catch mismatches between parts, e.g. a query that the real database rejects. They usually stop below the user interface, so they do not check what a customer gets.
An API test against one service and its real database is usually an integration test. When it follows a whole business flow through several services, some teams call it an end-to-end test.
What is an end-to-end test?
An end-to-end test runs the whole system through its real interface, the way a customer would use it. For a web app, the test drives a browser, often a headless browser, that loads pages and acts on them, e.g. with Playwright. For a command-line tool, the test runs the built program and checks its output. The app runs with its real database and services, in an environment close to production.
End-to-end tests catch failures that appear only when every part runs together, e.g. a web app that sends checkout requests to the wrong service address. They are the slowest level, often seconds to minutes per test. They are also the level most likely to produce flaky tests, because they depend on timing and shared data. When one fails, the cause can sit anywhere in the stack, so finding it takes longer.
When should a team use each level?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes update_address in the checkout service and writes the unit test below. A reviewer adds the integration test:
# Unit test: the order store is a stub.
def test_update_address_sets_address():
store = StubOrderStore(item_count=2)
order = update_address(store, "A-1042", "12 Elm Street")
assert order.address == "12 Elm Street"
# Integration test: the order store uses a real test database.
def test_update_address_keeps_items(db):
db.create_order("A-1042", item_count=2)
update_address(OrderStore(db), "A-1042", "12 Elm Street")
assert db.load_order("A-1042").item_count == 2
The unit test passes, because the stub accepts any write and never deletes items. The integration test fails with assert 0 == 2. In the real database, the changed code's REPLACE statement deletes the order row first, and a cascade rule deletes its item rows. The address saved, but the cart emptied.
An end-to-end test would show the same failure from the customer's side. It adds 2 items in a browser, changes the address, and checks the order summary. Its failure says only that the summary was wrong.
Acme tests each failure at the lowest level that can catch it, which gives this split:
- Unit tests. Use them for rules inside one function, e.g. which address formats are valid.
- Integration tests. Use them wherever code writes to or reads from another part, e.g. the order database.
- End-to-end tests. Use a few for the journeys customers depend on, e.g. checkout from cart to confirmation.
This example is simplified. A real checkout would also need tests for payment and for stock levels.
What changes when a coding agent writes the code?
A coding agent often tests at the level its session can run, not where a change can fail. Unit tests need no running database or browser, so they fit the agent's loop between edits. In the Acme example, the agent's test ran a database write against a stub that never deletes items. Stubs of that kind let AI-written tests pass on broken code.
An agent's test can also carry the name of a level it does not reach. A test in an integration folder that replaces the database with a test double does not check the database write. Reviewers can check what each test replaces, not only its file name.
A team can give the agent's session a test database, so integration tests run inside the loop. For a change that crosses a boundary, the team can also require one test above the unit level, written from the request before the agent starts.
Can the three levels be used together?
The three levels work together in one test suite, because each catches failures the others miss. The test pyramid is the usual guide to the mix. The International Software Testing Qualifications Board (ISTQB) defines it as a model of how much testing each level gets, with more at the bottom than at the top.
Unit tests are fast enough for every save or a pre-push hook, and continuous integration (CI) runs them again with the integration tests on each push. End-to-end tests usually run before merge or release.
Smoke and regression tests name a purpose, not a scope. A smoke test is a quick check that a build starts and its basic functions work. It runs the built software, so it is usually an integration or end-to-end test. A regression test checks that behavior that worked before still works, and it can be written at any level.
Even all three levels together check only the behavior their tests name, so passing tests are evidence, not proof. The pyramid is also a heuristic. Some teams write more integration tests than unit tests, because their bugs tend to sit between parts.
How do unit, integration, and end-to-end tests compare?
The three levels differ on the same few attributes:
| Attribute | Unit test | Integration test | End-to-end test |
|---|---|---|---|
| Scope | One function or class | Two or more parts, e.g. code and a database | The whole system through its interface |
| Dependencies | Slow or outside parts replaced by test doubles | Real at the boundary under test | Real, in an environment close to production |
| Typical run time | Milliseconds | Milliseconds to seconds | Seconds to minutes |
| Catches | Logic mistakes in one piece | Mismatches between parts | Broken journeys and wrong configuration |
| Misses | How parts connect | Whole journeys and the interface | Edge cases and the exact cause |
| Cost to maintain | Low | Medium | High |
| Flaky results | Rare | Occasional | Most frequent |
| Where it usually runs | Every save and every CI run | CI on each push | Before merge or release |
FAQs
How do unit, integration, smoke, and regression tests differ?
Unit and integration tests are named for their scope, from one piece to several parts together. Smoke and regression tests are named for their purpose, and a regression test can run at any level to check that old behavior still works. A smoke test checks that a build starts, so it usually runs at the integration or end-to-end level.
Which tests should run on every change?
Unit tests should run on every local change, because they are quick enough to run on each save. CI runs them again with the integration tests on each push, and slower end-to-end tests usually run before merge or release.
Are integration tests slower than unit tests?
Integration tests are usually slower than unit tests, because each one waits on a real dependency, e.g. a database. A unit test replaces that dependency with a test double, so it usually finishes in milliseconds.
Is an API test an integration test?
An API test is usually an integration test when it calls one service that runs with its real database. Some teams count the same kind of test as end-to-end when it follows a whole business flow, e.g. checkout, across several services.