What is end-to-end testing?
End-to-end testing is a testing method that runs complete user journeys through the whole running system, from the interface to the stored data. It is also called E2E testing. Each end-to-end test acts through the same surface a user touches, e.g. a web page, and checks the result that user would get.
The International Software Testing Qualifications Board (ISTQB) defines end-to-end testing as testing business processes from start to finish under conditions close to production. Engineers often define it by technical scope instead, as a test that runs the whole stack through the user interface or a public API. Under both definitions, a test runs the assembled software, not one part alone.
Among the levels of software testing, end-to-end tests sit at the top of the test pyramid. Each one costs more to run and to keep working than a unit test, so a suite usually has far fewer of them. In return, an end-to-end test can catch a broken journey that the tests of each separate part missed.
How does an end-to-end test work?
An end-to-end test runs against the assembled software and acts only through its outside surface. A typical run of a web app test has these steps:
- The software runs with its real database and services, in an environment close to production, e.g. a staging environment.
- The test loads known seed data, e.g. a customer account with an empty cart.
- The test opens the app in a browser, often a headless browser, then clicks and types as a user would.
- The test waits for each page or response before its next action.
- The test checks what the user would see, and sometimes the stored data too, e.g. the order row in the database.
- The test's cleanup code resets the data. With test isolation, each test sets up its own state, so other tests cannot change its result.
For a command-line tool, the same run starts the built program in a shell, passes it arguments, and checks what it prints and the files it writes. Because the test touches only the outside surface, it does not need to share the software's language. A Python test can drive a web app written in Go.
What is an example of an end-to-end test?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Before the agent starts, the developer writes an end-to-end test of that journey with a browser test framework, e.g. Playwright:
import { test, expect } from "@playwright/test";
test("customer edits the delivery address at checkout", async ({ page }) => {
await page.goto("https://shop.example.com/products/oak-chair");
await page.getByLabel("Quantity").fill("2");
await page.getByRole("button", { name: "Add to cart" }).click();
await expect(page.getByTestId("cart-count")).toHaveText("2 items");
await page.getByRole("link", { name: "Checkout" }).click();
await page.getByLabel("Delivery address").fill("12 Elm Street");
await page.getByRole("button", { name: "Save address" }).click();
await expect(page.getByText("12 Elm Street")).toBeVisible();
await expect(page.getByTestId("cart-count")).toHaveText("2 items");
});
The agent changes the checkout code and adds a unit test for its setAddress function. The unit test passes, but the end-to-end test fails on its last line, because the cart shows 0 items. The address saved, but the cart emptied.
The cause sits between parts. In this version of the example, saving the address starts a fresh checkout session, and the cart service returns no items for that session. No unit test covered that wiring, because each part worked on its own.
The agent changes the session handling, and the same end-to-end test passes. This example is simplified. A real checkout suite would also cover payment and the order confirmation, each with its own seed data.
What changes when a coding agent writes the code?
A coding agent that writes an end-to-end test after its change often scripts the journey it built. The agent reads the changed code and targets the buttons and fields it finds there. If the change missed part of the request, the test misses the same part and still passes.
An agent that writes or repairs an end-to-end test can also remove parts of the system from the test. A browser test framework can replace a network call with a fixed response, e.g. with Playwright's page.route. If the Acme test stubbed the browser's cart request, it would pass, because the fixed response still lists 2 items.
A practical adjustment is to write the critical journeys from the request before the agent starts, and to keep those files out of the agent's edits. Reviewers can treat any stub added to an end-to-end test as a change to what the test covers. Running those journeys on the whole app after each change is one form of runtime verification for AI-generated code.
Why are end-to-end tests slow and flaky?
An end-to-end test is slow because it needs the whole system running, sends real requests over the network, and waits for real responses. A unit test calls one function in memory and usually finishes in milliseconds.
End-to-end tests also cost more to keep working. A test that finds an element by its place in the page, e.g. the third div in a form, breaks when the layout changes. Locators that use a role, a label, or a test ID, as the example does, survive layout changes.
The network calls and waits that slow these tests also make them flaky. A flaky test passes and fails on the same code. Common causes in end-to-end tests are these:
- Timing. The test waits a fixed time, then checks the page before a response arrives, e.g. a confirmation page that usually loads within 200 milliseconds and sometimes takes 250.
- Shared data. Two tests use the same customer account, so one test's leftovers change the other test's result.
- Outside services. A payment service in the test environment is slow or down, and the test fails for a reason outside the code.
A 2026 study found that for 58% of flaky end-to-end tests, the test code and continuous integration (CI) logs were not enough to find the cause. Finding those causes would need more evidence from running the tests.
When are end-to-end tests worth their cost?
Ham Vocke's The Practical Test Pyramid, published on martinfowler.com, advises teams to "reduce the number of end-to-end tests to a bare minimum." End-to-end checks are usually worth their cost where a broken journey would block customers or lose data:
- Critical journeys. A journey whose failure stops sales, e.g. checkout, gets one test that runs before each merge.
- Wiring after a deploy. A few journeys run as a smoke test to show that the deployed parts connect.
- Changes that cross parts. When a change touches the interface, a service, and the database at once, separate tests are least likely to catch a failure.
Other end-to-end tests can run on a schedule or before a release.
End-to-end tests check only the journeys someone scripted, so a green run shows that those journeys work, not that the whole app does. Exploratory testing covers some of the rest, because the tester designs checks while using the software.
How is end-to-end testing different from integration testing?
An integration test checks that two or more parts work together, e.g. a service and its database, usually without the user interface. An end-to-end test runs the whole system through the surface a user touches, so it also covers the wiring an integration test leaves out. Integration tests run faster and point closer to the fault, while end-to-end tests stay closer to what a customer does. The comparison of unit tests with integration and end-to-end tests covers when a team should use each level.
How does RunStory help with end-to-end testing?
A coding agent can break a journey that no end-to-end test follows. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working.
Suppose the developer had not written the checkout test first. The prompt still names the journey to try, editing the address with items in the cart, and a run of that journey shows the empty cart. This is an illustrative example, not a live run. RunStory is in private alpha for CLIs and web apps. Your team keeps the final release decision.
FAQs
How do end-to-end tests keep up with AI-generated code?
End-to-end tests keep up with AI-generated code when the suite stays small and covers the critical journeys. A person writes them from the request before the agent starts and runs them after each change.
Should end-to-end tests use a real database?
End-to-end tests usually use a real database in a test environment close to production, because the stored data is part of the journey they check. Each test loads known seed data and resets it afterward.
Can end-to-end tests run against a system written in another language?
End-to-end tests can run against a system written in another language, because they act only through the surface a user touches, e.g. a web page. The test never imports the app's code.
How many end-to-end tests does a project need?
Most projects need only a few end-to-end tests, one for each journey whose failure would stop sales or lose data, e.g. checkout, plus a few smoke tests after each deploy. Unit and integration tests check other behavior more cheaply.