Skip to content

What is test isolation (hermetic tests)?

Test isolation means each test sets up its own state and data, so its result does not depend on other tests, the order they run in, or leftover files.

Last updated , 8 min read

What is test isolation (hermetic tests)?

Test isolation is a property of a test that sets up its own state, so other tests and leftover data cannot change its result. The strictest form is a hermetic test, which also avoids outside services and anything it did not declare. Its result then does not change with what else is installed on the machine.

Shared state is anything that outlives one test and that another test can read or change, e.g. a row in a database. A whole test run is isolated when it starts on a machine with nothing left from earlier runs, e.g. a fresh sandbox. Isolation also has a second meaning in unit testing, where a test double replaces a dependency so that a test checks one piece of code alone.

Broken isolation is a common cause of flaky tests, which pass and fail on the same code. The first large study of flaky tests, by Luo and colleagues (2014, 51 projects), found the most common causes were asynchronous waits, concurrency, and dependence on test order.

How does test isolation work?

An isolated test controls its own state from start to finish:

  1. A test fixture creates the state the test reads, e.g. a customer with an empty cart.
  2. The test runs the code and checks the result against that state.
  3. The fixture removes what it created, or the test runner discards it, e.g. by rolling back a database transaction.

In an essay on nondeterministic tests, Martin Fowler prefers rebuilding each test's starting state over cleaning up after each test. When a test cleans up badly, a later test fails, and the cause is hard to find. Broken isolation usually shows up as test-order dependence, a flaw that makes a test pass or fail depending on which tests ran before it. A test can also fail when run alone, because it needs data that an earlier test created.

Shared state compared with isolated tests Shared state Test A leaves 1 item Shared cart finds 3, not 2 Test B fails Isolated Test A Own cart A Test B Own cart B Each test creates its cart, then removes it
With shared state, the test that fails is not the test that left the extra item. With isolation, each test reads only data it created.

Isolation works at three levels:

  • Single test. Each test gets its own data, files, and browser state. In pytest, the tmp_path fixture gives each test its own temporary directory. In browser tests, Playwright Test creates a separate browser context for each test, with its own cookies and local storage.
  • Whole run. Each run starts from a clean checkout and an empty database, with no leftover processes. Hosted continuous integration and delivery (CI/CD) runners often start each job on a fresh virtual machine. Loading known data into that database is part of test data management.
  • Parallel runs. Runs at the same time need separate copies of anything they both write, e.g. a test database.

What is an example of a test isolation failure?

Here is an illustrative example. Acme Co. sells furniture online, and its checkout tests share one test account, shopper@example.com. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The work happens in this order:

  1. The agent writes test_edit_address, which adds 1 lamp at $90.00 to the shared account's cart and changes the address to 12 Elm Street.
  2. The agent's test does not empty the cart when it finishes.
  3. The agent runs pytest tests/test_address.py, which reports 1 passed test, and the agent's output says "done."
  4. In the full suite, that file runs before tests/test_cart.py, which uses the same account.
  5. The older test_cart_total adds 2 chairs at $120.00 each and expects a total of $240.00, but the cart also holds the lamp.

The run ends with this failure:

E       AssertionError: assert Decimal('330.00') == Decimal('240.00')
FAILED tests/test_cart.py::test_cart_total

Run alone, test_cart_total passes. The failing test is correct, and the agent's test caused the failure. Luo and colleagues call a test that changes shared state and leaves it changed a "polluter." The fix gives each test its own customer through a fixture that deletes the customer afterwards:

from decimal import Decimal
from uuid import uuid4

import pytest

@pytest.fixture
def customer(db):
    c = db.create_customer(email=f"{uuid4()}@example.com")
    yield c
    db.delete_customer(c.id)

def test_cart_total(customer):
    customer.cart.add("chair", quantity=2)
    assert customer.cart.total() == Decimal("240.00")

With the fixture, both tests pass in either order. This example is simplified. A real suite would also isolate orders and addresses, e.g. inside a database transaction that rolls back.

What changes when a coding agent writes the code?

A coding agent may run only the tests closest to its change, as the Acme agent did. Those tests can pass while the agent's test breaks an older one in the full suite. The agent's "done" can then rest on a run that never tried the order that fails.

Test order can also change between runs, e.g. when tests are split across parallel workers. That nondeterminism lets a rerun pass with no fix. A runner that shuffles tests usually prints its seed, so the agent can replay the failing order instead of rerunning until the suite passes.

Parallel agents on one machine often share one test database, so one agent's test can delete rows that the other agent's test created. A second development server cannot bind a port that the first one holds, so it fails to start or moves to a port the tests do not expect.

A practical adjustment is to give each agent its own Git worktree, which separates the working files. A worktree does not separate a database or a port, so each agent also needs its own test database and port. Then run the whole suite, not only the nearest tests, before accepting "done."

What are the limits of test isolation?

Isolation has costs and blind spots:

  • It costs time. A unit test that works on objects in memory is cheap to isolate. An end-to-end test drives the whole running app, so a fresh database and server for each test is slow.
  • It lowers fidelity. Fidelity is how closely a test matches production, and a more hermetic test usually has less of it. An integration test that swaps a real service for a fake is more isolated and checks less of the real connection.
  • It removes only one cause of flakiness. An isolated test can still fail on timing, e.g. a wait that is too short for a slow page.
  • It adds no checks. An isolated test that checks the wrong thing passes on broken code as reliably as any other test.

A staging environment runs the real services together, so it can catch problems that isolated tests leave out.

How is test isolation different from a sandbox?

A sandbox is an isolated environment that limits which files and network hosts the code inside it can reach. Test isolation keeps each test away from the state of other tests. The two overlap at the run level, because a fresh sandbox for each run gives that run a clean start. They differ inside the run. Tests in one sandbox can still share a database and depend on each other's order, so a sandbox alone does not make a suite isolated.

How do you run each check from a clean state?

Start each run from a clean checkout and an empty test database, and let each test create the data it reads. Also run the suite in a random order, e.g. with the pytest-randomly plugin, which shuffles tests on each run, so hidden order dependence fails where someone can see it.

A check that runs on a developer's own machine shares its ports and processes with other programs on that machine. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. You keep working. It is in private alpha for CLIs and web apps, and the alpha tests your software in isolated sandboxes.

Join the RunStory alpha →

FAQs

What is a hermetic test?

A hermetic test is a test that depends only on what it declares and sets up itself. It uses no shared state, no outside services, and no files left by earlier runs. Its result does not depend on what else is on the machine or on test order.

How do you stop tests from sharing state?

Tests stop sharing state when each test creates the data and files it reads, under names that no other test uses. Rolling back a database transaction and using a fresh temporary directory keep one test's changes away from other tests. Running the suite in a random order on each run can expose dependence that a fixed order hides.

Do integration tests need isolation?

Integration tests need isolation, and they are harder to isolate than unit tests because they often share a real database or service. Each integration test can create its own records inside a transaction that rolls back. A fully hermetic integration test replaces outside services with fakes, so it is more isolated but checks less of the real connection.

How do parallel agents collide in tests?

Parallel agents collide in tests when their runs share one checkout, one test database, or one port. A test can then fail because of the other agent's run, not because of either change. A separate worktree, database, and port for each agent keeps their runs apart.