What is test impact analysis?
Test impact analysis is a technique that uses a map from code to tests to run only the tests that exercise the changed code. It is a form of regression test selection, also called test selection. The other tests are skipped, which saves time when the whole suite is too slow to run on each change.
In Paul Hammant's 2017 article, a tool builds the map by running each test on its own with code coverage switched on. The map lives in text files in the code's repository.
Teams usually run it in continuous integration and delivery (CI/CD) pipelines on each pull request. A test that the map does not link to the change does not run, even when the change breaks it.
In software testing, impact analysis usually means change impact analysis, a wider practice. It lists the parts of a system that a change can affect, often by hand, so a team can plan what to retest.
How does test impact analysis work?
A selection step runs before the test runner starts:
- A tool builds the map from earlier runs or from the code's dependencies.
- A developer or a coding agent pushes a change, and the CI job lists the changed files, e.g. with
git diff --name-only origin/main...HEAD. - The selector looks up each changed file in the map and collects the tests linked to it.
- The test runner runs the selected tests and skips the rest.
- For a map built from runs, the results update it.
A selector links code to tests in four common ways:
- Import graph. The tool follows the code's imports and selects the tests that reach a changed file, e.g. Jest's
--changedSinceflag. - Coverage map. Each test runs with coverage recording on, and the tool saves the lines it executed, e.g. coverage.py with
dynamic_context = test_function. - Build graph. A build tool reruns the test targets whose declared dependencies changed, e.g. Bazel.
- Predictive model. A model trained on past changes and test results estimates each test's chance of failing, and the selector runs the likely failures. This variant is called predictive test selection.
A map built from runs describes the code at its last recording, so it goes stale as the code changes. A test added since then has no entry, so a selector must run it until the map records it. When code moves to another file, the map still links its tests to the old file. Hammant's article says the map has to be updated regularly.
What is an example of test impact analysis?
Here is an illustrative example. Acme Co. sells furniture online. Its web store has 340 Jest tests, and some are end-to-end tests that drive the store in a browser. On each pull request, CI runs only the tests that Jest selects:
npx jest --changedSince=origin/main
The change then moves through these steps:
- A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout."
- The agent edits one file,
src/checkout/address.js. The edited code starts a fresh checkout session after it saves an address. - Jest finds the test files that import that file and selects
src/checkout/address.test.js, which holds 2 of the 340 tests. The end-to-end checkout test loads pages over HTTP and imports no source file, so Jest skips it. - The 2 selected tests pass in under a second, and the agent's output says the change is "done."
- A reviewer merges the pull request. That night the whole suite runs, and the checkout test fails. The address saved, but the cart emptied.
The selection followed its rules, and the failure appeared only after the merge. This example is simplified. A real project would also map its end-to-end tests, e.g. with coverage collected from the running app during each test.
What changes when a coding agent writes the code?
Coding agents push more changes, so the data a selector reads has to update faster. Anthropic reported that its CI job volume grew 25-fold in six months as Claude came to write about 80% of its code. That growth forced a redesign of how it chooses which tests to run.
Anthropic's selector reads past test results and picks the tests that run on each pull request. When the service that records those results fell behind, the selector worked from stale data. A test that had been added or fixed would not run until the records caught up.
An agent can also run a subset itself, e.g. with npx jest --onlyChanged, and then report that the tests pass. That report covers only the tests Jest picked. If the agent has already committed its change, that command finds no tests and still exits without an error.
A team can ask the agent to report the command it ran and keep one full run between the agent's change and a release. Selection keeps CI for coding agents fast, and the full run catches what it skips.
What can test selection miss?
Test selection makes software testing faster, not wider. It picks from existing tests, so it can miss a failure in these cases:
- Changes outside the tracked code. A coverage map or an import graph records code files. A change to a file that the map does not track, e.g. a configuration file, links to no tests.
- Dynamic links. Code that loads a module by a name built at runtime hides the link from an import graph. Jest's options for changed files need a static dependency graph for this reason.
- Breaks outside the diff. A change can break a feature through shared state, e.g. a session that the cart also reads. If no linked test checks that feature, nothing runs for it, which is one way coding agents break working features.
- Flaky results. Random failures from a flaky test add noise to the past results that some selectors rank tests by, so a useful test can be skipped.
- Weak tests. A selected test that checks too little still passes on broken code. Mutation testing measures how many injected bugs a suite detects.
A run with gaps does not establish that the whole app works.
How is test impact analysis different from running the whole suite?
Running the whole suite on each change is called retest all in regression testing. It catches any failure the suite can detect, but its run time grows with the suite. Test impact analysis runs a subset, so results arrive sooner and a missed failure shows up later.
Test prioritization also runs the whole suite, but it orders the tests so the ones most likely to fail run first. Test sharding splits the whole suite across machines that run at the same time, so each test still runs. Teams often combine sharding with selection, and a pre-push hook can run the selected tests before CI runs more.
What should run beyond the selected tests?
On each change, also run a short smoke test and the tests that failed on the last run. Run the whole suite on a schedule, e.g. nightly, and before each release. Use those full runs to refresh the map and find failures that selection skipped. Treat a change to configuration or dependencies as a reason to run the whole suite too.
Selection can pick only from tests that someone already wrote. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. It is in private alpha for CLIs and web apps.
FAQs
What is predictive test selection?
Predictive test selection is a form of test impact analysis that uses a trained model instead of a fixed map. From past changes and test results, the model estimates how likely each test is to fail, and the selector runs the likeliest failures. It can still skip a test that would have failed.
How is test impact analysis different from test sharding?
Test impact analysis cuts the number of tests in a run. Test sharding cuts the time a full run takes by splitting it across machines that run at the same time. Sharding skips no test, so it misses nothing the suite would catch, and teams often combine the two.
How often should the full suite still run?
The full suite should still run on a schedule, e.g. nightly, and before each release. Full runs catch the failures that selection skipped and refresh the map. A change to configuration or dependencies is also a reason to run the full suite, because a map of code files may not link it to any test.
Does test impact analysis work for end-to-end tests?
Test impact analysis works less well for end-to-end tests, because they usually drive the running app over HTTP or in a browser instead of importing its source files. An import graph then links them to nothing. A coverage map for them needs coverage collected from the running app during each test.