What is snapshot testing?
Snapshot testing is a technique in which a test framework records a program's output and fails the test when later output does not match. The recorded output is the snapshot. When the test fails, a person reads the difference and either fixes the code or accepts the new output.
A snapshot can hold any output that a test turns into text or an image, e.g. the HTML of a rendered component. In Jest, snapshot files end in .snap and sit in a __snapshots__ folder next to the test. Snapshot testing with screenshots is called visual regression testing.
Snapshot testing is a form of software testing in which the expected result comes from the program, not from a person. A snapshot check can sit in a unit test or an integration test. Snapshots make test automation cheap to write, because nobody types the expected output by hand. For the same reason, a snapshot test checks that the output stayed the same, not that it is correct.
How does snapshot testing work?
A snapshot test goes through five steps:
- The test runs the code with a fixed input and captures the output.
- The test framework turns the output into readable text, or saves an image for a screenshot.
- On the first run, the framework saves the output as the snapshot.
- On later runs, the framework compares the output with the snapshot. A match passes, and a difference fails with a diff.
- A person reads the diff. A bug gets a code fix, and the snapshot stays. For an intended change, the person runs the tests with an update flag, e.g.
jest --updateSnapshot, and commits the result.
Jest writes a missing snapshot and passes, except in continuous integration (CI), where it fails the test unless the run includes the update flag. Playwright writes a missing snapshot but fails the test by default.
For a command-line application, a useful snapshot holds the exit code and what the program writes to stdout and stderr, as in this Jest test:
const { spawnSync } = require("node:child_process");
test("acme export --format csv", () => {
const run = spawnSync("acme", ["export", "--format", "csv"], {
encoding: "utf8",
});
expect(run.error).toBeUndefined();
expect({
status: run.status,
stdout: run.stdout,
stderr: run.stderr,
}).toMatchSnapshot();
});
Such tests often replace machine details, e.g. a temporary folder path, with a placeholder before they compare.
A terminal UI redraws its screen with escape sequences, so its tests usually record the rendered screen instead.
What is an example of a snapshot test?
Here is an illustrative example. Acme Co. sells furniture online. A Jest snapshot test covers the checkout summary, and its snapshot was recorded before the change below:
test("checkout summary with 2 items", async () => {
const order = await startCheckout({ items: 2 });
await order.setAddress("12 Elm Street");
expect(renderSummary(order)).toMatchSnapshot();
});
The snapshot holds the order ID, Items: 2, Total: $240.00, and the delivery address. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes how checkout saves the address and adds an "Edit address" link to the summary. The test now fails:
● checkout summary with 2 items
- Snapshot - 2
+ Received + 3
Order A-1042
- Items: 2
+ Items: 0
- Total: $240.00
+ Total: $0.00
Deliver to: 12 Elm Street
+ Edit address
The developer reads the diff. The last line is the link that the request asked for. The item and total lines are a bug. The address saved, but the cart emptied.
The agent fixes the checkout code at the developer's request. Only the "Edit address" line still differs, so the developer accepts it with npm test -- -u and commits the updated snapshot. This example is simplified. A real checkout would record more states, e.g. a cart with a discount code.
What changes when a coding agent writes the code?
A snapshot failure has a fix that changes no code. When the suite runs with npm test, Jest's failure summary prints the update command, npm test -- -u, and Playwright's equivalent is --update-snapshots. A coding agent told to make the tests pass can run that command. The suite then passes, and the only trace is an edited snapshot file. This is a form of test tampering that a reviewer can miss.
On a run of the whole suite, the flag accepts every failing snapshot at once. In the Acme example, running it before the fix would have saved the link and the empty cart together. Jest's documentation says to fix any bug behind a failing snapshot before regenerating it.
A snapshot test that the agent adds has a second problem. Its first run records whatever the code produces, bug included, so later runs pass against the bug.
A practical adjustment is to have the agent report snapshot failures instead of updating them. A person then updates the snapshots one test at a time, e.g. with Jest's --testNamePattern flag, and approves each change in code review. Snapshots that the agent added get the same review.
Why are snapshot tests fragile?
Snapshot tests are reliable at detecting that output changed. They help most with stable output that is slow to write out by hand, e.g. the output of a command-line tool. They are fragile because they check each detail of the output, including details that no user would notice. Four problems follow from that:
- Harmless changes fail the test. A renamed CSS class fails the test the same way a bug does. When most failures are harmless, people start to accept updates unread.
- Values that change between runs fail the test. A timestamp in the output changes on each run, although nothing is broken. Jest can replace such a field with a type check,
expect.any(Date), before it compares. - Machines render output differently. Fonts and rendering can differ between a laptop and a CI runner, so a screenshot test can pass on one and fail in CI.
- Large snapshots go unread. A reviewer skims a diff that runs to hundreds of lines. A single value is clearer as a plain check that names the expected value.
What is visual regression testing?
Visual regression testing is snapshot testing in which the snapshot is a screenshot of a page or a component. The test compares the current screenshot with the stored one pixel by pixel. It catches changes that a text snapshot misses, e.g. a button pushed off the screen.
In Playwright, expect(page).toHaveScreenshot() runs this check. The name of each stored screenshot includes the browser and platform, e.g. chromium-darwin, because rendering differs between them. Playwright's documentation advises running the tests in the environment that produced the stored screenshots. The maxDiffPixels option lets a few pixels differ.
Playwright's aria snapshots compare the page's accessibility tree as text, with toMatchAriaSnapshot(). They ignore how pixels render but still fail when a heading that the snapshot lists disappears.
How is a snapshot test different from a golden file?
A snapshot test and a golden file use the same record and compare method, and both are forms of regression testing. The difference is who manages the stored output.
A snapshot framework names each snapshot after its test and stores a formatted copy of the value, e.g. an object with its keys sorted. Jest can also keep that copy in the test file with toMatchInlineSnapshot(). A golden file is usually a plain file that the team names, holding the exact bytes a user receives, e.g. a CSV export.
How do you check a snapshot update against the running app?
For each changed line of a snapshot diff, run the app, repeat the workflow behind that line, and compare the result with the request that caused the change. A diff shows what changed, not whether a customer can still finish checkout.
What passed before needs to remain true, unless someone explicitly decides otherwise. A snapshot holds one output of a workflow, so checking a fix also takes a run of the workflow itself, e.g. the Acme checkout. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps.
FAQs
When should you use snapshot testing?
Snapshot testing suits stable output that is slow to write out by hand but short enough to review, e.g. the HTML of a rendered component. Output that changes often on purpose is a poor fit, because each change fails the test and invites updates that nobody reads. A single value is clearer as a plain check.
Should snapshots be committed or generated in CI?
Snapshots should be committed with the code and tests they cover, so each change to them shows up in review. Jest does not write missing snapshots in CI, and Playwright fails a test that creates a missing snapshot by default. A snapshot generated in CI would pass whatever the code produced.
How do you review a snapshot diff?
A snapshot diff review compares each changed line with the request that caused the change. A line that the request asked for is intended, and any other changed line is either a harmless change or a bug. Fix the bugs first, then update the snapshot once the diff holds only explained changes.
Do integration tests replace snapshot tests?
Integration tests do not replace snapshot tests, because an integration test checks only the results its assertions name. A snapshot test checks a whole output, so it also fails on a change that nobody wrote an assertion for, e.g. a wrong item count on a checkout summary.