What is golden file testing (golden master)?
Golden file testing is a technique that saves a program's approved output in a file and fails a test when later output differs. It is also called golden master testing or approval testing. The golden file is the saved, approved copy that each later run is compared with.
A characterization test is a test that records what existing code does now, even when that behavior is wrong, so that later changes can be checked against that record. The term was coined by Michael Feathers, whose book "Working Effectively with Legacy Code" uses it for tests that capture the behavior of legacy code before anyone changes it. Many characterization tests keep what they record in a golden file.
The golden file is the test's test oracle, the source of its expected result. That result is whatever an earlier version of the program produced. The test therefore checks that the output stayed the same, not that it is correct. The technique suits long or structured output, e.g. the text a command-line application prints. Code coverage counts the lines a test ran, and a golden file checks the output those lines produced.
How does golden file testing work?
A golden file test runs in one of two modes. In compare mode, which is the default, the test follows five steps:
- The test runs the program with a fixed input, e.g. the same command arguments on each run.
- The test captures the output, e.g. what a command-line interface (CLI) writes to stdout and stderr.
- The test reads the golden file from the repository.
- The test compares the two outputs, often after masking values that differ between runs, e.g. timestamps.
- The test passes when the outputs match. It fails and prints a diff when they differ.
In update mode, the test writes the captured output over the golden file instead of failing on a difference. A developer turns update mode on with a flag, e.g. the -update flag in the Go project's own gofmt tests. The developer then reads the rewritten file and adds it to a commit with the code. That review is the approval that makes the file golden.
Golden files sit next to the tests that read them, in version control. In Go, they usually live in a directory named testdata, which the go tool ignores when it builds packages. The testscript package for Go can keep the expected output inside each test script, as a file in the same text archive. Its cmp command compares two files, e.g. a command's standard output and an expected file stored in the script.
What is an example of golden file testing?
Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, acme, exports orders. A developer at Acme writes a golden file test for acme export --format csv.
The test loads one fixed order, A-1042, runs the command, and compares its output with testdata/export.golden. It follows the common Go pattern, with an -update flag that turns on update mode:
var update = flag.Bool("update", false, "rewrite golden files")
func TestExportCSV(t *testing.T) {
got := runAcme(t, "export", "--format", "csv")
golden := filepath.Join("testdata", "export.golden")
if *update {
os.WriteFile(golden, got, 0o644)
}
want, _ := os.ReadFile(golden)
if !bytes.Equal(got, want) {
t.Errorf("export differs from %s:\n%s", golden, lineDiff(want, got))
}
}
The developer runs go test -update once to create the file, reads the CSV, and commits it.
Later, a developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent moves the address to a separate field on the order. The checkout tests pass, but the golden file test fails with this diff:
order_id,items,total,delivery_address
-A-1042,2,$240.00,12 Elm Street
+A-1042,2,$240.00,
The export still reads the old address field, so the column is empty. No checkout test covers the export, so this regression shows up only because the golden file compares the whole output. The right response is to fix the export code and leave the golden file as it is.
This example is simplified. A real export test would cover more than one order, e.g. an order with no delivery address.
What changes when a coding agent writes the code?
A coding agent can make a failing golden file test pass with one command, by running update mode. In the Acme example, go test -update would save the empty address column as the expected output, and no line of test code would change. This form of test tampering is hard to spot, because the change sits in a data file. A large golden file diff can pass code review unread.
Many test tools include this rewrite as an option. In testscript, the UpdateScripts option makes a failing cmp command succeed and rewrites the expected output inside the script.
Golden files also help an agent. A characterization test written before the agent edits old code records what that code does. An unintended change to the output then fails with a diff the agent can read and fix.
A team can keep that benefit and limit the risk. It can tell the agent to report a golden file mismatch instead of running update mode. A continuous integration run still passes with a golden file regenerated before the commit. So the team can also require a person to approve each golden file change, e.g. with a code owners rule on testdata.
What are the limits of golden file testing?
Golden file tests have four limits:
- Recorded output is not correct output. A golden file holds what the code produced when someone approved it. If that output was already wrong, the test protects the bug.
- Values that change between runs cause false failures. Output with a timestamp or a random ID differs on each run. The test has to control those values, e.g. with a fixed clock, or replace them with a placeholder before it writes or compares the output. Otherwise it fails when nothing is broken.
- Intended changes fail many tests at once. A golden file records every detail of the output, e.g. the order of the columns, so one intended format change fails each test that uses that format. A team that regenerates the files after each such change ends up with tests that record whatever the code does.
- The test covers only the inputs it records. A golden file checks the fixed inputs that someone chose. No failures is not the same as complete coverage of the inputs that users send.
How is a golden file different from a snapshot?
A snapshot test uses the same record and compare method, managed by a test framework. The framework names and stores each snapshot, e.g. in a __snapshots__ folder next to a Jest test, and writes a missing one on the first run. A golden file is usually a plain file that the team names and reads directly, often the exact output a user sees, e.g. a CSV export.
The terms overlap. "Golden file" is common in Go projects, and "snapshot" is common in JavaScript. Flutter's golden files are images of rendered widgets, and flutter test --update-goldens rewrites them.
How do you keep what passed before true?
What passed before needs to remain true, unless someone explicitly decides otherwise. A golden file update is that decision, so a person should make it and write down why the output changed.
A golden file holds that rule only for the output a test captured, not for the workflows people run in the app. When a workflow breaks, the useful check is to run that workflow again after the fix. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps.
FAQs
What is a golden file in a test?
A golden file in a test is a saved, approved copy of a program's output. The test runs the program again, compares the output with the golden file, and fails with a diff when any line differs.
When should you use golden tests?
Golden tests fit output that is long or structured, e.g. the text a command-line tool prints, where one file can hold the whole result. They also fit legacy code, where a characterization test records what the code does before anyone changes it. They fit poorly when intended changes to the output are frequent.
Where should a golden file be located?
A golden file should sit next to the test that reads it, in version control with the code, so each change to it shows up in review. Go projects usually keep golden files in a directory named testdata, which the go tool ignores when it builds packages.
What is a characterization test?
A characterization test is a test that records what existing code does, even when that behavior is wrong, so that later changes can be checked against that record. Teams write one before they change legacy code, and many keep its recorded output in a golden file.