Why should tests run on a clean checkout?
Tests should run on a clean checkout because a pass there describes the committed code alone, the code that other people and machines receive. A clean checkout is a folder that holds exactly the files of one commit. It has no uncommitted edits, untracked files, or ignored files, e.g. build output.
The working folder of a developer or a coding agent usually holds all three. A test can pass there because of a file that the commit lacks, then fail on a continuous integration (CI) runner. A fresh clone of the branch into an empty folder gives a clean checkout. A fresh sandbox around it also removes state outside the folder, e.g. global packages.
A clean build starts with no output or cache from earlier builds. The Reproducible Builds project calls a build reproducible when the same source code, build environment, and build instructions give identical artifacts, bit for bit. A clean checkout controls the first of those inputs. Test isolation applies the same idea to each test inside a run.
What causes failures that only show on a clean run?
A test that passes locally and fails on a clean run depends on a layer of the working folder outside the commit.
Is a file missing from the commit?
A file that nobody ever added stays untracked, so no clean checkout contains it. A .gitignore rule can also hide a needed file, e.g. a build/ rule that also matches a source folder at src/build. The import that needs the file then fails.
Did an edit stay out of the commit?
Tests run against the folder, not the commit. When someone commits some changed files and leaves another unstaged, the local run tests code that no commit holds.
Does the code read an ignored file?
Ignored files usually stay out of the commit. Code that reads settings from a local .env file passes locally and fails where the file is absent. A local database file, e.g. dev.sqlite3, can also hold rows that the tests rely on.
Is a package installed but not declared?
A package installed by hand, e.g. with pip install, stays in the environment with no entry in pyproject.toml. A clean install from the declared dependencies leaves it out, and the import fails.
Is old build output still on disk?
The TypeScript compiler does not delete a compiled file when its source file is removed. Code that loads from dist/ can then still find a renamed module under its old name. When dist/ is ignored, a clean checkout has no dist/ folder, so the old name fails there.
What does a failure on a clean checkout look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The work runs in this order:
- The agent edits
src/checkout/address.js, and its test fails because saving the address empties the cart. - The agent fixes
src/cart/cart.js, andnpm testreports 12 passed tests in its working folder. - The agent stages the address code and its test by name and commits them, so the cart fix stays uncommitted.
- The agent's output says "done," a false completion claim.
The developer clones the branch into an empty folder, installs from the lockfile, and runs the suite:
tmp=$(mktemp -d)
git clone --quiet --branch edit-address ~/src/acme-store "$tmp/acme"
cd "$tmp/acme" && npm ci && npm test
The agent's test fails on the clean checkout:
FAIL test/checkout/address.test.js
● changing the address keeps the cart
expect(received).toHaveLength(expected)
Expected length: 2
Received length: 0
Received array: []
The address saved, but the cart emptied. In the agent's folder, git status still lists src/cart/cart.js as modified and not staged. The agent commits the cart fix, and the clean run passes.
This example is simplified. A real project would also run this check on CI's operating system, e.g. Linux.
What changes when a coding agent writes the code?
A coding agent works in one folder for the whole task, and each attempt can leave packages or files behind. The agent's tests run against that folder, so its "done" describes the folder, not the commit.
Some agents stage files by name, as the Acme agent did, and a file left off the list stays out of the commit. A cloud agent often starts from a fresh clone, but its folder gathers the same leftovers. A clean install also fails on one kind of hallucination, a declared package name that nobody has registered.
A practical adjustment is to accept "done" only after two checks. First, git status --porcelain prints nothing, so no edit or file outside the ignore rules is left uncommitted. Second, the suite passes on a fresh clone of the branch, on a machine that starts empty, e.g. a microVM created for each run. The steps to verify a CLI start the same way.
How can teams stop depending on local state?
Each practice below removes one source of local state or makes it visible:
- Run from a fresh copy. The checks run in a fresh clone of the branch, or in a separate checkout made with
git worktree add, as the steps to run CI checks locally show. - Check how a local runner copies code. By default, act copies the working folder into the job in place of
actions/checkout, leaving out paths that.gitignorelists. Uncommitted edits and untracked files join the run unless act runs with--no-skip-checkout. - Install only from the lockfile. The
npm cicommand deletes an existingnode_modulesfolder and exits with an error whenpackage.jsonand the lockfile disagree. A cache of package downloads works with that install, but a restorednode_modulesor build folder brings the leftovers back. - Keep the clean step. On a runner that keeps its workspace between jobs, the checkout action for GitHub Actions runs
git clean -ffdxandgit reset --hard HEADunless a workflow setsclean: false. - Start the app, not only the tests. A smoke test on the clean checkout shows whether the app starts without the developer's local settings.
- Start the machine clean too. A fresh virtual machine for each run, or an ephemeral environment for the whole running app, removes state outside the folder.
A clean checkout removes leftovers, not differences between machines, e.g. in the time zone. Those gaps in environment parity are why tests that pass locally can fail in CI. A clean run also costs install time and adds no checks, so a suite that tests the wrong thing passes there too.
How is a clean checkout different from a fresh clone?
A fresh clone is one way to get a clean checkout. It copies a repository into an empty folder and checks out one branch, so the folder holds only committed files. A clone from the remote holds only pushed commits, while a clone from a local path also holds unpushed ones.
A clean checkout can also come from an existing folder. The git reset --hard command undoes edits to tracked files, which git clean leaves alone. Then git clean -fdx deletes untracked and ignored files. The Git documentation suggests -x with a reset "to create a pristine working directory to test a clean build."
Run git clean -ndx first to list what git clean -fdx would delete, because Git cannot restore those files.
In GitHub Actions, the checkout action fetches one commit by default, unlike a fresh clone, so a job whose tests read Git tags needs fetch-depth: 0.
How does RunStory help with testing from a clean state?
A test run on the machine where the code was written can pass because of packages or settings that only that machine has. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. You keep working.
RunStory is in private alpha for CLIs and web apps. The alpha tests your software in isolated sandboxes. Your team keeps the final release decision.
FAQs
Does git clean give you a clean checkout?
The git clean command gives a clean checkout only together with a reset. With the -x option, it deletes untracked and ignored files but does not touch edits to tracked files, which a hard reset undoes. Neither command removes packages installed outside the folder.
Should a run from a clean checkout use cached dependencies?
A run from a clean checkout can use cached package downloads, because the install still follows the committed lockfile. A restored dependency folder or build output can bring back the leftovers that the clean checkout removed.
Does a local CI runner start from a clean checkout?
A local CI runner does not always start from a clean checkout. By default, act copies the working folder into the job in place of the checkout step, so uncommitted edits and untracked files join the run. Running act from a fresh clone keeps them out.
How do uncommitted files break CI?
Uncommitted files break CI when the code or its tests depend on them. The developer's run reads the working folder, where the files exist, so it passes. The CI runner checks out the commit, where they are missing, so an import or a file read fails.