Why do tests pass locally but fail in CI?
Tests pass locally but fail in continuous integration (CI) when an input that the tests never set has a different value on the CI machine. Common differences are timing, test order, tool versions, machine settings, and state left on the laptop. The fix is to pin that input or make the test stop depending on it.
A CI/CD pipeline usually runs each job on a fresh machine, so these differences show up there. Each is a gap in environment parity, the degree to which two environments match.
Many such failures are flaky tests whose hidden input changes more in CI. The first large study of flaky tests (2014, 51 projects) found the most common causes were asynchronous waits, concurrency, and dependence on test order.
What causes tests to pass locally but fail in CI?
Is the CI machine slower or busier?
A CI runner often has fewer CPU cores than a laptop, and the whole suite competes for them. A test that waits a fixed time, e.g. 200 milliseconds, can then check before the page is ready. A race condition can stay hidden on a fast laptop and appear under load.
Do the tests run in parallel or in another order?
A developer often runs one test file locally, while CI runs the whole suite, sometimes in parallel worker processes. Tests that share one test account or database then collide, or pass only when another test runs first. Test isolation removes this cause.
Do the versions or the operating system differ?
A laptop can run different tool versions from CI, e.g. an older Node.js. GitHub also updates the tools in its hosted runner images, typically once a week. A Mac usually ignores the case of file names and Linux does not, so an import of ./Header finds header.js only on the Mac. Playwright names the screenshot baseline of each snapshot test after the browser and platform, e.g. chromium-darwin, so a Linux runner looks for a file that the Mac never wrote.
Does the code read settings from the machine?
Code that formats a time without naming a zone reads the machine's time zone. A CI runner often uses Coordinated Universal Time (UTC), while a laptop uses the developer's zone. GitHub Actions sets CI to true, so Jest fails a test with a missing snapshot and prints New snapshot was not written. CI steps also write to a pipe, so tools that check for a TTY drop their colors and prompts, and a test that expects them fails.
Random test data can reach inputs that fixed values miss. A CI failure from it is hard to replay unless the test fixes or prints its seed, or prints the failing value.
Does the laptop hold state that CI lacks?
A test can pass because of state the laptop kept, e.g. rows in a local database. A hosted runner starts from a fresh checkout, so an untracked file, e.g. .env, is missing there. CI usually runs npm ci, which installs the committed lockfile's exact versions, while a local npm install can drift from them. A runner that a team hosts itself can keep files from earlier jobs.
What does a CI-only failure look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds the time of the last address change to the checkout summary, with a Jest snapshot test that fixes the clock:
test("summary shows the saved address", () => {
jest.useFakeTimers();
jest.setSystemTime(new Date("2026-10-05T23:30:00Z"));
const html = renderSummary({ address: "12 Elm Street", items: 2 });
expect(html).toMatchSnapshot();
});
The work runs in this order:
- The agent runs
npx jeston the laptop, which uses Los Angeles time. Jest records "Address updated at 4:30 PM" in a snapshot file, and the test passes. - The agent commits the code, the test, and the snapshot. Its output says "done."
- CI runs the same commit on a runner that uses UTC. There the same moment formats as 11:30 PM, and the test fails.
- The agent runs
npx jest -uon the laptop. It records 4:30 PM again, and the next CI run fails the same way.
The CI log shows the difference:
● summary shows the saved address
expect(received).toMatchSnapshot()
Snapshot name: `summary shows the saved address 1`
- Snapshot - 1
+ Received + 1
- <p>Address updated at 4:30 PM</p>
+ <p>Address updated at 11:30 PM</p>
The developer pins the time zone in the test script, which the agent and CI both run with npm test:
{ "scripts": { "test": "TZ=UTC jest" } }
On Windows, set process.env.TZ = "UTC" in a Jest globalSetup file instead.
After a reviewed snapshot update, both machines record 11:30 PM and pass. This example is simplified. A real checkout would also test how customers in other time zones read the time.
What changes when a coding agent writes the code?
A coding agent runs tests where it works, on a laptop or in a cloud sandbox. Its "done" rests on that machine's settings and state, not the CI runner's. A 2026 study of four database systems found that tests generated by a large language model (LLM) were slightly more likely to be flaky than existing tests. The flaky tests often relied on an ordering the system did not guarantee, e.g. rows from an unsorted query.
When CI fails, an agent told to make the check pass can edit the test instead of removing the difference. A skip when process.env.CI is set turns CI green and hides the failure.
A team can have the agent run the check in a container that matches CI before its output says "done." Review test code that reads process.env.CI, and give the agent the failed CI log.
How can teams remove the difference for good?
Each fix below removes one difference or makes it visible:
- Use one script. The workflow and the developer call one script, as the steps to run CI checks locally show.
- Pin the machine. Local checks and CI jobs use one container image, e.g. the image that a dev container names.
- Pin the settings. The test script pins the time zone. Martin Fowler's essay advises wrapping the system clock so a test can replace it.
- Wait for the result. A test that waits until the page shows the result passes on a slow runner.
- Isolate test data. A test that creates its own data does not depend on test order or worker count.
- Reproduce CI's settings. A local run that changes one setting at a time, e.g.
CI=true, shows which one makes the test fail. - Keep the evidence. The run's page in the CI service holds each step's log. In GitHub Actions,
gh run view RUN_ID --log-failedprints only the failed steps, and an upload step keeps test reports and traces.
How is a CI-only failure different from a flaky test?
A CI-only failure from a fixed difference fails in CI and passes locally on every run, e.g. one caused by the time zone. A flaky test changes its result on one machine with no change to the code. The two overlap when a timing or order flake fails rarely on a laptop and often under CI load. Reruns with retries off on both machines tell them apart, the same way reruns separate a flaky test from a real bug.
How do you run a change somewhere other than your machine?
Let CI run the change on a branch, or run CI's script from a clean copy in a container that matches CI.
A clean run elsewhere exposes environment differences, but only for behavior that a test already checks. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps.
FAQs
How do you get the logs from a failed CI run?
The logs from a failed CI run are on the run's page in the CI service, one log for each step. In GitHub Actions, GitHub's command-line tool prints only the failed steps' logs. Test reports and traces stay available only when a workflow step uploads them.
Why do end-to-end tests fail only in parallel in CI?
End-to-end tests fail only in parallel in CI when parallel workers write to the same data, e.g. one test account. A developer often runs one test file at a time on a laptop, so the tests never collide there. Giving each test its own data removes the collision.
Why does a test pass on one CI runner and fail on another?
A test passes on one CI runner and fails on another when the runners differ, e.g. in their image version. Hosted runner images get regular tool updates, and a runner that a team hosts itself can keep files from earlier jobs. Running the job in a pinned container image removes many of these differences.
Are random values in unit tests a good idea?
Random values in unit tests are a good idea when the test fixes or prints its seed, because they can reach inputs that fixed values miss. Without the seed or the failing value, a failure in CI is hard to replay on the laptop.