What is a race condition, and how do you reproduce one?
A race condition is a bug that makes a result depend on the order in which concurrent operations run. Small differences in timing change that order, so the bug appears in some runs and not others. To reproduce one, run the operations together many times, or pause them in a test so the bad order happens every time.
The operations can be threads in one program, asynchronous tasks in one process, or separate requests to one server. In Node.js, two async functions that await network calls can interleave on a single thread. The same code and input can then give different results, a form of nondeterminism.
Races are hard to debug because the bad order needs two operations to overlap inside a short window. Load and thread scheduling decide whether they overlap, and the test controls neither. A race often shows up as a flaky test. The first large study of flaky tests (2014, 51 projects) found the most common causes were asynchronous waits, concurrency, and dependence on test order.
What causes a race condition?
A race needs operations that overlap in time, state that they share, and at least one write to that state. When nothing forces a safe order, timing decides the order. Races usually take one of three shapes:
- Lost update. Two operations read the same value, and each writes a value based on it. The later write erases the earlier one.
- Check then act. One operation checks a condition, e.g. that the last chair is in stock, and another changes it before the first acts. Security writers call this time of check to time of use (TOCTOU).
- Order violation. Code depends on one event coming first, e.g. data loading before a click handler runs, and nothing enforces that order.
A data race is a related bug. Two threads access the same memory at the same time without a lock or atomic operation, and at least one access is a write. A data race and a race condition overlap, but neither contains the other. Two database requests can race with no shared memory, and a data race can leave the result correct.
The diagram shows a lost update in a shopping cart.
What does a race condition look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent writes a handler that reads the whole checkout record, sets the address, and writes the whole record back.
When the address changes, the checkout page also sends a request that recalculates shipping. That handler clears the cart's items and writes them back with updated prices.
The agent's tests pass. Then a customer reports a failure. The address saved, but the cart emptied. The server log for order A-1042 shows the order of events:
12:00:01.004 shipping A-1042 cleared items (count 0)
12:00:01.006 address A-1042 read checkout (count 0)
12:00:01.009 shipping A-1042 wrote items (count 2)
12:00:01.011 address A-1042 wrote checkout (count 0)
The address handler read the checkout during the 5 milliseconds when the cart was empty, then wrote that empty copy back last. On a laptop the requests rarely overlap this way, and a debugger or extra logging can shift the timing enough to hide the bug. A heisenbug is a bug that changes or disappears when someone observes it.
This example is simplified. A real checkout has more shared state, e.g. the stock count for each item.
How do you reproduce and test a race condition?
Reproducing a race means making the bad order happen more often, and then making it happen on purpose. Useful reproduction steps for a race name the operations that must overlap. These techniques run from cheapest to most exact:
- Repeat the test. Run it many times, e.g. with Playwright's
--repeat-each=100option. The failure count is a baseline for checking the fix. - Add pressure. Run the operations in parallel with more workers, or during load testing, so they overlap more often.
- Widen the window. In a test build, add a short delay between the read and the write, so a rare overlap becomes common.
- Force the order. Add test hooks that pause the operations at known points, and then release them one at a time in the bad order. The bad order then happens on every run, whatever the timing.
At Acme, a test that sends both requests at once fails in some runs:
test("editing the address keeps the cart", async () => {
for (let run = 0; run < 200; run++) {
const order = await startCheckout({ items: 2 });
await Promise.all([
order.setAddress("12 Elm Street"),
order.recalculateShipping(),
]);
const saved = await loadCheckout(order.id);
expect(saved.items).toHaveLength(2);
}
});
A test that pauses both handlers and releases them in the logged order fails on every run until the code is fixed:
test("an address save during recalculation keeps the cart", async () => {
const order = await startCheckout({ items: 2 });
const cleared = pauseAt("shipping:after-clear"); // test build only
const saving = pauseAt("address:before-write");
const recalc = order.recalculateShipping();
await cleared.reached;
const save = order.setAddress("12 Elm Street");
await saving.reached;
cleared.release();
await recalc;
saving.release();
await save;
expect((await loadCheckout(order.id)).items).toHaveLength(2);
});
In end-to-end testing, two browser sessions on one account can send the requests together. For data races in memory, a race detector records memory accesses while tests run. Go has one built in (go test -race), and Clang and GCC offer ThreadSanitizer (-fsanitize=thread). A detector finds data races only in code that the tests run.
Some property-based testing libraries run generated operations in parallel and check that the result matches some serial order of the same operations. Other tools run one test under many random thread schedules, which applies fuzzing to timing instead of input.
What changes when a coding agent writes the code?
A coding agent's tests often run one operation, wait for its result, and then run the next. The operations never overlap, so the race window never opens. In the Acme example, each unit test called one handler on its own and passed.
Timing failures also make a wrong fix look right. Asked to fix an intermittent failure, an agent can add a short sleep before the address save. The two requests then overlap less often, and the race is still in the code. A green run of the concurrent test, or a pass on retry, is weak evidence of a fix.
A team can give the agent a target that waiting cannot meet. Ask for a test that forces the bad order and fails before the fix. After the fix, run the concurrent test many times, and reject changes that only add sleeps or retries.
How do you fix or prevent a race condition?
Common fixes remove the unsafe order instead of making it rarer:
- Make the update atomic. Replace a separate read and write with one operation, e.g.
UPDATE carts SET item_count = item_count + 1, or lock the row before reading it. - Use a lock. A mutual exclusion lock, or mutex, lets one operation at a time run code that touches shared state.
- Check a version. Reject a write based on an older version of the record, so the caller retries with fresh data.
- Make repeats safe. Idempotency means a repeated request has the same effect as one request, e.g. a unique order key that stops a double click from creating two orders.
- Write only what changed. At Acme, the address handler can update the address field, not the whole checkout record.
A lock slows code that waits for it. Two locks taken in opposite orders can cause a deadlock, where each operation waits for the other forever. Many passing runs show only that the bad order did not happen in those runs. A test that forces the order is the stronger check.
How is a race condition different from a flaky test?
A race condition is a defect in code. A flaky test is a test that passes and fails on the same code, and a race is one of several causes.
A race in the product makes a correct test flaky and is a real bug. A race in the test itself, e.g. a fixed sleep that sometimes ends before the page loads, makes the test flaky while the product works. Other flaky tests have no race at all, e.g. a test that fails for lack of test isolation. Treating every flaky test as noise hides races that customers can hit.
FAQs
How do you unit test multithreaded code?
Unit testing multithreaded code works best when the test controls the order of threads instead of waiting for a bad one. Pause each thread at a known point with a test hook, and then release them in the order that fails.
Can a race condition happen in single-threaded code?
A race condition can happen in single-threaded code when asynchronous tasks interleave. In Node.js, two async functions that await network calls can run in an order that changes between runs, so the same code and input give different results.
Can retries hide a race condition?
Retries can hide a race condition, because a retry runs the test again and the second run can get a safe timing. The test turns green while the bug stays in the code. Customers can still hit the bad order when their requests overlap.
Can a debugger hide a race condition?
A debugger can hide a race condition, because breakpoints and stepping change how long each operation takes. The bad order may stop happening while someone watches, which makes the race a heisenbug. A test that forces the order does not depend on timing, so it still fails under a debugger.