What is spec drift?
Spec drift is a mismatch between a software specification and its code, caused by one of them changing without the other. It is also called specification drift. A spec with drift still reads as correct, so people who trust it, and coding agents that read it, build on behavior the software no longer has.
Drift runs in both directions. The code can change while the spec stays the same, e.g. when an endpoint returns a field that its OpenAPI file does not list. The spec can also change while the code stays the same.
The spec itself raises no error when drift opens, because a document does not run. Spec drift is one kind of documentation drift, the general case in which any document, e.g. a README, stops matching the software.
A spec that stays current is what spec-driven development calls spec-anchored. Birgitta Böckeler describes that level as a spec "kept even after the task is complete" and used for later changes to the feature. She also found that the approaches she reviewed often leave open how a spec should be maintained over time.
How does spec drift start?
Spec drift usually starts after the task that the spec was written for. In a spec-first workflow, where nobody plans to update the spec after that task, drift opens in these steps:
- A person and a coding agent write a spec with acceptance criteria, and the agent builds the code from it.
- The task ends, and the spec stays in the repository as a Markdown file.
- A later prompt asks the agent to change a behavior, and the agent edits the code without opening the spec.
- The agent or a person updates the tests to match the changed code, so the suite still passes.
- The spec still describes the old behavior, and no check reports the difference.
- The next session reads the spec as current and builds on the old behavior.
Drift can also start on the spec side, when someone edits a requirement and nobody changes the code to match. It can even start on the first task. When the spec is silent on a detail, the agent fills the gap with a default, so the code holds a decision the spec never recorded. The same rule can also have copies outside the spec, e.g. a scenario in behavior-driven development, and each copy can drift on its own.
What is an example of spec drift?
Here is an illustrative example. Acme Co. sells furniture online. Its checkout spec has this criterion for editing the delivery address:
3. The cart keeps its items and total when the address changes.
The spec and the code then move apart:
- A coding agent builds the feature, and an acceptance test checks criterion 3 with an order total of $240.00.
- Later, a developer asks the agent to "Add the shipping cost to the total when the delivery address changes."
- The agent changes the address handler. The test order's total becomes $265.00, and the test for criterion 3 fails.
- The agent edits the test's expected total to $265.00, and the suite passes. Nobody edits the spec.
- In a later session about saved addresses, the agent reads criterion 3 and writes a test that expects the total to stay $240.00.
- That test fails. The agent removes the shipping cost and resets the older test to $240.00.
The suite passes, and checkout no longer charges shipping. The drift opened in step 3, when the code changed and the spec did not. Step 4 hid it.
Editing criterion 3 in step 4 would have closed the gap, e.g. "The cart keeps its items when the address changes, and only the shipping cost changes the total." This example is simplified. A real project would also update the rule's other copies, e.g. a user journey test of checkout.
Why do coding agents make drift worse?
A coding agent treats a spec as input, not as background. The agent reads project files into the model's context, e.g. a spec that AGENTS.md names, and the model generates code from that context. A stale line can therefore become a code change, as in step 6 of the example. A person may notice that a document is old, but the model gets no such signal unless the context contains one.
Agents also change the software in many small steps, and the reason for each change stays in the session, not in the spec. A plan from plan mode covers one task, and it changes the spec only if the plan includes that step. An agent can still skip an instruction to update the spec. In Böckeler's trials, the agent often did not follow all the instructions it was given.
To keep the spec in sync with the code, a team can put the spec edit in the same pull request as the code change. In spec-anchored work, each change starts with that edit.
A line in AGENTS.md can ask the agent to update the spec with each behavior change, and a reviewer reads both together. When the agent edits the spec after the code, the edit describes what the code now does. The reviewer checks that it is also what the team wants.
What are the limits of keeping specs current?
Keeping a spec in sync with the code has limits of its own:
- Two edits for each change. Each behavior change needs a spec edit and a review of that edit. For a fix of one line, the spec edit can take longer than the fix.
- Agreement is not correctness. A spec and code that agree can both be wrong about what users need. Checking code against its spec is verification, and checking it against users' needs is validation.
- Checks treat the spec as right. The
/speckit-convergecommand in Spec Kit uses the spec, plan, and tasks as its "sole source of intent" and adds a task for each gap it finds. Its instructions forbid edits to the spec, so a change made on purpose can come back as a gap to close. - Text does not show behavior. A spec and a diff are both text. Running the software, a core step when teams test AI-generated code, shows what the code does with the inputs that were tried.
How is spec drift different from a regression?
A regression is a defect in which something that worked before a change stops working after it. Spec drift is a disagreement between the spec and the code, and the software may still work as the team wants. Regression tests compare behavior with what worked before the change. A check for drift compares the code with its written spec.
The two meet in the checkout example. Adding the shipping cost was intended, so steps 3 and 4 created drift, not a regression. When the next session followed the stale spec and removed the shipping cost, the drift caused a regression.
What should run when the spec and the code disagree?
Run the software before editing either side. Acceptance tests written from the spec show where the running software and the spec differ, for the criteria they check. The prompt that changed the behavior often shows what the team meant, e.g. the checkout prompt that asked for the shipping cost. A person then decides which side is right and changes the other to match.
RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps. Your team keeps the final release decision.
FAQs
Who should update the spec when the code changes?
The spec should be updated by whoever makes the behavior change, in the same pull request as the code. A coding agent can draft the edit, and the reviewer checks that the spec states what the team wants, not only what the code now does.
Can a coding agent keep the spec up to date?
A coding agent can keep the spec up to date when its instructions ask it to edit the spec with each behavior change, e.g. a line in AGENTS.md. The agent can still skip that instruction, so a reviewer checks that each pull request that changes behavior also changes the spec.
Is spec drift the same as documentation drift?
Spec drift is one kind of documentation drift, which covers any document that stops matching the software, e.g. a README. A drifted spec also steers the code when a coding agent reads it as instructions, because the agent can change working code to match it.
Can tests detect spec drift?
Tests detect spec drift only for the criteria they check, and only while their expected values still come from the spec. When someone edits a test to match changed code, as in the checkout example, the test passes and the drift stays hidden.