Skip to content

What is spec-driven development?

Spec-driven development is a way of working where a written specification, not a chat prompt, drives what a coding agent builds and how it is checked.

Last updated , 8 min read

What is spec-driven development?

Spec-driven development is a way of working in which a written specification comes first, and a coding agent writes the code from it. The spec states what the software should do and why, and the team uses it to guide and check the work. It is also called SDD or specification-driven development.

A spec differs from a chat prompt in two ways. It is a file that people review before the code is written, and it lists conditions that the finished change must meet. GitHub's open-source Spec Kit states the order as "Define what and why before deciding how to build it."

SDD is often presented as the alternative to vibe coding, where a person prompts an agent and accepts its code without reading it. A written spec gives the reviewer something to compare the result with. The spec does not run anything, so the result still has to be checked when the software runs.

How does spec-driven development work?

Most spec-driven workflows move one feature or change through the same documents in order:

  1. A person describes the change, and the agent drafts a spec with acceptance criteria that a person or a test can check.
  2. The person reviews the spec and edits its criteria, because the agent's draft fills gaps with defaults that nobody chose.
  3. The agent writes a plan that says how to build the change, e.g. which files and data it touches.
  4. The agent splits the plan into an ordered list of tasks.
  5. The agent carries out the tasks and writes the code and its tests.
  6. A person or a test runner checks the result against the criteria and against the team's definition of done, which applies to every change.
How a spec reaches a coding agent Spec what and why Plan how to build it Tasks ordered steps Coding agent writes the code Running software the built change Check each criterion, on a real run acceptance criteria
The acceptance criteria skip the plan and the tasks. They come back only at the end, when the running software is checked against them.

Spec Kit names its steps specify, plan, tasks, implement, and converge, and adds one file of project principles that it calls a constitution. Its spec template has sections for user scenarios, requirements, success criteria, and assumptions. Its converge step compares the code with the spec, plan, and tasks, and adds a task for each gap it finds.

A product requirements document (PRD) usually describes a whole product or release for the people who build it. A spec usually covers one feature or change in enough detail for an agent to work from. The spec records what was decided. Most of the exchange that led there, e.g. the options a person turned down, stays in the session history.

What are spec-first, spec-anchored, and spec-as-source?

Birgitta Böckeler describes three levels of spec-driven development, set by what happens to the spec after the first task:

  • Spec-first. The spec is written and reviewed before any code, and it guides the task at hand. Keeping the spec current afterward is not part of the method.
  • Spec-anchored. The team keeps the spec after the task and updates it along with the code. Each later change to the feature starts from the current spec.
  • Spec-as-source. The spec is the main source file. People edit only the spec, and the agent generates the code from it.

Every approach in her review was spec-first, and not all of them aimed to be spec-anchored or spec-as-source. A spec that nobody updates leads to spec drift, where the written spec no longer describes what the software does. She also compares spec-as-source with older methods that generated code from structured designs, e.g. diagrams, instead of prose. She wonders whether it could end up with their inflexibility.

What is an example of spec-driven development?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The agent drafts a spec, and the developer edits its acceptance criteria to read:

Spec: Edit the delivery address during checkout

Acceptance criteria
1. The customer can change the delivery address on the review step.
2. The order confirmation shows the new address.
3. The cart keeps its items and total when the address changes.
4. An empty address shows an error and keeps the old address.

The change then moves through these steps:

  1. The agent writes a plan that edits the checkout page and the order's address field.
  2. The agent completes 4 tasks and ends with a completion claim that the work is "done" and its tests pass.
  3. The agent's tests check the address field. None of them checks the cart.
  4. The developer runs end-to-end tests, written from the 4 criteria before the agent wrote any code, against a staging copy of the store at staging.shop.example.com.
  5. The test for criterion 3 fails. The address saved, but the cart emptied.
  6. The developer sends the failing steps to the agent, and the agent fixes the code until the test passes.

The spec did not catch the bug. It put criterion 3 in writing before any code existed, and the run against that criterion caught it. This example is simplified. A real spec would need more criteria, e.g. one for a shipping cost that changes with the address.

Is spec-driven development a return to waterfall?

Spec-driven development repeats waterfall's order of requirements, design, and code, but usually for one feature at a time. A waterfall project finishes each phase for the whole system before the next one starts. In SDD, a team can revise the spec after a run and go through the loop again, as agile teams do with each iteration. It looks more like waterfall when specs get long and are treated as fixed.

The approach also has limits of its own:

  • Review cost. Specs, plans, and task lists add documents for a person to read, and Böckeler found one toolkit's files repetitive and tedious to review.
  • Skipped instructions. Böckeler often saw the agent miss instructions in the files it was given. The same spec can also produce different code on different runs.
  • Existing code. On a codebase that already exists, the spec has to describe the behavior around the change too. Böckeler found that 2 of the 3 toolkits she tried took more work to introduce there.
  • Small changes. A fix of one line rarely needs a spec, and a full spec for it adds more review than it saves.
  • Wrong spec. Code can match a spec that is wrong. Standards call a check against the spec verification and a check against users' needs validation. Code built from a wrong spec passes only the first.
  • Unwritten behavior. A spec covers only the behavior someone wrote down. Exploratory testing looks for the rest.

How is spec-driven development different from test-driven development?

Test-driven development (TDD) starts each change with a failing test, and SDD starts it with a written spec. A TDD test is code that covers one small behavior, while a spec is prose that covers a whole change. The two combine when each criterion becomes a test that is written before the code. A 2004 paper by Ostroff, Makalsky, and Paige described an agile form of SDD that combined TDD with design by contract.

How do you check what a spec-driven change does when it runs?

Turn each acceptance criterion into an acceptance test that runs the software, e.g. an end-to-end test, and keep those checks in files the agent does not edit. Run them after the agent reports the work finished, then test what the spec left out.

RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Does spec-driven development work on existing code?

Spec-driven development can work on existing code, but the spec has to cover more than the change itself. It also has to state which current behavior must stay the same, e.g. that the cart keeps its items. Some toolkits take more work to set up there, and a fix of one line rarely needs a spec.

What is the difference between a spec, a plan, and a PRD?

A spec, a plan, and a PRD differ in scope and in what they settle. A PRD usually covers a whole product or release, while a spec usually covers one feature or change and states its behavior and purpose. A plan comes after the spec and says how the agent will build the change, e.g. which files it edits.

Who should write the spec, the developer or the agent?

The agent often drafts the spec from a short request, but a person should own it. The developer reads the draft and edits the acceptance criteria, because the agent fills gaps with defaults. The change later passes or fails its checks against those criteria.

How do you keep a spec and the code in step?

A spec and the code stay in step when each later change starts with an edit to the spec and then to the code. Böckeler calls a spec kept this way spec-anchored. A spec that nobody updates drifts until it describes behavior the software no longer has.