# What is test-driven development (TDD)?

Test-driven development is writing a failing test before the code, then the least code that passes it, then refactoring, in short repeated cycles.

Last updated September 29, 2026, 7 min read

## Learning objectives

After reading this article you will be able to:

-   Define test-driven development and red-green-refactor
-   Explain what TDD protects against
-   Identify how TDD changes when an agent writes code

## Related content

-   [What is software testing?](https://specstory.com/learning/testing/software-testing)
-   [What is unit testing?](https://specstory.com/learning/testing/unit-testing)
-   [What is the difference between mocks, stubs, and fakes?](https://specstory.com/learning/testing/mocks-vs-stubs)
-   [What is the difference between unit, integration and end-to-end tests?](https://specstory.com/learning/testing/unit-vs-integration-vs-e2e-tests)

## What is test-driven development (TDD)?

Test-driven development (TDD) is a programming practice in which a developer writes a failing test first, then the smallest code change that makes it pass. The developer then cleans up the code through [refactoring](https://specstory.com/learning/glossary#refactoring), with the test still passing. Each cycle covers one small piece of behavior, and the cycles repeat until the feature is complete.

Red-green-refactor is the TDD cycle of writing a failing test, making it pass, and cleaning up the code while the test stays green. The colors come from test runners, many of which show a failing test in red and a passing test in green.

TDD is a way to write code, not a separate testing phase. It relies on [software testing](https://specstory.com/learning/testing/software-testing), and it is most often used for [unit testing](https://specstory.com/learning/testing/unit-testing). The practice protects against tests that cannot fail, because each test is seen failing before its code exists. It also protects against later breakage, because the finished tests stay in the suite and run after each change. Many teams use one part of the cycle, e.g. a failing test that reproduces a bug before the fix, which then stays in the suite for [regression testing](https://specstory.com/learning/testing/regression-testing).

## How does test-driven development work?

One TDD cycle runs in five steps:

1.  The developer writes a short list of behaviors the change needs, e.g. "an edited address is saved."
2.  The developer picks one item and writes a test for it, with an assertion that states the expected result.
3.  The test runner runs the test, and the test fails. This is the red step, and the failure should come from the missing behavior, not from a typo.
4.  The developer writes the smallest code change that makes this test and every earlier test pass. This is the green step.
5.  The developer refactors the code and the tests where the design needs it, then runs the suite again to confirm that it is still green.

Diagram: The red-green-refactor cycle

Each cycle adds one failing test and the least code that passes it. The loop back to red starts the next behavior on the list.

The red step checks the test itself. A test that has never failed may pass whether or not the code works, e.g. when it compares a value with itself. The green step asks for the least code on purpose, so that each behavior in the code has a test that required it. Writing the test first also makes the developer call the code before it exists, which can show early when its interface is hard to use.

## What is an example of test-driven development?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme gets the request "Let customers edit their delivery address during checkout."

The first item on the developer's test list is "an edited address is saved and the cart keeps its items." The developer turns that one item into a test and runs it with pytest:

```python
from checkout import Cart, Checkout

def test_edit_address_keeps_cart():
    checkout = Checkout(cart=Cart(item_count=2))
    checkout.set_address("12 Elm Street")
    assert checkout.address == "12 Elm Street"
    assert checkout.cart.item_count == 2
```

The test fails, because `set_address` does not exist yet:

```text
>       checkout.set_address("12 Elm Street")
E       AttributeError: 'Checkout' object has no attribute 'set_address'
```

That is red, for the expected reason. The developer adds a `set_address` method that stores the address, and the test passes. That is green.

During the refactor, the developer rewrites the method's body as a call to an existing `restart_checkout` helper that stores the address the same way. The helper also replaces the cart with an empty one, so the test fails with `assert 0 == 2`. The address saved, but the cart emptied. The developer reverts the refactor, the test turns green again, and the bug never reaches a customer.

This example is simplified. A real project would add more items to the list, e.g. rejecting an empty address.

## What changes when a coding agent writes the code?

When a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) writes the test and the code in one edit, the red step is lost. The agent can also copy expected values from the code's output. Kent Beck's [Canon TDD](https://newsletter.kentbeck.com/p/canon-tdd) guide lists that copying as a mistake. Such [AI-written tests](https://specstory.com/learning/verification/ai-written-tests-always-pass) pass whether or not the code is right.

In [Dan Luu's 2026 experiment](https://danluu.com/agentic-testing/), agents told to use named techniques mostly did not beat the default prompt. A TDD prompt led to more tests and iteration, yet fewer runs passed every hidden test. Few runs followed TDD closely. Iterating on their own tests tended to make agents write more tests that enforced incorrect behavior.

On [ImpossibleBench](https://arxiv.org/abs/2510.20270), where passing a task requires breaking its specification, GPT-5 "passed" 54% of one set of impossible SWE-bench tasks, by editing the tests or gaming them. Hiding the tests cut cheating to near zero.

Hidden tests do not fit TDD, because the agent runs each test. Read-only tests do. A person can write or approve each test and see it fail before the code exists, so no expected value comes from the code's output. A review of the diff checks for [test tampering](https://specstory.com/learning/verification/test-tampering), a form of [reward hacking](https://specstory.com/learning/verification/reward-hacking), and for special cases added for the tested inputs.

## What are the disadvantages of TDD?

TDD has costs, and it leaves some failures unchecked:

-   **The test shares the code's assumptions.** The source of each expected result, the [test oracle](https://specstory.com/learning/test-quality/test-oracle), is usually the same person who writes the code. If that person misreads the request, the test and the code agree and both are wrong.
-   **Doubles hide failures between parts.** TDD tests often replace databases and services with [test doubles](https://specstory.com/learning/testing/mocks-vs-stubs). The suite can pass while the real parts fail together.
-   **Coverage can overstate the checks.** Code written this way often has high [code coverage](https://specstory.com/learning/test-quality/code-coverage), but coverage counts the lines that ran, not the results that were checked. [Mutation testing](https://specstory.com/learning/test-quality/mutation-testing) measures whether the tests fail when the code is broken on purpose.
-   **It costs time up front.** Each change starts with a test, and tests tied to implementation details need rework during refactoring. [A pooled analysis](https://doi.org/10.1109/TSE.2012.28) of TDD studies found a small positive effect on quality and little to no clear effect on productivity. In industrial studies, the quality gain and the productivity drop were both larger than in academic ones.

## How is TDD different from BDD?

[Behavior-driven development](https://specstory.com/learning/glossary#behavior-driven-development) (BDD) grew out of TDD and keeps its order of test before code. The main difference is who reads the tests. BDD scenarios are Given-When-Then steps in plain language that the team agrees on and a tool runs as tests. TDD tests are code that developers write, one small behavior at a time. The two work together, e.g. a BDD scenario for the whole checkout, with TDD for each function it needs.

[Spec-driven development](https://specstory.com/learning/ai-coding/spec-driven-development) applies a similar order to coding agents, with a written specification before the code.

## FAQs

### Is TDD with a coding agent still TDD?

TDD with a coding agent is still TDD when each test fails before the agent writes the code that passes it. If the agent writes the test and the code in one edit, or copies the code's output into the test, the red step is missing.

### Can a person write the tests while an agent writes the code?

A person can write the failing tests while a coding agent writes the code, and the split means the agent does not write the expected results. The person runs each test to see it fail first, and a review of the diff checks that the agent did not edit the tests.

### Do teams practice TDD?

Teams practice TDD to different degrees. Some run the full red-green-refactor cycle on each change. Many keep one habit from it, a failing test that reproduces a bug before the fix, which then stays in the suite as a regression test.

### Does TDD slow development down?

TDD adds work at the start of each change, because each change begins with a test. Studies of TDD differ on whether it slows a team overall, and the tests it leaves behind keep checking later changes.

---

Source: [What is TDD? | Test-driven development | SpecStory](https://specstory.com/learning/testing/test-driven-development)
