Skip to content

What is behavior-driven development (BDD)?

Behavior-driven development (BDD) is a practice in which a team agrees on plain Given-When-Then scenarios before coding, and tools run them as tests.

Last updated , 9 min read

What is behavior-driven development (BDD)?

Behavior-driven development (BDD) is a software practice in which business people, developers, and testers describe a feature's behavior as examples in plain language. Each example is a scenario in the Given-When-Then format, which names a starting state, an action, and an expected outcome. The team agrees on them before coding, and a tool can run each as a test.

The International Software Testing Qualifications Board (ISTQB) defines behavior-driven development as a collaborative approach to development. In that approach, the team focuses on delivering the expected behavior of a component or system for the customer, and that behavior forms the basis for testing.

Dan North described how the practice began in Introducing BDD. He kept meeting the same confusion while teaching test-driven development (TDD). Talking about behavior instead of tests cleared up much of it. With a business analyst, he later applied the idea to requirements, and Given-When-Then scenarios came from that work.

In BDD, the tests come out of building the software, not from a separate phase of software testing. BDD scenarios usually serve as the acceptance criteria for a change. The people who ask for the change and the people who build it then check against the same examples.

How do Given-When-Then scenarios work?

A Given-When-Then scenario is one concrete example of a rule, written as steps. Product owners and testers can write the steps, and a developer writes the code behind them, called step definitions. Cucumber and other BDD tools read scenarios in Gherkin, a text format whose files end in .feature. Cucumber's Gherkin reference defines each step keyword:

  • Given. A Given step puts the software in a known state. Its step definition builds the scenario's test fixture.
  • When. A When step describes the action or event under test.
  • Then. A Then step states the expected outcome, which should be observable outside the system. That outcome is the scenario's test oracle.
  • And or But. These keywords continue the step type above them.

A Scenario Outline repeats one scenario for each row of an Examples table, so the same steps check several values. A test runner handles each scenario in this order:

  1. The runner reads the feature file and splits each scenario into steps.
  2. The runner matches each step's text to a step definition, a function with a pattern for that text. Cucumber ignores the keyword when it matches.
  3. The Given definitions set up the state.
  4. The When definition drives the software, often through its user interface as an end-to-end test.
  5. The Then definitions compare the actual outcome with the expected one, and the runner reports the first step that fails.
How a Given-When-Then scenario runs Scenario (plain text) Step definitions (code) Software Given 2 items in the cart Setup creates the checkout Known state When the address changes Action calls setAddress Action applied Then the cart still holds 2 Assertion pass or fail Actual result
The scenario text holds no checks of its own. The check happens in the Then step's definition, which compares the actual result with the expected one.

Given, When, and Then follow the arrange-act-assert layout of many unit tests, but as sentences for the whole team instead of code.

What is an example of a BDD scenario?

Here is an illustrative example. Acme Co. sells furniture online. Its product owner asks the team to "Let customers edit their delivery address during checkout."

Before the work starts, the owner, a developer, and a tester talk through examples in a meeting often called the Three Amigos. The tester asks what happens to the cart, and the owner says it must keep its items. The developer writes the agreed example in checkout.feature:

Feature: Delivery address at checkout

  Scenario: Customer edits the delivery address
    Given a customer has 2 items in the cart
    When the customer changes the delivery address to "12 Elm Street"
    Then the order shows the delivery address "12 Elm Street"
    And the cart still holds 2 items

Next, the developer writes the step definitions in JavaScript for Cucumber. Each pattern matches one step, and {int} and {string} pass the step's values to the function:

import { Given, When, Then } from "@cucumber/cucumber";
import { expect } from "expect";
Given("a customer has {int} items in the cart", async function (count) {
  this.checkout = await startCheckout({ items: count });
});
When("the customer changes the delivery address to {string}",
  async function (text) {
    await this.checkout.setAddress(text);
  });
Then("the order shows the delivery address {string}", function (text) {
  expect(this.checkout.address).toBe(text);
});
Then("the cart still holds {int} items", function (count) {
  expect(this.checkout.items).toHaveLength(count);
});

The developer runs the scenario before changing any code. It fails at the When step, because setAddress does not exist yet. The developer writes the change and runs the scenario again. The first three steps pass, and the shortened report shows the last one failing:

1) Scenario: Customer edits the delivery address
   And the cart still holds 2 items
     Expected length: 2
     Received length: 0

The address saved, but the cart emptied. The change rebuilt the checkout session and dropped the cart. The step that caught it came from the tester's question. The developer keeps the existing cart when the address changes, and all four steps pass.

The feature file stays in the repository as living documentation, a readable record of the rule. This example is simplified. A real feature would have more scenarios, e.g. one for an empty address field.

What changes when a coding agent writes the code?

Given-When-Then scenarios give a coding agent a precise target it can run. When people approve the scenarios before the agent starts, as in spec-driven development, each Then step holds an expected result that did not come from the agent's code. A failing scenario also names the step and the values, which the agent can use to find the fault.

The weak point moves into the step definitions. A reviewer can approve the feature file without opening the code behind it, which an agent may write or edit. A step that reads "the cart still holds 2 items" passes on broken code if its definition only checks that the cart exists, e.g. with toBeDefined() in place of toHaveLength(count). An agent that repairs a failing scenario can also loosen a definition instead of fixing the code, and the scenario text stays the same.

A practical adjustment is to treat step definitions as tests, not as glue code. A person writes or approves the definition of each Then step, because it holds the check. That person sees the step fail once, e.g. against code that empties the cart. Reviewers read each later change to a step definition next to the step it serves.

What are the limits of BDD?

BDD has costs, and a passing feature file leaves some failures unchecked:

  • Scenarios cover only the examples someone wrote. A case nobody discussed, e.g. a customer who edits the address twice, has no scenario.
  • Steps that script the screen break often. A step that names a button or a field ties the scenario to one layout, so a redesign can fail it while the behavior still works.
  • Scenarios need upkeep. Each step needs a definition whose pattern matches its wording, so rewording a step means editing code too. A scenario that is skipped instead of updated after the rule changes describes behavior the software no longer has, a form of spec drift.
  • Scenarios are slow and broad. Many run through the user interface, so they take longer than unit tests and point less precisely at the faulty function. Teams usually keep unit tests for the functions a scenario uses.
  • The format alone adds little. A team that writes Gherkin after the code, without the people who asked for the feature, gets the upkeep without the shared examples.

How are BDD scenarios different from acceptance tests?

An acceptance test is defined by its level and purpose. It checks finished software against the needs of its users and owners, so they can decide whether to accept it. A BDD scenario is defined by its form and its timing. It is a Given-When-Then example that the team agrees on before the code exists.

The two often overlap. Many teams write their acceptance tests as BDD scenarios before the code, one form of acceptance test-driven development (ATDD). A scenario that follows one whole task through the running software is also a user journey test. They differ at the edges. A product owner's manual trial is an acceptance test with no scenario, and a scenario can check one small part well below the acceptance level.

FAQs

Can people who do not code write BDD scenarios?

People who do not code can write BDD scenarios, because each step is a plain sentence. In practice they usually write them with a developer or tester, who makes the wording match the patterns in the step definitions.

How is Given-When-Then different from arrange-act-assert?

Given-When-Then and arrange-act-assert split a test into the same three parts, which are setup, action, and check. Given-When-Then is written as sentences that the whole team can read. Arrange-act-assert is a layout for the code inside one unit test.

What is Gherkin?

Gherkin is the text format of the feature files that Cucumber and other BDD tools read. Each step starts with a keyword, e.g. Given, and the tool matches the rest of the step's text to a step definition, the code that runs it.

Do BDD scenarios replace unit tests?

BDD scenarios do not replace unit tests. In the Acme example, the scenario found the empty cart, and a unit test on the checkout code would point more precisely at the function that dropped it. Teams usually keep both kinds of test.