Skip to content

What is the difference between black-box and white-box testing?

Black-box testing checks software from the outside against what it should do, while white-box testing designs tests from the code's internal structure.

Last updated , 9 min read

What is the difference between black-box and white-box testing?

Black-box testing designs tests from what the software should do, while white-box testing designs tests from the code itself. A black-box test checks only what a user or caller can observe. A white-box test picks inputs so that chosen branches and paths run. Black-box tests catch missing or wrong behavior, and white-box tests catch faults in the code's branches.

The International Software Testing Qualifications Board (ISTQB) defines black-box testing as a test approach based on the specification of a component or system. Its glossary defines white-box testing as one based on the internal structure of the component or system.

Both terms describe how tests are designed, not a level of software testing. A unit test written from a function's documented behavior is black-box, and one written to run each of its branches is white-box.

What is black-box testing?

Black-box testing, also called specification-based testing, is a test approach that designs tests from the specification, e.g. a user story, not from the code. The tests reach the software only through its interface:

  • Web app. A browser test clicks and types through the pages, as an end-to-end test does.
  • API. An API test sends a request and checks the status code and the response body.
  • Command-line tool. A test runs a command-line application with arguments and checks its output and exit code.

Test design techniques choose which inputs to try. The ISTQB Foundation syllabus names four black-box techniques, which are equivalence partitioning, boundary value analysis, decision table testing, and state transition testing. Many of the inputs they pick sit off the happy path, e.g. an empty address field. The syllabus puts exploratory testing in a third group, based on the tester's experience. It usually runs from the outside.

Black-box tests can also check non-functional qualities, e.g. how long the checkout page takes to load. They stay useful when the code is rewritten and the required behavior stays the same. They cannot show which parts of the code ran.

What is white-box testing?

White-box testing, also called structural or glass box testing, is a test approach that designs tests from the internal structure of the code. The tester reads a function, finds its decisions, and picks inputs that send a run down each chosen route. Progress is measured with code coverage at one of these depths:

  • Statement coverage. Statement coverage is the share of executable statements that the tests ran.
  • Branch coverage. Branch coverage is the share of branches that the tests ran, e.g. both outcomes of an if statement. Full branch coverage includes full statement coverage, but not the reverse.
  • Path coverage. Path testing runs chosen routes through a function from entry to exit. Paths multiply with each decision, so path testing picks a subset. Cyclomatic complexity counts the linearly independent paths, which is the minimum number of tests for basis path testing.

White-box testing covers the code that exists, so it can find defects even when the specification is vague. The same structure can be checked without a run, which is what static analysis tools do. White-box tests depend on the code, so a rewrite that keeps the behavior can still break them.

When should a team use black-box or white-box testing?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Both kinds of test run on the change:

  1. The agent writes change_address and 3 pytest tests, one for each route through its code. They pass a valid address, an empty one, and one outside the delivery area.
  2. The tests pass with no missed statements or branches in checkout.py.
  3. A reviewer writes a black-box test from the request, which implies that editing the address keeps the cart. It adds 2 oak chairs to the cart, saves a delivery address, and counts the cart items.
  4. The test finds 0 items where it expected 2. The address saved, but the cart emptied.
  5. Saving the address starts a fresh checkout session, and no code copies the cart into it. No branch existed for the unit tests to run.
  6. The agent edits change_address to copy the cart into the fresh session, and the black-box test passes.

The coverage report from step 2 reads:

Name          Stmts   Miss Branch BrPart  Cover
-----------------------------------------------
checkout.py      11      0      4      0   100%
-----------------------------------------------
TOTAL            11      0      4      0   100%

The reviewer's test from step 3 reads:

import { test, expect } from "@playwright/test";
import { addToCart } from "./helpers";

test("changing the address keeps the cart", async ({ page }) => {
  // baseURL in playwright.config.js points at the build under test
  await addToCart(page, "oak chair", 2);
  await page.goto("/checkout");
  await page.getByLabel("Delivery address").fill("12 Elm Street");
  await page.getByRole("button", { name: "Save address" }).click();
  await expect(page.getByTestId("cart-count")).toHaveText("2 items");
});

Use white-box tests for the logic inside a function. Use black-box tests for each behavior someone asked for, because they can catch one that is missing. This example is simplified. A real checkout would also test the order total.

What changes when a coding agent writes the code?

When a coding agent writes its tests from the same prompt as the code, with the code open, those tests are usually white-box in practice. The ISTQB syllabus notes that when the code does not implement a requirement, white-box tests may not detect the gap.

An agent that runs only its own tests is checking its own code, and its tests share the code's gaps. Black-box tests do not come from the agent's code, and they let a person judge a change by its behavior, which is how a team can check unread code.

Black-box checks need only the request, so a team can write them before the agent starts, which white-box tests cannot do. Keep them in files the agent does not edit. Run them on a build that records coverage, and read the report. Code that no black-box check reaches may be behavior that nobody asked for.

Can black-box and white-box testing be used together?

A team can run both as separate suites, as Acme did. Gray-box testing, also spelled grey-box, mixes the two inside one test. The ISTQB glossary defines it as a test approach that combines elements of black-box and white-box testing. A gray-box test reaches the software through its interface and also uses some knowledge of the inside:

  • Inputs chosen from the code. The tester reads the code to find a risky branch, then reaches it through the page, e.g. an address outside the delivery area.
  • State checked from the inside. The test uses the page, then reads the database or the logs to check what the software saved.
  • Coverage used to pick inputs. The team runs its black-box tests on a build that records coverage, then adds tests for code no run reached. Go's go build -cover makes such a binary, which needs GOCOVERDIR set.
What each kind of test reaches The software Interface page, API, CLI Code functions, branches, and paths Database and logs Black-box test from the request Gray-box test also reads the inside White-box test from the code
Only the black-box test is limited to what a user or caller can reach. The gray-box test adds a check on saved state, and the white-box test aims at the code's branches.

Used together, neither approach shows that a program is correct, because each checks only the inputs and paths someone chose.

How do black-box and white-box testing compare?

The two approaches differ on these points:

PointBlack-box testingWhite-box testing
Tests are designed fromThe requirements and expected behaviorThe structure of the code
Needs the source codeNoUsually
Can be writtenAs soon as the requirements existOnly after the design or code exists
FindsMissing features and wrong resultsWrong branches, and code that coverage shows no test ran
MissesCode paths the requirements never mentionRequirements the code never implemented
Measure of progressRequirements or workflows checkedStatement, branch, or path coverage
After a rewrite with the same behaviorTests still applyTests may need changes
Most often used forEnd-to-end, API, and command-line testsUnit tests

How do you test the running app from the outside?

A running app needs a check from the outside, not only tests written with the code open. Run the workflows a user relies on against a fresh build, and check what the app shows at each step. Keep the steps, so a failure can be repeated.

RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

What is gray-box testing?

Gray-box testing is a test approach that combines black-box and white-box testing. The tester drives the software through its interface and can also read its database or logs, e.g. to check what a checkout saved.

Is unit testing white-box testing?

Unit testing is a test level, not a design approach, so a unit test can be black-box or white-box. Many unit tests are white-box in practice, because their author writes them with the code open. A unit test written only from a function's documented behavior is still black-box.

Is exploratory testing black-box testing?

Exploratory testing is usually black-box testing in practice, because the tester works through the interface without reading the code. The ISTQB syllabus still puts it in a third group of techniques, based on the tester's experience.

Does code coverage apply to black-box tests?

Code coverage can apply to black-box tests when the software runs as a build that records coverage. The test still ignores the code, but the report shows which code the run reached. Without such a build, black-box testing gives no measure of code coverage.