# What is a test harness?

A test harness is the code and tools that run tests against a system, with drivers that start it, test doubles for its dependencies, and checks on results.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   Define a test harness
-   Explain how a harness drives the system under test
-   Compare a test harness with an agent harness

## Related content

-   [What is software testing?](https://specstory.com/learning/testing/software-testing)
-   [What is test automation?](https://specstory.com/learning/testing/test-automation)
-   [What is unit testing?](https://specstory.com/learning/testing/unit-testing)
-   [What is the difference between mocks, stubs, and fakes?](https://specstory.com/learning/testing/mocks-vs-stubs)

## What is a test harness?

A test harness is the code around a system under test that runs tests on it and checks the results. Its main parts are fixtures for setup, drivers that start the system, [test doubles](https://specstory.com/learning/testing/mocks-vs-stubs) for its dependencies, and checks on its output. In AI coding, an [agent harness](https://specstory.com/learning/ai-coding/agent-harness) is a different program that runs a language model's tools and loop.

The International Software Testing Qualifications Board (ISTQB) [defines a test harness](https://glossary.istqb.org/en_US/term/test-harness) more narrowly, as the drivers and test doubles needed to run a test suite. A driver is code that calls or controls the part being tested in place of its real caller. The system under test (SUT) is the part of the software that a test runs and checks, as opposed to the harness around it.

In [software testing](https://specstory.com/learning/testing/software-testing), teams build a harness when a system cannot run alone in a test. They also build one when its real dependencies are slow or unsafe to use, e.g. a payment service that charges real cards. Harnesses usually serve [test automation](https://specstory.com/learning/testing/test-automation), so the same checks can repeat on each change without a person. A test result is only as trustworthy as the harness that produced it.

## How does a test harness work?

A test harness takes one test through six steps:

1.  A [test fixture](https://specstory.com/learning/glossary#test-fixture) sets up the state the test needs, e.g. a temporary folder with a sample order file.
2.  The harness replaces dependencies that are slow or outside the team's control with test doubles, e.g. a stub that returns a fixed shipping rate.
3.  A driver starts the system under test or calls it with the test input, e.g. a command with its flags.
4.  The harness captures what the system returns, e.g. its output and its exit code.
5.  Checks compare the captured results with the expected results from a [test oracle](https://specstory.com/learning/test-quality/test-oracle).
6.  The harness removes what the fixture created, and the test runner reports the test as passed or failed.

Diagram: The parts of a test harness

Everything inside the dashed line except the system under test is harness code. A bug in any of those parts can change the result.

The driver depends on the kind of system. In [unit testing](https://specstory.com/learning/testing/unit-testing), the test function is usually the driver, and the test framework supplies most of the rest. A harness for a [command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications) usually starts the program as a separate process. A harness for a [terminal UI](https://specstory.com/learning/cli-and-web/testing-terminal-uis) often runs the program in a pseudoterminal and checks the screen it draws.

Some harnesses run on generated inputs. In [fuzzing](https://specstory.com/learning/test-quality/fuzzing), a fuzzer generates the inputs, and the harness is often one small function that passes each input to the code under test, e.g. libFuzzer's `LLVMFuzzerTestOneInput`.

## What is an example of a test harness?

Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, `acme`, exports orders. A developer asks a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) to add a `--status` flag, so that `acme export` can list only paid orders. The developer writes this harness with pytest:

```python
import json, os, subprocess


def test_export_lists_only_paid_orders(tmp_path):
    store = tmp_path / "orders.json"  # fake order store
    store.write_text(json.dumps([
        {"id": "A-1041", "status": "unpaid", "total": 95.00},
        {"id": "A-1042", "status": "paid", "total": 240.00},
    ]))
    env = {**os.environ, "ACME_ORDERS_FILE": str(store)}
    result = subprocess.run(  # driver
        ["acme", "export", "--format", "csv", "--status", "paid"],
        env=env, capture_output=True, text=True)
    assert result.returncode == 0  # checks
    assert result.stdout.splitlines()[1:] == ["A-1042,240.00"]
```

The harness runs the test in five steps:

1.  The built-in `tmp_path` fixture of pytest gives the test its own temporary folder.
2.  A fake order store, a JSON file with one unpaid and one paid order, replaces the order database.
3.  The environment variable `ACME_ORDERS_FILE` points the tool at the fake store.
4.  The driver calls `subprocess.run`, which starts the real `acme` command as a separate process and captures its output and exit code.
5.  The checks expect exit code 0 and, after the header row, one data row, `A-1042,240.00`.

The agent adds a filter function and a unit test for it, and the unit test passes. The harness reports a failure, because the output still holds both orders. The command parses `--status` but never passes it to the filter, and the unit test called the filter directly. This example is simplified. A real harness would also check an error path, e.g. an order file that does not exist.

## What changes when a coding agent writes the code?

When a coding agent writes the code, it often writes the harness in the same session. The drivers, test doubles, and checks then come from the same prompt as the code, so they can share its gaps. In the Acme example, the agent's unit test drove the filter function it wrote, not the command a user types.

The agent receives the harness's result through its own tools, and a zero exit code often ends the task with a "done" message. A harness bug that passes a broken program then becomes a false "done." The agent can also edit the harness instead of the code, e.g. by marking the failing test as skipped. The run then passes, and the system under test has not changed.

A team can keep the drivers and checks for its main workflows in a folder that is read-only for the agent. A person then reviews each change to those files separately from changes to the code.

## When can the harness itself be the bug?

A harness is code, so it can have bugs. A harness bug can fail a correct program. It can also pass a broken one, which is worse, because nobody looks further. Four harness bugs are common:

-   **It reports the wrong exit code.** In Bash, a pipeline returns the exit status of its last command unless the `pipefail` option is set. A harness that runs `acme export | head -n 5` gets the status of `head`, so a crash in `acme` passes.
-   **It never runs the checks.** When pytest searches a folder, it collects only files named `test_*.py` or `*_test.py` by default. A file of checks named `export_checks.py` never runs, and the rest of the suite still passes.
-   **It leaks state between tests.** A fixture that does not clean up leaves data that the next test reads, which breaks [test isolation](https://specstory.com/learning/environments/test-isolation). The result then depends on the order the tests run in.
-   **It never reaches the code.** A fuzz harness that returns early on most inputs rarely reaches the parser it targets. The run reports no crashes, and a quiet run looks the same as a clean one.

The fix is to test the harness too. A team can run it once against a build with a known bug and confirm that it fails. The code that grades results belongs outside the system under test, so a bug in the system cannot change the grade. In the words of the post on [testing an Electron app](https://specstory.com/blog/ten-ways-to-let-an-agent-test-an-electron-app), "A harness that grades itself has the same bugs as the code it grades."

## How is a test harness different from a test framework?

A test framework is a reusable library for writing and running tests, e.g. pytest. It usually ships a test runner, the program that finds tests, runs them, and reports each result. A test harness is the setup for one system under test, and it is usually built from a framework plus project code.

The ISTQB draws the line the other way. It defines a [test automation framework](https://glossary.istqb.org/en_US/term/test-automation-framework) as a set of test harnesses and test libraries, so in its terms the framework is the larger whole.

## How is a test harness different from an agent harness?

An agent harness is the program around a language model that gives it tools, context, memory, and a loop, which turns the model into a working agent. A test harness runs a system under test and checks its results.

The two meet when a coding agent tests its own work. The agent harness runs the test harness as a tool and returns its output to the model. The test harness produces the evidence, and the agent harness passes it on. A benchmark's evaluation harness is a third sense. It is a test harness whose system under test is a model or an agent.

## FAQs

### What is a system under test?

A system under test is the part of the software that a test runs and checks, e.g. the Acme export command. The harness around it, including its fixtures and test doubles, is not part of the system under test, so a bug there is a harness bug.

### When should you use a test harness?

A test harness is worth building when a test cannot run the code alone or safely. A command-line tool usually needs a driver that starts it as a separate process. Code that calls a slow or unsafe service needs a test double in its place.

### Do unit tests need a test harness?

Unit tests need a test harness, but a test framework usually supplies most of it. In pytest, the test function acts as the driver, fixtures set up state, and the runner reports results. A team adds harness code of its own when a unit needs test doubles or special setup.

### Is a test runner the same as a test harness?

A test runner is not the same as a test harness, because the runner is only one part of it. The runner finds tests, runs them, and reports each result. The rest of the harness is the fixtures, drivers, test doubles, and checks that each test needs.

---

Source: [What is a test harness? | Test vs. agent harness | SpecStory](https://specstory.com/learning/testing/test-harness)
