What is harness engineering?
Harness engineering is the practice of setting up a coding agent's tools, instructions, checks, and loop so that the agent's mistakes are found and fixed. It covers more than prompt engineering, which changes only the wording a model receives. When a mistake gets through, a harness engineer changes the setup, not only the request.
An agent harness runs a model in a loop and supplies its tools, context, and memory, so it sets what the agent can reach and check. In an April 2026 article, Birgitta Böckeler writes that "harness" has become shorthand for "everything in an AI agent except the model itself." The word has an older meaning in testing, where a test harness is the code that runs tests against software.
How does harness engineering work?
Harness engineering runs as a slower loop around the agent's own loop. Böckeler calls the parts that act before the agent does "guides," e.g. an instruction file. She calls the parts that act after it "sensors," e.g. a test suite. One round of the slower loop runs in these steps:
- A person gives the agent a task inside the current harness.
- The guides shape the first attempt, e.g. a rule that names the test command.
- The sensors check the result, some inside the loop and some when the agent stops or later.
- A failed check goes back to the agent with its output or reproduction steps, and the agent tries again.
- A mistake that gets past the sensors, or repeats across tasks, leads a person to change a guide or a sensor.
- The person reruns the same tasks with the same model and compares results from checks the agent did not write.
Böckeler describes the person's job as steering the agent by changing the harness. Guides without sensors give an agent that "encodes rules but never finds out whether they worked." Sensors without guides let the same mistakes repeat.
A check's cost in time and tokens sets where it can sit. A type checker can run in the loop after each edit. A browser test can run when the agent stops. Mutation testing can run outside the session, after the change merges. A team's definition of done lists which checks a change must pass.
What is an example of harness engineering?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout."
The agent reports "done" with 42 passed. The address saved, but the cart emptied. The developer treats the miss as a harness problem:
- The developer finds the gap. No instruction mentioned the cart, and no check loaded the checkout page.
- The developer adds a guide to AGENTS.md, the line "Changing the delivery address must not change the cart."
- The developer adds a sensor, a browser test in
acceptance/checkout-address.spec.tsthat changes the address and checks that the cart still holds 2 items. - The developer adds a deny rule for
acceptance/to the agent's permission settings, so the agent's edit tools cannot change that test. - The developer adds a Stop hook that runs the browser test when the agent tries to end its turn.
- The developer reverts the agent's change and gives the agent the same request. The browser test fails with
Expected: 2andReceived: 0, and the agent fixes the handler before its final message. - The developer keeps the request as a test task for later harness changes.
This example is simplified. A real team would try a harness change on several tasks, e.g. recent checkout requests, before keeping it.
What does a harness engineer change?
A harness engineer changes the parts of the harness that users control. Böckeler separates the harness built into the agent, e.g. its system prompt, from an "outer harness" that users build with the features the agent offers.
The changes fall into these parts of the outer harness:
- Instructions. The engineer adds a line to the instruction file when a mistake repeats.
- Skills and tools. The engineer packages a procedure as an agent skill, which the agent loads when a task needs it. A tool, e.g. a browser tool, adds an action the model can request with a tool call.
- Checks. The engineer adds tests, type checkers, and linters, which give the same result for the same code unless a test is flaky. A model reviewer, e.g. an LLM-as-a-judge setup, can grade what no fixed rule states, but its results vary.
- Check output. The engineer writes a custom check's failure message to say how to fix the problem, because the message goes into the agent's context.
- Hooks. The engineer adds a hook that runs a check at a fixed point in the loop, even when the model does not request it.
- Permissions. The engineer writes permission rules that set which files the agent can change and which commands it can run, including the files that hold the checks.
Anthropic's post on long-running agents describes changes of this kind. A first session writes a feature list and marks each feature as failing. Later sessions are told to edit that list only by changing a feature's passes field, not by removing features.
What are the limits of harness engineering?
A harness makes a correct change more likely, not certain, for these reasons:
- A check finds only what it checks. Each sensor is only as good as its test oracle, the source of its expected result. No findings is not the same as complete coverage.
- Behavior is the hardest part. Böckeler writes that most people who give agents high autonomy check functional behavior with tests the agent generated plus manual testing. She calls that "not good enough yet." Beyond plain failures, e.g. a crash, a sensor cannot judge behavior that nobody specified.
- A stronger harness can cost more. In a small September 2026 study, a review harness of several passes found about 1.6 times as many verified bugs as a single prompt, at about 10 times the tokens.
- Rules have gaps. A deny rule blocks only the calls it matches, so a script can still change a protected file.
- Failures can repeat. A fix can break something else or bring back an earlier failure. This AI bug-fix loop runs until a person or an attempt limit stops it.
How is harness engineering different from agentic engineering?
Harness engineering is one part of agentic engineering, the wider practice of building software with coding agents while keeping engineering practices, from written specs to review. Harness engineering sets up what the agent works inside. The rest of agentic engineering is work that people do outside the harness, e.g. deciding what to build.
Context engineering overlaps with both. Böckeler calls engineering a user's harness "a specific form of context engineering." Guides and check output reach the agent through its context, but harness engineering also covers permissions and when each check runs.
How do checks feed back into the harness?
A failed check does two jobs. It sends the agent the steps that caused the failure. If the failure repeats, it leads a person to change the harness, e.g. by adding a rule. The agent can edit checks inside its session, and deny rules have gaps. Keep at least one check outside the agent's session, where the agent cannot edit it.
RunStory runs your software against the change. It sends reproducible failures back to your coding agent, with the actions RunStory took and evidence of the unexpected result. RunStory is in private alpha for CLIs and web apps.
FAQs
Is harness engineering a kind of prompt engineering?
Harness engineering is broader than prompt engineering. Prompt engineering changes the wording a model receives, while harness engineering also changes the tools, checks, permissions, and loop around it. A reworded instruction is one small harness change.
Who does harness engineering on a team?
Harness engineering on a team usually falls to the developers who use the coding agent. A platform team can also keep shared instruction files, skills, and checks. The agent's maker builds the part that ships inside the agent.
How do you measure whether a harness change helped?
A harness change is measured by running the same tasks with the same model before and after the change. Compare the results of checks the agent did not write. Also compare the tokens and time each run used, since a harness change can add cost as well as findings.
Can a team change the harness of a coding agent it did not build?
A team can change much of the harness of a coding agent it did not build. The maker controls built-in parts, e.g. the system prompt. Users can add instructions, skills, tools, checks, hooks, and permission rules through the features the agent offers.