What is an agent harness?
An agent harness is the program that runs a large language model (LLM) in a loop and supplies its tools, context, and memory. It is also called agent scaffolding. The model only generates text, including requests to use tools. The harness carries out those requests and returns the results, and the two together form an AI agent.
Every coding agent, e.g. Codex, is a model inside a harness. Claude Code's documentation describes Claude Code as "the layer around the model that provides the tools and manages the context the model sees." It calls that layer the agentic harness.
An agent framework is a library that developers use to build agents, while a harness is the running software around a model. The line blurs when a framework ships with a loop and tools, and Anthropic's engineering post on long-running agents calls its Claude Agent SDK an agent harness.
The harness changes results even when the model stays the same. A small September 2026 study compared a review model in a harness that runs several review passes with the same model given a single prompt. The harness found about 1.6 times as many verified bugs, at about 10 times the tokens.
How does an agent harness work?
A harness usually does six jobs around the model:
- Context. The harness assembles what the model reads on each turn, from its own instructions and project files, e.g. AGENTS.md, to earlier tool output. Selecting that input is called context engineering.
- Tools. The harness tells the model which tools exist, e.g. one that runs shell commands. When the model returns a tool call, the harness runs it and adds the output to the context.
- Loop. The harness sends the updated context back to the model until the model returns no tool call, or until a limit or a person stops it.
- Permissions. The harness sets which calls run without asking and which wait for a person's approval. It can run commands inside a sandbox, which limits the files and hosts they reach.
- Memory. The harness keeps state that outlasts one context window, e.g. a progress file for the next session. When the context fills, it can summarize older messages, which is called compaction.
- Hooks. The harness can run scripts at fixed points in the loop, called agent hooks, e.g. a script that runs the tests when the model tries to stop.
Some model APIs run a few tools on the provider's servers, e.g. web search, but those tools still run outside the model. A harness can also start a subagent, a second agent with its own context that returns only its result.
What is an example of an agent harness?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a terminal coding agent to "Let customers edit their delivery address during checkout." Its harness offers file and shell tools, but no browser tool. The harness handles each step around the model:
- The harness loads Acme's
AGENTS.md, which says to runnpm test, into the context with the task. - The model requests an edit to
checkout/AddressForm.tsx. The harness checks its permission rules, which allow edits inside the project, and writes the file. - The model requests
npm test. The rules require approval for shell commands, so the developer approves it. The harness runs it and returns42 passed. - The model returns a summary with no tool call. A Stop hook runs
npx tsc --noEmit, which exits with code 0, so the harness ends the turn.
The developer opens the checkout in a browser and changes the address. The address saved, but the cart emptied. The harness offered no tool that loads the checkout page, and neither the tests nor the hook checked the cart after an address change. This example is simplified. A real session would include more tool calls, e.g. reads of the files around the checkout.
How does MCP relate to an agent harness?
The Model Context Protocol (MCP) is a standard way to connect a harness to outside tools and data. In MCP terms, the harness is the host. It starts one MCP client for each configured server, asks each server for its tools, and offers them to the model next to its built-in tools. When the model calls an MCP tool, the harness sends the call to that server and adds the result to the context.
MCP covers only that connection. The loop, the context, memory, and permission rules stay with the harness. Two harnesses that use the same MCP server can still load and approve its tools differently.
What are the limits of an agent harness?
A harness cannot make the agent's work correct. It has these limits:
- It runs requests without checking that they are correct. A permission rule can stop a risky command, but not a wrong edit. The final summary is still text from the model, so a turn can end with a "done" that no check supports.
- Its tools limit what can be checked. Anthropic's post reports that its agent could not detect the browser's own alert dialogs through its browser tool, and features that used those dialogs tended to have more bugs.
- Its checks can be weakened from inside. A hook protects a check only while the agent cannot edit it. An agent that can edit the tests a Stop hook runs can make the hook pass.
- More harness work costs more. Each extra pass or check adds tokens and time to a task.
- It widens what a mistake can reach. A shell tool runs commands with the user's permissions unless a sandbox limits them.
How is an agent harness different from a test harness?
A test harness runs tests against software. The International Software Testing Qualifications Board (ISTQB) defines a test harness as "a collection of drivers and test doubles needed to execute a test suite." An agent harness runs a model instead, and the model is the part that makes requests, not the part under test.
In a coding agent, the agent harness often starts a test harness, e.g. by running npm test, and the result becomes the model's next input. Wikipedia's article on agent harnesses also names benchmark "evaluation harnesses" as a related sense, so it helps to say which sense is meant.
What should a harness let an agent check?
An agent can check only what its harness gives it tools to reach. Designing the harness, including its checks, is called harness engineering. Anthropic's post names a failure mode in which the agent marked a feature "as complete without proper testing." The post reports that browser automation tools, and a prompt to use them, let the agent find and fix bugs "that weren't obvious from the code alone."
A harness can offer these checks:
- Tests and type checks. A shell tool lets the model run the test suite and the type checker after each edit.
- The running software. A browser tool, or a command that starts the app, lets the model use the changed feature, e.g. change the delivery address and read the cart.
- Checks at fixed points. A hook runs a check whether or not the model asks for it, e.g. before the turn ends.
- A second reader. A subagent or an LLM-as-a-judge setup reviews the diff with its own context. A judge that only reads the diff does not run the code.
- Checks the agent cannot edit. Tests and hook settings outside the agent's write access keep a check from being edited until it passes.
These checks still report to the same model that wrote the code. A team can also add a check outside the agent's session, e.g. one that runs on the pull request.
FAQs
Is Claude Code an agent harness?
Claude Code is an agent harness. Its documentation calls it the layer around the model that provides the tools and manages the context. A Claude model generates each step, and the two together make up the coding agent.
How is an agent harness different from a framework?
An agent framework is code that developers build an agent with, while an agent harness is the running program that wraps a model in a loop with tools. A framework can supply most of a harness, and Anthropic calls its Claude Agent SDK an agent harness.
Why does one model perform differently in two harnesses?
One model performs differently in two harnesses because each harness gives it different tools, context, and loop rules. A model with a tool that runs the tests can check an edit, while the same model without that tool can only read it.
Does the harness or the model run the tools?
The harness runs the tools, not the model. The model generates a tool call, and the harness runs it and returns the output. Some model APIs run a few tools on the provider's servers, which are still outside the model.