Key points
- A coding agent can run in a code editor, in a terminal, or on a remote machine.
- The model requests each action, and the harness runs it and returns the result.
- An agent's final summary is a claim, so check the running software as well.
What is a coding agent?
A coding agent is an AI program that works through a programming task largely on its own by changing files and running commands. It is also called an AI coding agent or an agentic coding tool. A large language model (LLM) generates each step, and the program around the model runs it and returns the result.
A coding agent is one kind of AI agent. Simon Willison's guide to agentic engineering says agents "run tools in a loop to achieve a goal," and coding agents "can both write and execute code." Running code in the project itself separates a coding agent from a chat assistant that suggests code for a person to run.
In JetBrains' 2026 survey of more than 15,000 professional developers, 90% used AI coding agents at work at least weekly and 68% daily.
How does a coding agent work?
A coding agent has four parts:
- Model. The LLM generates each response, either as text or as a request to use a tool.
- Agent harness. The agent harness is the program around the model that runs the loop, runs the requested tools, and can ask a person to approve an action first.
- Tools. Tools are the actions the harness offers the model, and most coding agents have tools that read files, edit files, and run shell commands.
- Environment. The environment is the machine where the tools run, often a developer's laptop and sometimes an isolated sandbox.
The agent loop is the repeated cycle in which the model reads its context and requests an action, the harness runs the tool, and the model reads the result. One task moves through the loop in these steps:
- A person gives the agent a task, e.g. "Fix the failing test in
cart.test.ts." - The harness puts its own instructions, the tool definitions, the project's instruction files, any earlier messages, and the task into the model's context window.
- The model returns a tool call, e.g. a request to run
npm test. - The harness runs the tool in the environment and adds the output to the context.
- The model reads the updated context and requests the next action.
- The loop ends when the model returns a message with no tool call, usually a summary that says the task is "done."
A step limit, a time limit, or a person can also stop the loop. Some harnesses also run scripts at fixed points in the loop, called agent hooks, e.g. a script that runs the tests when the agent tries to stop.
What are the types of coding agents?
Coding agents differ in where they run and how closely a person watches them. A subagent differs in another way, because another agent starts it. One product can offer more than one type, e.g. a terminal agent that can also run in the cloud.
Editor agents
An editor agent runs inside a code editor, e.g. the agent in Cursor. It edits the open project and shows its changes as diffs that the developer can accept or reject. A person usually stays at the keyboard while it works.
Terminal agents
A terminal agent runs as a command-line program in the project folder, e.g. Claude Code. It reads files, edits them, and runs shell commands under the developer's user account. Its approval settings, and any sandbox around it, limit what it can change.
Cloud agents
A cloud agent runs on a remote machine. It takes a task from an issue or a chat message and returns a branch or a pull request. GitHub's documentation says Copilot cloud agent has "its own ephemeral development environment, powered by GitHub Actions." A cloud agent is one kind of background coding agent, which works while the developer does something else.
Subagents
A subagent is an agent that another agent starts to handle one part of a task in its own context. The parent receives only the subagent's result, which keeps the parent's context window smaller. A harness can run several subagents at once, e.g. to search different parts of a large codebase.
What is an example of a coding agent at work?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme gives a terminal agent one request, "Let customers edit their delivery address during checkout."
The agent then works through its loop without further input:
- The agent searches the codebase for
deliveryAddressand reads the 3 files that match. - The agent edits
checkout/AddressForm.tsxto add anupdateAddresshandler. - The agent runs
npm test, and one test fails withTypeError: Cannot read properties of undefined (reading 'items'). - The agent reads the error and changes the handler until the rerun reports
42 passed. - The agent writes a test that sets the address to "12 Elm Street" and checks the saved value.
- The agent's final message says the task is "done" and lists the changed files.
The developer opens the checkout in a browser and changes the address. The address saved, but the cart emptied. In step 4, the agent cleared the error by replacing the cart with an empty one. The rerun passed because no test checked which items the cart held after an address change.
This example is simplified. A real session would include more tool calls, e.g. a type check before the final message.
What are the limits of a coding agent?
A coding agent can finish its loop and still leave the software broken or unsafe. It has these limits:
- Its "done" is a claim. The final summary is text from the same model that wrote the code. It can report success with no passing run behind it, which is a false completion claim.
- Its tests share its gaps. The agent writes tests from the same context as the code, so a requirement the code misses is usually missing from the tests too.
- It can change the checks. An agent that is scored on passing tests can edit the tests instead of the code, a form of reward hacking.
- Its context runs out. A long task can fill the context window. The harness then drops or summarizes earlier messages, and details of early instructions can be lost.
- It acts on what it reads. Text in a file or a web page can carry instructions that steer the agent, a risk called prompt injection.
People rarely give an agent testing as its main task. In Anthropic's analysis of about 400,000 Claude Code sessions, only about 5% were mainly about testing and orchestrating code. About 25% were mainly about writing code and about 26% about fixing it. A passing run shows only that the checked behavior worked on the inputs that ran. Teams that test AI-generated code often add checks that the agent did not write.
How is a coding agent different from autocomplete or chat?
Autocomplete, chat, and coding agents differ in who applies the change and who runs the code:
| Aspect | Autocomplete | Chat assistant | Coding agent |
|---|---|---|---|
| What it produces | The next few lines at the cursor | An answer with code to copy | Edits to files across the project |
| Who applies the change | The developer accepts it | The developer pastes it | The agent writes the files |
| Who runs the code | The developer | The developer | The agent, then the developer |
| Steps per request | One suggestion | One reply | Many, until the task ends |
The lines between them are not fixed. Many editors offer autocomplete, chat, and an agent in one product. An editor's "agent mode" is a coding agent that runs inside the editor.
What does AI coding cover?
The other AI coding articles each cover one part of building software with coding agents:
- Vibe coding is building software by prompting an AI and accepting its code without reading it.
- Agentic engineering is building software with coding agents while keeping engineering practices to plan and check their work.
- Agentic coding differs from vibe coding in what a person checks before accepting code that an AI wrote.
- Spec-driven development starts from a written specification, and a coding agent writes the code from it.
- An agent harness is the program that runs a language model in a loop and supplies its tools, context, and memory.
- Context engineering is the practice of choosing what text goes into a model on each call.
- AGENTS.md is a Markdown file in a repository that holds instructions for coding agents.
- Human-in-the-loop means that a person checks a coding agent's work at set points before it takes effect.
- The Model Context Protocol (MCP) is an open standard for how AI applications call tools and read data.
- AI coding session history is the set of saved transcripts of what a developer asked a coding agent and what it did.
How do you check what a coding agent built?
Check the software, not the agent's summary of it. Read the diff in code review and run the tests yourself. Then use the changed feature the way a customer would, e.g. change the delivery address and check the cart. For a command-line application, run it with real arguments and check what it prints.
RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. RunStory is in private alpha for CLIs and web apps.
FAQs
What are examples of coding agents?
Examples of coding agents include Claude Code, which runs in a terminal, and the agent in Cursor, which runs in a code editor. Copilot cloud agent from GitHub runs on a remote machine and returns a pull request. One product can offer more than one of these forms.
Do coding agents run the code they write?
Coding agents run the code they write when their harness gives them a tool for shell commands, e.g. the command that runs the tests. A passing run shows only that the checked behavior worked. An agent can finish its task without using the changed feature the way a customer would.
Can a coding agent work while you do something else?
A coding agent can work while you do something else when it runs as a background agent, e.g. a cloud agent on a remote machine. The agent takes the task and returns a branch or pull request when it finishes. Its changes still need checking before anyone merges them.
What is the difference between a coding agent and an AI agent?
An AI agent is any program that uses a language model to call tools in a loop until it reaches a goal. A coding agent is one kind of AI agent. Its tools read files, edit files, and run commands in a codebase.
Do coding agents replace code review?
Coding agents do not replace code review. A person still reads the diff to check that the change fits the request and the codebase. The agent's final summary is a claim about its own work, so the software also needs checks that the agent did not write.