What is a computer-use agent?
A computer-use agent is an AI agent that operates a computer through its graphical interface, by reading screenshots and sending mouse and keyboard input. It is also called a computer-using agent (CUA). It can use apps that have no API, because it acts only on what the screen shows.
Computer use is the name model providers give this capability. Anthropic's computer use tool is one example. With it, a model requests screenshots and mouse and keyboard actions, and the developer's own application performs them on a desktop that the developer controls.
A coding agent changes code through files and a shell. A computer-use agent acts on the running app instead, so it can use a feature after a coding agent builds it. Some services give a background coding agent its own virtual machine, where it can open the app it changed and record screenshots of the run.
How does a computer-use agent work?
The agent harness runs the model in a loop and performs the actions the model requests, so the model never connects to the machine. The loop moves in this order:
- The harness sends the task, and the model often requests a screenshot first.
- The model returns one or more tool calls, each for one action, e.g. a left click at pixel coordinates 412, 318.
- The harness performs the actions in order, often on a virtual machine.
- The harness returns the results, usually with a screenshot.
- The model requests the next actions, or returns a final message with no tool call.
Computer use is built on tool calling, in which a model requests a named function and the harness runs it. A typical tool call returns data, e.g. an order as JSON, while a computer-use agent gets screenshots and has to locate each button by position.
A browser agent is an AI agent that operates a web browser instead of a whole desktop. Many read the page's accessibility tree instead of a screenshot. A coding agent can do the same through Playwright MCP, a Model Context Protocol (MCP) server. Scripted browser automation drives the same browsers with no model.
MDN describes the accessibility tree as the browser's record of the name, description, role, and state of most HTML elements. A button with no name appears there with a role only, a fault that accessibility testing reports. A computer-use agent clicks by position, so it can use that button and report no fault.
What is an example of a computer-use agent?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The developer then asks a computer-use agent in a virtual machine to change the address as a test customer and report the cart afterward.
The run has 5 steps:
- The model requests a screenshot, which shows the checkout page with 2 items in the cart.
- The model returns a triple click on the Delivery address field, which selects the old address, then a type action with "12 Elm Street" as the text.
- The model returns a click on the Save address button.
- The next screenshot shows the saved address and the message "Your cart is empty" under the cart heading.
- The model returns a final message that reports the saved address and the empty cart.
The harness logged these calls:
screenshot checkout page, cart shows 2 items
triple_click coordinate [412, 318] Delivery address field
type text "12 Elm Street"
left_click coordinate [540, 402] Save address button
screenshot 12 Elm Street, "Your cart is empty"
The address saved, but the cart emptied. The agent reported the empty cart because its task named the cart, not only the address. The saved screenshots let the developer check the report against what the screen showed before and after the change.
This example is simplified. A real run would also try the payment step, because an address change can alter the order total.
Can computer-use agents test software?
Computer-use agents can test software through its screen, which suits desktop apps and faults that show only in a picture, e.g. a label drawn over a price. Anthropic's documentation names automated software testing in trusted environments among the uses where the tool's slower pace is acceptable. A run in which an agent chooses, runs, and judges each check itself is agentic testing. Testing agents for web apps often use a browser tool instead, e.g. Playwright's test agents, which explore an app and turn a test plan into test files.
What are the limits of computer-use agents?
Computer use has these limits:
- Speed. Each group of actions waits for a screenshot and a model call before the next group runs. A script sends the same clicks with no model in between. An API or an agent-native CLI returns exact values in one tool call.
- Accuracy. Anthropic's documentation lists wrong coordinates and wrong tool choices as limits. In OSWorld, a 2024 benchmark of real computer tasks on Ubuntu, the agents tested failed mainly at locating screen elements and at operating apps.
- Repeatability. The model can take a different path on each run, and a misclick can pass for a bug. A check run this way can become a flaky test.
- Security. The agent can act on anything the machine shows, including accounts that are signed in. Text on the screen reaches the model, so a web page can carry prompt injection. Anthropic advises a dedicated virtual machine or container with minimal privileges, and no access to sensitive data, e.g. account logins.
- Data. Each screenshot goes to the model provider with whatever the screen shows, e.g. an open email. The tool and its settings control what leaves your machine.
How is a computer-use agent different from RPA?
Robotic process automation (RPA) is software that repeats a business task through an app's interface, following steps that a person recorded or wrote in advance. A computer-use agent generates each step from the most recent screenshot instead. The two differ on these points:
| Point | Computer-use agent | RPA bot |
|---|---|---|
| Who sets the steps | The model, during the run | A person, before the run |
| How it finds a button | Its position in the screenshot | Usually a selector recorded with the steps |
| When a screen changes | The model can often adjust | The bot usually stops until someone updates it |
| Cost of each step | A model call | Machine time only |
Both work through the interface when an app has no API. RPA suits frequent tasks on screens that rarely change. A computer-use agent suits tasks whose steps vary between runs. Some RPA products let a model choose steps that a fixed script cannot handle.
What does it take for an agent to test an app?
An agent that tests an app needs the built app running where the agent can reach it, with test accounts instead of real logins. Its goal comes from the change the developer asked for, and it names what must stay the same, e.g. the cart. Each action and its result go into a record, so a person can repeat a failure and tell a bug from a misclick.
RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. RunStory is in private alpha for CLIs and web apps. Desktop apps, mobile apps, and browser extensions are not in the alpha.
FAQs
What is a browser agent?
A browser agent is an AI agent that operates a web browser instead of a whole desktop, e.g. to fill in a form. Some browser agents read screenshots, and many read the page's accessibility tree instead, which records the name and role of most elements on the page.
Why is a computer-use agent slower than a script?
A computer-use agent is slower than a script because each step waits for a screenshot and for a model call that returns the next actions. A script sends the same clicks straight to the app, with its steps written before the run.
Is it safe to let a computer-use agent use your desktop?
Letting a computer-use agent use your everyday desktop is risky, because the agent can act on anything the screen shows, including accounts that are signed in. Text on the screen can carry prompt injection, and each screenshot goes to the model provider. A dedicated virtual machine with minimal privileges and no access to real accounts limits the risk.
How is computer use different from tool calling?
Computer use is a form of tool calling in which the tool acts on the screen, e.g. a click at pixel coordinates. The model checks the result in a screenshot. A typical tool call runs a function that returns data, e.g. an order as JSON from an API.