Skip to content

What is Playwright MCP?

Playwright MCP is a server that lets an AI agent drive a real browser through the Model Context Protocol, using page snapshots instead of screenshots.

Last updated , 8 min read

What is Playwright MCP?

Playwright MCP is an open-source Model Context Protocol (MCP) server that lets an agent operate a real web browser through tools. The agent reads each page as a structured text snapshot of its elements, not as a screenshot. An action on an element, e.g. a click, usually names its target by a reference from that snapshot.

The Playwright project publishes the server as the npm package @playwright/mcp. It builds on Playwright, an open-source browser test framework. Each snapshot is a text form of the page's accessibility tree, a structured view of elements with their names and roles. The official documentation says the server "operates on the accessibility tree, not pixels."

A coding agent that edits a web app can use Playwright MCP to open the app and try its change. Playwright MCP is not a test runner, and a session leaves no test file behind unless the agent writes one. The server works on web pages. An agent checks a command-line application by running it in a shell instead.

How does Playwright MCP work?

A session with Playwright MCP runs in six steps:

  1. A person registers the server with the agent, e.g. claude mcp add playwright -- npx @playwright/mcp@latest in Claude Code.
  2. The agent's MCP client asks the server for its tools, e.g. browser_click, and offers them to the model.
  3. The model generates a tool call, e.g. browser_navigate with the app's URL.
  4. The server runs the matching Playwright action in the browser, which opens in a visible window by default.
  5. The server returns the result, including the Playwright code it ran and the page's state.
  6. The model generates the next call from the last snapshot and names its target element by reference.
How an agent drives a browser through Playwright MCP Coding agent model and MCP client tool call Playwright MCP server actions Browser shop.example.com result and snapshot with element references
The agent reaches the browser only through tool calls and reads back text, plus an image when it asks for a screenshot. Each next call names an element from the last snapshot.

In a snapshot, each element carries a reference, e.g. e5. A reference stops working when its element leaves the page, e.g. a button that a click removed. A call that uses it fails with an error that asks for a fresh snapshot, so the agent takes one.

The --headless option runs a headless browser, for a machine with no display. By default the server keeps a persistent browser profile, so logins survive between sessions. The --isolated option keeps the profile in memory only. Options go after the package name. Other MCP clients list both in a configuration file:

{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest", "--isolated"]
    }
  }
}

What is an example of Playwright MCP?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The developer then asks it to try the change through Playwright MCP, as a test customer with 2 items in the cart.

The agent calls browser_navigate with https://shop.example.com/checkout, and the result includes a snapshot. This excerpt shows the form and the cart:

- textbox "Delivery address" [ref=e14]
- button "Save address" [ref=e15]
- list "Cart" [ref=e20]:
  - listitem [ref=e21]: Oak chair
  - listitem [ref=e22]: Pine desk

The agent calls browser_type with the text "12 Elm Street" and the reference e14. It then calls browser_click with e15. The click's result includes the Playwright code that ran, await page.getByRole('button', { name: 'Save address' }).click(), and a fresh snapshot. The address is on the page, and the cart list is gone:

- paragraph [ref=e31]: 12 Elm Street
- paragraph [ref=e32]: Your cart is empty

The agent reports the failure with the calls it made, so the developer can repeat them. This example is simplified. A real session would also read the browser console with browser_console_messages and save a screenshot as evidence.

What changes when a coding agent writes the code?

When the coding agent that wrote a web change tries it through Playwright MCP, the run can catch a failure the agent's tests missed. The run starts from the same context as the code, so it tends to try the path the agent built.

The browser can also carry real logins. The default profile persists, so the agent can act with any account a person signed into there. The --extension option, which needs Playwright's browser extension, connects the agent to a running Chrome or Edge browser where a person is already signed in. Page text enters the model's context, so a page the agent opens can carry prompt injection.

A practical adjustment is to run the server with --isolated and give the agent a test account, so a session holds no real login. When the agent finds a failure, turn its steps into a committed end-to-end test so the check repeats.

What does it cost in tokens and time?

A live session has these costs and limits:

  • Tool definitions. The server's tool definitions usually load into the model's context when the session starts.
  • Snapshots. By default each action also captures a snapshot, which runs to many lines on a long page. Each one the agent reads takes context away from the code. The browser_snapshot tool can limit the depth of the tree or save it to a file.
  • Time. Each step waits for the model to generate a call and read the result, so the agent's conversation stays busy until the session ends. A script that runs the same steps without a model finishes much sooner.
  • Repeatability. A failure can mean the app broke or the agent acted on the wrong element. That mixed signal is the one a flaky test gives.

By default a snapshot lists names and roles, not positions, so a layout fault, e.g. a label drawn over a price, shows in a screenshot, not in the snapshot. A session that finds nothing is not the same as complete coverage.

The Playwright project also publishes Playwright CLI (playwright-cli), a command-line tool that a coding agent runs through its shell. It is headless by default and writes snapshots to files. It is not the npx playwright test command. It installs with npm install -g @playwright/cli@latest.

The Playwright MCP README suggests the CLI for coding agents, because CLI calls "avoid loading large tool schemas and verbose accessibility trees into the model context." The same README says MCP still suits agent loops that benefit from "persistent state."

How is Playwright MCP different from committed Playwright tests?

A committed Playwright test is a script in the repository that runs the same steps and checks each time, e.g. in continuous integration (CI). A Playwright MCP session has no script, so it can reach states no test covers but does not repeat reliably. The table compares them:

AttributePlaywright MCP sessionCommitted Playwright test
Who sets the stepsThe model, during the runA person or an agent, in advance
Needs a modelYesNo
ResultA report in the agent's chatA pass or fail for each test
Leaves behindTool calls and the Playwright code they ranA test file that CI can run again

A session is a form of agentic testing, and without a script it works as exploratory testing.

Playwright's test agents connect the two. They are definitions made of instructions and MCP tools, added with a command, e.g. npx playwright init-agents --loop=claude for Claude Code. The planner explores the app and writes a Markdown test plan, and the generator turns the plan into test files. The healer replays a failing test and suggests a patch, or marks the test skipped when the feature looks broken, so a person should review each patch.

When should an agent explore and when should a test run?

Use an agent in a browser to try a change that no test covers yet, or to reproduce a reported bug. Once a workflow works, keep it working with a Playwright test in the repository that CI runs. The session's Playwright code gives that test a starting point.

While a coding agent drives a browser, its own conversation stays busy. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

Should the browser agent run tests in CI?

A browser agent makes a weak CI check, because its steps can change between runs and each run needs a model. Its failures also mix broken apps with actions on the wrong element. Committed Playwright tests give CI a stable pass or fail for workflows that must keep working.

What are Playwright test agents?

Playwright test agents are agent definitions that ship with Playwright and are made of instructions and MCP tools. The planner writes a Markdown test plan, the generator turns it into test files, and the healer suggests patches for failing tests. A person reviews each patch, because the healer can mark a test skipped instead.

Should you use Playwright MCP or the Playwright CLI?

The Playwright CLI usually suits a coding agent better than Playwright MCP, because CLI calls avoid loading large tool definitions and page snapshots into the model's context. Playwright MCP suits agent loops that benefit from persistent state.

Does Playwright MCP need a visible browser?

Playwright MCP does not need a visible browser, although it opens a visible window by default. A headless option runs the browser without a window, which suits a machine with no display.