Skip to content

What is browser automation?

Browser automation is controlling a web browser with code or an agent to open pages, click, type, and read results, for testing or repetitive tasks.

Last updated , 9 min read

What is browser automation?

Browser automation is software control of a web browser, in which code or an agent opens pages, clicks, types, and reads what the page shows. Tests use it to check a web app the way customers use it, often in several browsers. Scripts use it to repeat tasks on websites, e.g. filling in the same form each week.

Code usually drives the browser through an automation library, which sends commands over a control protocol. Playwright, Puppeteer, and Selenium are three open-source libraries of this kind. The browser can show a window or run as a headless browser with none.

In testing, browser automation is how test automation reaches a web app's interface, usually in end-to-end tests. A browser test framework is a library of this kind built for tests.

An HTTP client, e.g. curl, downloads a page without running its JavaScript, so it is not browser automation. Web scraping, the collection of data from websites, often uses such plain HTTP requests. Scrapers use browser automation when a page builds its content with JavaScript. A command-line application has no page to load, so its tests run the program and read its text output.

How does browser automation work?

A script drives a browser through a library in five steps:

  1. The library starts the browser, or attaches to a running one, and opens a control connection.
  2. The script calls a library method, e.g. a click on the Save address button.
  3. The library turns the call into protocol commands that find the element and act on it.
  4. The browser carries out the commands, runs the page's JavaScript, and sends back results.
  5. The script reads a value from the page, e.g. the cart count.

Most libraries connect through one of three protocols:

  • Chrome DevTools Protocol (CDP). The Chrome DevTools Protocol is an interface that lets a program drive and inspect Chrome and other Chromium-based browsers. Chrome's documentation describes it as a JSON-based protocol of commands and events. A client sends the commands over a WebSocket or a pipe. Puppeteer and Playwright use it to drive Chromium. For Firefox, Playwright uses its own patched build.
  • WebDriver. A WebDriver library sends each command as an HTTP request to a driver, e.g. ChromeDriver, which controls the browser and returns one response. The W3C standard grew out of Selenium, which uses it.
  • WebDriver BiDi. WebDriver BiDi is a W3C draft that adds messages in both directions to WebDriver over a WebSocket, so the browser can send events without a request. Firefox and Chrome support it.
How a library reaches the browser CDP or WebDriver BiDi Script or agent calls the library Library e.g. Puppeteer commands one connection results and events e.g. a page error, sent with no request Browser e.g. Chrome WebDriver Script or agent calls the library Library e.g. Selenium HTTP Driver e.g. ChromeDriver Browser the installed browser one HTTP request and one response for each command
CDP and WebDriver BiDi let the browser send events without a request. WebDriver answers each request with one response and sends nothing on its own.

Over CDP or WebDriver BiDi, the connection also carries JavaScript errors from the page. A script gets them only if it listens, which is how tests catch browser console errors.

What is an example of browser automation?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." Before any test exists, the developer asks the agent to try the change in a browser and report the cart afterward. The agent writes this Node.js script, try-address.mjs, with Playwright's library:

import { chromium } from "playwright";

const browser = await chromium.launch();
const page = await browser.newPage({ storageState: "test-customer.json" });
await page.goto("https://shop.example.com/checkout");
await page.getByLabel("Delivery address").fill("12 Elm Street");
await page.getByRole("button", { name: "Save address", exact: true }).click();
await page.getByText("Address saved").waitFor();
const items = await page.getByTestId("cart-item").count();
console.log(`cart items: ${items}`);
await browser.close();

The file test-customer.json holds the saved session of a test customer who is signed in and has 2 items in the cart. Playwright starts Chromium without a window and controls it over CDP. The click becomes CDP commands that move the mouse to the button, press it, and release it. In a shortened form, the press looks like this:

{
  "id": 58,
  "method": "Input.dispatchMouseEvent",
  "params": { "type": "mousePressed", "x": 212, "y": 388, "button": "left" }
}

The agent runs node try-address.mjs, and the script prints one line:

cart items: 0

The address saved, but the cart emptied. The script is the record of what the agent tried, so the developer can run it again. This example is simplified. A real project would turn the script into a committed test that fails whenever the cart count is not 2.

What changes when a coding agent writes the code?

A coding agent often writes the page and the automation that checks it from the same prompt and context. Its script then finds elements through markup the agent added a moment earlier, e.g. a data-testid attribute. A test ID matches an element whatever its visible label says. If the agent's address field has no label, a fill by test ID still passes.

When a step fails, the agent can also edit the script to match the page instead of fixing the page. It changes the text the step looks for to the text the page shows, e.g. "Place order" on the button that saves the address. The step then passes, and the script now expects the wrong label.

A team can ask the agent to find elements by role and exact visible label, as the script above does with getByRole("button", { name: "Save address", exact: true }). A missing or wrong label then fails the step. The team can also ask the agent to report each change to the text a script expects, so a reviewer checks it.

What are the limits of browser automation?

Browser automation runs steps and reads results. It does not judge whether those results are right. These limits apply to scripts and agents alike:

  • It reads only what someone asked for. A script that counts cart items misses a wrong order total. A hydration error can leave the page looking right while React reports an error that no step reads.
  • Timing makes runs unstable. The browser, the page's scripts, and the network each run at their own pace, so on some runs a step acts before the page is ready. The check is then a flaky test.
  • The protocols change. Chrome's documentation says CDP is not a public or supported API and does not guarantee backwards compatibility. A script that sends CDP commands directly can break after a Chrome update.
  • Some websites block it. The blocks aim to stop bulk copying of content and fake accounts. Under the WebDriver standard, a controlled browser sets navigator.webdriver to true, which a page can read. Bot defenses check other signals too.

A run with no failures is not the same as complete coverage.

How is scripted automation different from agent-driven browsing?

Scripted automation runs steps written in advance, so it repeats the same actions and needs no model. Agent-driven browsing adds a model to each step. The agent reads the page, usually as a text snapshot of its accessibility tree or a screenshot, and generates the next action. It often sends that action through a Model Context Protocol (MCP) server, e.g. Playwright MCP.

Both reach the browser through the same protocols. A script names its targets with locators, which are queries that match elements, and fails when one stops matching. An agent names its targets from the most recent snapshot, so it can follow a changed page or act on the wrong element. Checks run this way are a form of agentic testing.

A computer-use agent works at the level of the screen. It reads screenshots and sends mouse and keyboard input, so it can also drive apps outside the browser.

How do you check a change in a real browser?

Run the built app in a real browser and follow the changed workflow from start to finish. Check the results a customer reads, e.g. the cart after the address changes, not only the field that changed. Write those checks from the request, not from the markup the agent added.

RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. You keep working. It is in private alpha for CLIs and web apps.

Join the RunStory alpha →

FAQs

What is the Chrome DevTools Protocol?

The Chrome DevTools Protocol (CDP) is an interface that lets a program drive and inspect Chromium-based browsers with JSON commands and events. Puppeteer and Playwright turn a click in Chromium into CDP commands. Chrome's documentation says CDP is not a public or supported API, so its commands can change.

What is WebDriver BiDi?

WebDriver BiDi is a W3C draft standard that keeps one WebSocket connection open for commands and events. Classic WebDriver answers each HTTP request with one response, so the browser cannot report an event on its own. With BiDi, which Firefox and Chrome support, the browser sends events without a request.

Is browser automation the same as web scraping?

Browser automation is not the same as web scraping, although the two overlap. Browser automation operates a browser for tests and repeated tasks. Web scraping collects data from websites, and a scraper turns to a browser mainly when a page builds its content with JavaScript.

Can websites detect browser automation?

Websites can detect browser automation through more than one signal. A browser under WebDriver control reports that fact to each page through a standard property. Bot defenses add other signals, and some websites block automated browsers to stop bulk copying of content and fake accounts.