Skip to content

What is agentic testing?

Agentic testing is software testing done by AI agents that choose, run, and judge checks themselves, not agents that only write test scripts for later.

Last updated , 8 min read

What is agentic testing?

Agentic testing is a testing approach in which an AI agent runs the software, chooses each check, and judges each result itself. It is also called agentic quality assurance (QA), and the agent is often called an AI QA agent. Under this definition, an agent that only writes test scripts for a later run does test generation instead.

The term has no standard definition, and test tool vendors also use it for AI features that write or repair test scripts. In the narrower sense above, the agent acts on the running software, so it can reach states that no script covers.

Agentic testing is one approach within software testing. It differs from classic test automation, in which a person writes the steps and expected results in advance and a test runner repeats them on each change.

How does agentic testing work?

An agentic test run usually follows six steps:

  1. A person or a pipeline gives the testing agent a goal, e.g. the request behind a change.
  2. The software starts in a sandbox with test data, not real customer data.
  3. The agent reads the current state through a tool, e.g. a snapshot of a web page.
  4. A large language model (LLM) returns the next action from the goal and that state, and the tool performs it.
  5. The agent compares the result with the expected behavior and records the action and its evidence.
  6. Steps 3 to 5 repeat until the goal is covered or a step budget runs out. The agent then reports each finding with the reproduction steps it recorded.
An agentic test run with a deterministic check Goal the request Testing agent chooses each action Browser or shell acts on the software Software in a sandbox state after each action Findings steps and evidence Deterministic check exit code or saved test same answer each run
The dashed line closes the loop, in which each action depends on the state the last action left. The deterministic check sits outside that loop and reads the software directly.

Because each action follows from the last result, agentic testing is often a form of exploratory testing.

The agent acts through the pages a user sees, as an end-to-end test does, or through a shell for a command-line tool. A browser tool for agents, e.g. Playwright MCP, returns a structured snapshot of the page by default, not a screenshot. Its project README says no vision model is needed.

A deterministic check, e.g. a command's exit code, gives the same answer for the same state, while a judgment by the model can differ between runs.

What is an example of agentic testing?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's unit tests pass.

The developer then gives a testing agent a goal from the request, which is that editing the address must not lose order data. The example has six steps:

  1. The agent opens the store in a sandbox, adds 2 items to the cart, and saves "12 Elm Street" as the delivery address.
  2. The goal names order data, so the agent reloads the checkout page to check the items.
  3. After the reload, the cart shows 0 items.
  4. A second run with 1 item also ends with 0 items, so a misclick is unlikely.
  5. The agent reports the finding with its recorded actions, the expected 2 items, and the actual 0 items.
  6. The coding agent fixes the bug, and a developer turns the recorded actions into a Playwright test.

The first run recorded these actions:

goto    https://shop.example.com/products/oak-chair
fill    Quantity = 2
click   "Add to cart"      -> cart-count "2 items"
click   "Checkout"
fill    Delivery address = "12 Elm Street"
click   "Save address"     -> address "12 Elm Street"
reload  /checkout          -> address "12 Elm Street", cart-count "0 items"

The reload that exposed the bug came from the goal, not from a script. This example is simplified. A real run would also try the other journeys the change touches, e.g. payment.

Is agentic testing the same as testing AI agents?

Agentic testing is not the same as testing AI agents, because it uses an agent to test software, e.g. a web store. Testing an AI agent checks the agent itself, e.g. a support chatbot. Questions about how agentic workflows are tested usually mean this second job.

Teams test an AI agent with evals, which run a fixed set of tasks and score the results. The scoring often uses LLM-as-a-judge, in which a language model grades each output against written criteria.

A testing agent is an AI agent too, so a team can run an eval of it on a copy of the app with planted bugs. The bugs it misses show where its checks are weak.

What changes when a coding agent writes the code?

A coding agent with a browser or terminal tool can test its own change before it reports "done." That run starts from the same prompt and context as the code. A goal the coding agent writes for itself can describe what it built instead of what was asked, e.g. "the address form saves."

Naming a test technique in the prompt may not help. In Dan Luu's 2026 experiment, agents that each wrote a Zstd decoder and were told to fuzz it built useful structured inputs in only 10 of 160 runs. About half of those found real bugs. Agents told to use named techniques mostly did not beat the default prompt. The experiment did not show that a separate testing agent does better.

A practical adjustment is to write the testing goal from the request, including what must not change, e.g. the cart. The team then reads the run's recorded actions and evidence, not the agent's summary.

What are the limits of agentic testing?

Agentic testing can reach states that scripts miss, but it has limits:

  • The model can be wrong. It can accept a wrong result. On the next run it can take a different path or judge the result differently, which can make the check a flaky test.
  • The agent can cause the failure. A wrong click can pass for a bug in the software, so each finding needs steps that someone can follow again on the same build and data.
  • Judgment needs written intent. Apart from crashes and error codes, the agent has to guess what the right result is unless the expected behavior is written down.
  • Each run costs more. Each step sends the page state to the model, so an agent-run check is slower and costs more than a scripted test.
  • Coverage is unknown. No findings is not the same as complete coverage.

For known behavior that must stay true, a deterministic check, e.g. a committed test, should set the pass or fail result. The agent is better used to find failures nobody scripted.

How are agent-run tests different from agent-written tests?

An agent-written test is a script, e.g. a Playwright test file, that an agent writes once and a test runner repeats. An agent-run test has no fixed script, because the agent chooses each step during the run. The two differ on these points:

QuestionAgent-run testAgent-written test
Who acts during the runThe agentA test runner
Same steps each runNoYes
Cost of each runA model call per stepRunner time only
What it findsFailures off the scripted pathsBreaks on the scripted paths
Main weaknessResults vary between runsChecks only what the script names

A self-healing test mixes the two, because a tool repairs a scripted step when the app changes, which can hide a real break. One workflow uses both. An agent explores a change, and a person commits a test for each path that must keep working. The person checks its assertions first, because a test written from what the software did can lock in a bug.

How does RunStory help with agentic testing?

An agentic test run needs a goal from the request and a sandbox to run in. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps. Your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Is agentic QA the same as test automation?

Agentic QA differs from classic test automation, because the agent chooses its steps during the run instead of following a script written in advance. Each step needs a model call, so teams keep scripted tests for the checks that repeat on each change.

How do you repeat a check an agent ran?

To repeat a check an agent ran, start from the actions it recorded, not from its summary. Follow those steps on the same build and data. When the failure repeats, turn the steps into a committed test so later runs check it the same way.

Should a team commit tests an agent explored?

A team should commit a test for a path an agent explored when that path must keep working, e.g. checkout. A person checks the test's assertions against the request first, because a test written from what the software did can lock in a bug.

Is agentic testing deterministic?

Agentic testing is not deterministic, because the model behind the agent can choose a different path or judge a result differently on each run. Teams get repeatable results by pairing the agent with deterministic checks, e.g. a committed test, for behavior they already know.