Skip to content

What is exploratory testing, and can an AI agent do it?

Exploratory testing is designing and running tests at the same time, guided by what the tester learns about the software as they use it.

Last updated , 8 min read

What is exploratory testing, and can an AI agent do it?

Exploratory testing is a testing approach in which a tester learns about the software, designs tests, and runs them at the same time. No steps are scripted in advance. An AI agent can do part of this work, because it can operate software and record its actions. It needs people to state what the software should do.

The International Software Testing Qualifications Board (ISTQB) defines exploratory testing as designing and running tests on the fly, based on the tester's knowledge, exploration of the software, and earlier results. ISTQB does not require sessions or charters, but many practitioners use both, in a method called session-based test management.

A test charter is a short mission statement for one session that names the area to explore, the risks to look for, and a time box. A common form is Explore TARGET with RESOURCES to discover INFORMATION.

Exploratory testing is one approach within software testing. In scripted testing, someone writes the steps and expected results before any test runs.

How does exploratory testing work?

An exploratory session usually runs in six steps:

  1. The tester starts from a charter that names one area and one risk, narrow enough for one session.
  2. The tester uses the software in that area and tries one action.
  3. The tester compares the result with a test oracle, the source of what should happen, e.g. the original request.
  4. The tester designs the next action from that result, and steps 2 to 4 repeat until the time box ends.
  5. The tester takes notes, including the exact reproduction steps behind each problem.
  6. A short debrief sorts the findings into bugs, questions for the feature's owner, and areas that need another session.
Scripted testing and an exploratory session Scripted testing Write the steps before any run Run the steps same path each run Compare results pass or fail Exploratory session Charter area, risk, time Design a test Run it Read the result surprise or not Notes and findings the result shapes the next test
A script fixes its steps before the first run. An exploratory session designs each test from the last result, inside the limits of its charter.

Step 4 is how exploratory testing finds bugs that scripted tests miss. A scripted test, e.g. an end-to-end test, takes the same path on each run and catches failures only on paths someone planned. An exploratory tester follows a surprise into states that no script reaches. Useful findings then become scripted checks for regression testing.

A bug bash is a group form of exploratory testing, with people from across a team on the same build.

What is an example of exploratory testing?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes the checkout code, and its tests pass. A tester at Acme runs one session with this charter:

Explore the delivery address edit in checkout,
with carts of different sizes, to find ways
the change loses order data.
Time box: 60 minutes.

The session and the fix run in six steps:

  1. The tester adds 2 items to the cart and saves the address "12 Elm Street." The page shows the address and 2 items.
  2. The checkout URL changes from /checkout/s-81 to /checkout/s-82, which suggests that the save starts a fresh session.
  3. That clue leads the tester to reload the page. The address saved, but the cart emptied.
  4. The tester repeats the steps with 1 item and with 5 items. The cart empties both times, so its size is not the cause.
  5. The tester files a bug report with the steps, the expected 2 items, and the actual 0 items.
  6. The coding agent changes the session handling. On a retest the cart keeps its 2 items, and the steps become a regression test.

The agent's tests checked the page right after the save and passed. The reload exposed the failure. This example is simplified. A real session would also explore the payment step.

Can an AI agent do exploratory testing?

An AI agent can run an exploratory session when it can operate the software and read the results. It drives a web app through a browser, often a headless browser, and a command-line tool through a terminal. It generates each next action from the output of the last one. Testing that AI agents carry out themselves is often called agentic testing.

Given a charter, an agent can explore these parts of a change:

  • Workflows. The agent tries a workflow in different orders and states, e.g. a reload right after a save.
  • Input. The agent generates unusual input faster than a person can type it, e.g. an address with a line break.
  • Records. The agent logs each action it took, so a finding comes with the steps that led to it.

An agent's testing is often shallow. In Dan Luu's 2026 experiment, agents told to fuzz a Zstd decoder built useful structured inputs in only 10 of 160 runs. About half of those found real bugs. Agents told to use named techniques mostly did not beat the default prompt. A charter points the agent at one area and one risk, but Luu's experiment did not test charters.

Two parts stay hard. Apart from crashes, a surprise is a bug only when it breaks what the software should do, so the agent needs that intent in writing, e.g. the request. Whether customers find a page confusing is still a question for a person.

What changes when a coding agent writes the code?

A coding agent's own tests usually check only the path its prompt described, and they pass. The bugs that remain sit where the prompt said nothing, e.g. a reload in the middle of checkout. Many bugs in AI-generated code are of this kind.

The agent can explore its own change, but it works from the same prompt and context, so it tends to skip the same cases. The diff shows where else to look. In the Acme example, the request named only the address, but the change also started a fresh checkout session on each save.

A practical adjustment is a short session on each change that touches a user-facing workflow, run by a person or a separate agent session. Its charter comes from the request and the diff, not from the coding agent's summary. That session is a form of runtime verification, because it checks the change by running the software.

What are the limits of exploratory testing?

Exploratory testing finds problems, but it leaves real gaps:

  • Hard to repeat. A session depends on the tester and the path taken, so a second session rarely follows the same steps. Findings need written steps before anyone can check them again.
  • Unknown coverage. A session samples a small part of the states the software can reach. No findings is not the same as complete coverage.
  • Depends on the tester. Results vary with the tester's knowledge of the software and with the charter.
  • Needs an oracle. Without the requirements, a tester can report a surprise but cannot always say whether it is a bug.

How is exploratory testing different from monkey testing?

Monkey testing feeds software random clicks, keystrokes, or input without a script. It does not use one result to choose the next action, and it checks only for crashes and hangs, so it rarely finds a wrong total. Exploratory testing chooses each test from what the tester learned, and it compares each result with what should happen. Monkey testing needs no judgment, so a tool can run it without a person.

Ad hoc testing sits between the two, with a person trying the software without a script, a charter, or notes. Adversarial testing shares the aim of breaking the software and can use either technique.

How does RunStory help with exploratory testing?

Exploratory testing can find what the scripts missed, but only when someone uses the software after the change. RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Can exploratory testing replace manual functional tests?

Exploratory testing cannot fully replace manual functional tests, because scripted tests repeat known checks the same way on each run and a session does not. Exploratory sessions find problems nobody predicted, and the useful ones become scripted tests.

How do you write a test charter?

A test charter is one sentence that names the area to explore, the resources to use, and the risk to look for. An example is "Explore the delivery address edit, with carts of different sizes, to find ways it loses order data." Add a time box that fits one session.

Are exploratory findings bugs or design choices?

Exploratory findings can be bugs or behavior that nobody specified. A finding is a bug when it breaks what the software should do, e.g. a cart that empties after an address change. When no requirement covers it, the tester reports it as a question for the feature's owner.

How do you test for unknown issues?

Testing for unknown issues means using the software in ways no script covers and following each surprise until it is narrowed down. Exploratory sessions aimed at risky areas do this. Adversarial testing and monkey testing add input that nobody planned.