What is test automation?
Test automation, also called automated testing, is the use of software to run tests and compare each actual result with an expected result. The steps and expected results live in code, so the same tests can repeat on each change. Manual testing is the alternative, in which a person performs each step and judges the result.
The International Software Testing Qualifications Board (ISTQB) defines test automation as "the conversion of test activities to automated operation." Under that definition, a script that loads test data is test automation too.
Test automation is a way to carry out software testing, not a separate kind of test. It makes a check cheap to repeat, which lets a team test each change through continuous integration and delivery (CI/CD). Someone still has to choose what to check.
How does test automation work?
An automated test run usually follows six steps:
- A developer or a tester writes each test as code, with its steps and expected result.
- After a code change, a developer or a CI service starts a run.
- The test runner sets up the software and its test data.
- Each test acts on the software through an interface, e.g. a function call, and records the actual result.
- Each test compares the actual result with its expected result and passes or fails.
- The runner reports the results, and one failed test fails the whole run.
A test automation framework is the set of libraries and tools that a team writes and runs its tests with. In the ISTQB glossary, it is a set of test harnesses and test libraries. Python projects often use pytest, and web apps often use a browser test framework, e.g. Playwright.
Fast tests usually run on each change, e.g. in a GitHub Actions workflow. Slow tests often run on a schedule or before a release.
What is an example of test automation?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change goes through five steps:
- The agent edits the checkout code and adds its own tests, which pass.
- A tester at Acme follows the manual test case below. The address saved, but the cart emptied.
- The agent fixes the session handling, and the manual steps pass.
- A developer turns the manual case into the browser test below, and Acme's CI service runs it on each change.
- Later, the agent changes how the cart stores items for another request. The browser test fails on that change, because the cart count shows
0 items, and the agent fixes the cart before the change merges.
The manual test case reads:
Test case CO-7: edit the delivery address
1. Add 2 oak chairs to the cart.
2. Open checkout and save the address "12 Elm Street".
3. Expected: the page shows the new address and 2 items in the cart.
The browser test runs the same steps from code:
import { test, expect } from "@playwright/test";
test("CO-7: edit the delivery address", async ({ page }) => {
// baseURL in playwright.config.js points at the build under test
await page.goto("/products/oak-chair");
await page.getByLabel("Quantity").fill("2");
await page.getByRole("button", { name: "Add to cart" }).click();
await page.getByRole("link", { name: "Checkout" }).click();
await page.getByLabel("Delivery address").fill("12 Elm Street");
await page.getByRole("button", { name: "Save address" }).click();
await expect(page.getByText("12 Elm Street")).toBeVisible();
await expect(page.getByTestId("cart-count")).toHaveText("2 items");
});
Each part of the manual expected result becomes its own check, and the cart-count check fails in step 5. This example is simplified. A real checkout suite would also check the order total.
Which tests are worth automating?
A test is usually worth automating when it will run many times and its expected result can be written down exactly. Good candidates have these traits:
- Repeated runs. A check that runs on each change pays back its cost soonest, which is why regression tests are strong candidates for automation.
- Exact results. Code can compare an item count, but it cannot judge whether a page is confusing.
- Stable behavior. A screen that changes each week can break its tests each week.
- High cost of failure. A journey that loses sales when it breaks, e.g. checkout, earns an automated check.
Automate each check at the lowest level that can catch its bug. A unit test checks one function without starting the app. An end-to-end test through the user interface (UI) runs slowly and can break when the page changes, so teams keep those tests to a few critical journeys.
Manual and exploratory testing fit better for checks that run once or need a person's judgment. Upkeep of automated tests usually falls on the developers who change the code, sometimes with test automation engineers who own the framework.
What changes when a coding agent writes the code?
A coding agent makes automated tests cheap to write. It can write a test file with each change, run the suite, and read the failures before it reports "done." In Stack Overflow's 2025 survey of about 49,000 developers, 84% said they use or plan to use AI tools, and about half of professional developers use them daily.
Writing and fixing scripts used to limit what a team automated. With an agent, the cost moves to choosing what the tests check. The agent writes its tests from the same prompt and code as the change, so a gap in the change usually reappears in the tests. When a test fails, the agent can also edit the expected result instead of the code. Agents that run the software and choose and judge the checks themselves are doing agentic testing.
A practical adjustment is to take expected results from test cases a person wrote from the request, e.g. CO-7 in the Acme example. The agent can script each case, and a reviewer checks each edited test against its case.
Does automating manual tests give good tests?
Automating manual tests step for step usually gives slow, brittle tests that check less than the tester did. A person fills the gaps in a manual case with judgment, and a script does only what its code says. Four things get lost:
- Unwritten checks. A tester also notices a wrong price. The script compares only the values in its code, so its test oracle is narrower than the tester's.
- Independent steps. A long manual case run as one script stops at its first failure, so automated tests are kept short and set up their own data.
- Timing. A person waits for the page. A script that waits a fixed time fails when the page is slow, which makes it a flaky test.
- Upkeep. A script that finds a button by its position breaks when the layout changes. A self-healing test updates how it finds the button on its own, which cuts upkeep but can hide real breakage.
A passing suite shows that its checks passed, not that the software works.
How is automated testing different from manual testing?
Automated testing moves the steps and the judgment from a person into code. That changes what each run costs and what it checks:
| Question | Manual testing | Automated testing |
|---|---|---|
| Who runs the steps | A person | A test runner |
| Cost of one more run | A person's time | Machine time once the test exists |
| What it checks | Whatever the tester notices | Only the results its code compares |
| Repeatability | Varies with the tester | The same steps on each run |
| Best for | Unfinished features, usability, and exploration | Known checks on each change |
| Main cost | Time on each run | Writing and maintaining the scripts |
Most teams use both. Automation repeats the known checks, and people look for problems nobody has scripted yet, often as part of quality assurance (QA) testing.
What should an automated check look at besides the tests?
Besides the test files, an automated check can look at the running software and the original request. Run the changed workflows from start to finish on a clean setup, and compare the results with what was asked. A check written from the request, not from the code, can catch gaps that the agent's tests share with its code.
RunStory considers your prompts and code changes to decide what to test. A testing agent uses your software in a separate environment, tries relevant workflows, and checks the results while you keep working. RunStory is in private alpha for CLIs and web apps.
FAQs
What is a test automation framework?
A test automation framework is the set of libraries and tools that a team uses to write and run automated tests. The ISTQB defines it as a set of test harnesses and test libraries. Python projects often use pytest, and web apps often use a browser test framework.
What problem does automated UI testing solve?
Automated UI testing removes the need to repeat known checks through the interface by hand after each change. UI tests run slowly and can break when the page changes, so teams usually keep them to a few critical journeys.
Is manual QA coming back because of AI?
Manual QA keeps a clear role when coding agents write the code, because an agent's own automated tests usually check only what the agent built. People explore the changed workflows and judge what no script can check, e.g. whether a page is confusing. Automation still repeats the known checks.
Who maintains automated tests?
Automated tests are usually maintained by the developers who change the code they check, sometimes with test automation engineers who own the framework. When a coding agent changes the code, it can edit the tests too, so a reviewer should check each edited test.