Key points
- Dynamic testing runs the software, and static testing, e.g. a code review, examines it without running it.
- Tests are grouped by level, from one function to the whole system, and by type.
- When a coding agent writes the code and the tests, keep one check outside its work.
What is software testing?
Software testing is the practice of checking what a program does against what it should do, by running it or by examining it. Its goals are to find defects and to give evidence about whether the software is ready to release. A test that runs the program fails when the actual result differs from the expected result.
The International Software Testing Qualifications Board (ISTQB) separates two approaches. Dynamic testing runs the software, and static testing examines it without running it, e.g. a code review.
Testing can check that software was built as specified and that it meets its users' needs. Those are the two questions behind verification and validation. This learning center uses "verification" in a wider sense, for any check that produces evidence about software, including running it.
How does software testing work?
An automated test checks one behavior in five steps:
- A developer writes a test case that names an input, the steps to take, and the expected result.
- The test runner sets up the software and the data the test needs.
- The test runner runs the steps and records the actual result.
- The test compares the actual result with the expected result.
- The test passes if the two match. If they differ, the test fails. A person or a coding agent then debugs the cause.
Running a program is a test only when something checks the result against an expectation. A crash is one such check, and a developer watching the screen is another. An expected result written in the test makes the check repeatable. Static testing skips the run. A reviewer or a tool reads the code and compares it with rules or requirements, e.g. a type checker that flags text passed where a number belongs.
What are the types of software testing?
A test type groups tests by the quality they check, by how they are designed, or by how they run. Types and levels describe different things, so one test has both, e.g. a functional test at the unit level. The first three groups below cover the four test types that the ISTQB syllabus names, and it notes that each type can run at every level. The last two groups sort tests by when they run and how they run.
Functional testing
Functional testing checks what the software does against what it should do, e.g. that a discount code lowers the order total. It answers whether each feature gives the right result for its inputs.
Non-functional testing
Non-functional testing checks how well the software behaves, e.g. how long the checkout page takes to load. Performance testing and security testing are two common kinds.
Black-box and white-box testing
Black-box testing designs tests from the requirements without reading the code. White-box testing designs tests from the structure of the code, e.g. so that both branches of an if statement run. The two terms describe how a tester chooses test cases, not what the software does.
Regression and confirmation testing
Regression testing checks that a change has not broken behavior that worked before. Confirmation testing, also called retesting, repeats a failed test after a fix to check that the failure is gone. The ISTQB treats both as testing that follows a change.
Manual and automated testing
A person runs a manual test by hand, and a test runner executes an automated test from code, e.g. a pytest function. Automated tests are cheap to repeat, so teams usually run them on each change. Many teams shape their automated tests as a test pyramid, with many fast unit tests and a few slow end-to-end tests.
What are the levels of software testing?
A test level is a stage of testing with its own scope and goal. The levels run from one small part of a program to the whole system in use. The common levels, from smallest to largest, are:
- Unit testing. A unit test checks the smallest testable part of a program, e.g. one function, in isolation. The ISTQB calls this level component testing.
- Integration testing. An integration test checks that separate parts work together, e.g. a service and its database.
- System testing. System testing checks the complete, integrated system against its requirements.
- Acceptance testing. Acceptance testing checks whether the system meets its users' needs, so that its users and owners can accept it.
The ISTQB splits integration testing into component integration and system integration, which makes five levels. End-to-end tests usually sit at the system or acceptance level, because they follow a whole user journey through the running software. Many developers use a shorter set of three levels, which are unit, integration, and end-to-end tests. Developers usually run unit tests on each change. The slower levels often run in continuous integration (CI) or before a release.
What is an example of software testing?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change goes through three kinds of testing, a unit test, a code review, and an end-to-end test:
- The agent edits the checkout code and adds a unit test for
order.setAddress. The unit test passes. - The full unit suite runs in CI, and 214 tests pass.
- A reviewer reads the diff and approves it. The review is static testing and finds nothing wrong.
- The developer writes an end-to-end test from the request. It adds 2 items to the cart, starts checkout, and sets the address to "12 Elm Street."
- The test expects 2 items on the review page and finds 0. The address saved, but the cart emptied.
- The developer sends the failing steps to the agent, and the agent edits the code so the cart keeps its items.
- The end-to-end test runs again and passes. It stays in the suite as a regression test.
The failing run in step 5 printed this:
1) checkout.spec.ts:3:5 › customer edits delivery address
Error: expect(locator).toHaveCount(expected) failed
Locator: getByTestId('cart-item')
Expected: 2
Received: 0
Each kind of testing checked something different. The unit test checked the address alone, and the review read the code without running it. Only the test that ran the whole checkout found the empty cart. This example is simplified. A real project would add more checks, e.g. one for the order total.
What changes when a coding agent writes the code?
The ISTQB syllabus recommends several levels of independence in testing for most projects, so the author of the code is not its only tester. Testers who did not write the code tend to find different kinds of failures. The syllabus's example splits the work by level, with developers on unit tests and business representatives on acceptance tests.
A coding agent can write the code and all of its tests in one session, from one prompt. The agent is then both the author and the tester at each level it covers. If the code misses part of the request, the tests usually miss the same part, so they pass. In the Acme example, the agent's unit test checked the address and never checked the cart.
The practical adjustment is to give at least one level a different author. In the Acme example, that was the end-to-end test the developer wrote from the request, which found the empty cart.
What are the limits of software testing?
Testing gives evidence about software, not proof that it is correct. Its main limits are:
- Tests show that bugs exist, not that they are absent. A passing test shows that the inputs that ran gave the expected results. It says nothing about the inputs that did not run.
- Most programs accept too many inputs to test them all. Teams choose a sample, so a bug can hide in a case nobody picked.
- A test is only as good as its expected result. A test with a weak or missing test oracle, the source of its expected result, can pass on broken code. High code coverage does not fix this, because coverage counts lines that ran, not checks that were made.
- Tests cost time to write, run, and maintain. A slow or unreliable suite often gets skipped, and a skipped test checks nothing.
No findings is not the same as complete coverage. Teams therefore agree in advance on when to stop, based on risk. The conditions they agree on are called exit criteria, e.g. that the planned tests have run and no severe defect is open.
How is software testing different from quality assurance?
Quality assurance (QA) improves the process a team uses to build software, so that fewer defects are made. Testing examines the product and finds the defects already in it, and the ISTQB calls it a form of quality control.
What does software testing cover?
The articles in this category explain these techniques and comparisons:
- Unit testing checks one small part of a program in isolation.
- Integration and end-to-end tests run more of the system than a unit test.
- End-to-end testing checks a whole user journey through the running software.
- Smoke testing checks quickly that a build starts.
- Regression testing checks that behavior that worked before still works.
- Exploratory testing designs and runs tests while using the software.
- Test-driven development writes a failing test before the code.
- Mocks, stubs, and fakes replace real parts of a system during a test.
- Integration testing checks that connected parts of a system work together.
- Acceptance testing checks software against the needs of its users and owners.
- The test pyramid models a suite with fewer tests at each higher level.
- QA testing is the testing a quality assurance team does to find defects.
- Test automation uses software to run tests and compare their results.
- A test harness is the code that runs tests on a system and checks the results.
- Snapshot testing fails a test when output no longer matches a recorded copy.
- User journey tests follow one user task from start to finish.
- Agentic testing has an AI agent run the software and judge each result.
- Pre-launch testing runs the tasks customers need on day one on a clean copy of the app.
How do you test software a coding agent wrote?
Test it at more than one level, as with code a person wrote, and have someone other than the agent write at least one of the checks. Then run the finished software, not only its tests. The page on how to test AI-generated code lists the steps.
RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures to your coding agent and verifies the fix. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
Does running the code count as testing?
Running the code counts as testing when someone or something checks the result against an expectation. A developer who starts the app and watches it is running a manual test with an unwritten expected result. Writing down the steps and the expected result makes that test repeatable.
What kind of testing should a developer do before handing off a build?
A developer handing off a build should at least run the unit tests and a smoke test, to check that the build starts and its basic functions work. Checking the changed feature and running the regression tests catches more failures before a tester or reviewer spends time on the build.
Can testing prove software is bug-free?
Testing cannot prove that software is free of bugs, because a suite runs only a sample of the possible inputs. A bug can sit in an input nobody picked, or behind a weak expected result. Each added level or type of testing gives more evidence, not proof.
When do you stop testing?
Testing stops when the remaining risk is low enough to release, not when the bugs run out. Teams agree on exit criteria before testing starts, and testing ends when the planned tests are done with no severe defect left open.