What is the difference between performance, load, and stress testing?
Load testing and stress testing are two kinds of performance testing, which measures how fast and stable software is under a workload. A load test runs the software at the traffic its team expects, from the quietest hour to the busiest. A stress test goes past that traffic, or removes resources, to find where the software breaks.
Performance testing also covers checks with one user, e.g. timing a page load. All three are forms of non-functional software testing, which checks how well software works rather than what it does. Accessibility testing is another non-functional type.
A good load test and a good stress test also end differently. A load test that meets its targets is evidence that the software can serve the expected traffic. A stress test that ends in errors has done its job, because it shows where the limit is.
What is performance testing?
Performance testing is a test type that measures how fast, stable, and efficient software is under a given workload. The International Software Testing Qualifications Board (ISTQB) defines performance testing as a test type that determines performance efficiency, which is how well software uses time and resources.
A performance test usually records response time, throughput, error rate, and resource use. Response time is often reported as the median and a high percentile, e.g. the 95th, the time within which 95 of every 100 requests finish. Each result is judged against a target set in advance, e.g. a 95th percentile under 300 milliseconds. The software testing life cycle puts these targets in the test plan.
Two more kinds of performance testing differ in how the load changes over time. Spike testing sends a sudden burst of traffic, e.g. the first minute of a sale, and checks that the software recovers to its normal state. Soak testing, also called endurance testing, holds a steady load for hours to find problems that grow over time, e.g. a memory leak.
What is load testing?
Load testing is a performance test that runs software under the traffic expected in production, to measure response times and find where it slows down. The ISTQB defines load testing as testing behavior under varying loads, usually between low, typical, and peak usage.
A load test follows a load profile, which states how many virtual users act at once and what each one does. A load tool, e.g. the open-source Apache JMeter, runs the virtual users and sends their requests over HTTP or another protocol. JMeter does not render pages or run their JavaScript. Other open-source load tools define the virtual users in code, e.g. JavaScript.
Many load tests reuse the requests from API testing and send many copies at once. Requests that arrive together can expose a race condition, e.g. two requests that change one cart.
What is stress testing?
Stress testing is a performance test that runs software at or beyond the limits of its expected workload, or with fewer resources, to find how and where it fails. The ISTQB defines stress testing as testing at or beyond the limits of anticipated workloads, or with reduced resources, e.g. less memory.
A stress test shows three things:
- Breaking point. The breaking point is the load at which errors or timeouts start.
- Failure mode. The failure mode is how the software fails, e.g. with a "try again" message instead of a lost order.
- Recovery. Recovery is whether the software returns to normal when the load drops.
When should a team run each type?
Short performance checks can run on each change in a continuous integration and delivery (CI/CD) pipeline. Load tests belong before launch and before each expected peak. A change to a busy path, e.g. checkout, can need one too. Stress tests fit the times a team needs to know its margin.
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's code looks up the delivery price with one database query per cart item, the pattern known as the N+1 query problem, and its tests pass. The team tests the change in its staging environment:
- A test with one user sends one address change at a time. The median response is 90 milliseconds.
- A load test sends 50 address changes at once, the expected peak, with ApacheBench (
ab), an open-source command-line tool that repeats one HTTP request. - The 95th percentile is 2.4 seconds, against a target of 300 milliseconds. The queries for each cart item wait in line for database connections.
- The agent rewrites the lookup as one query, and the load test passes with a 95th percentile of 180 milliseconds.
- A stress test raises the load in steps, to 100, 200, and then 400 requests at once. At 400, some requests fail with
503, and checkout recovers when the load drops.
The load test in step 2 ran this command:
ab -n 1000 -c 50 -p address.json -T application/json \
https://staging.shop.example.com/api/checkout/address
ApacheBench printed this summary, shortened:
Concurrency Level: 50
Complete requests: 1000
Failed requests: 0
Requests per second: 21.40 [#/sec] (mean)
Time per request: 2336.449 [ms] (mean)
Percentage of the requests served within a certain time (ms)
50% 2210
95% 2410
100% 3050 (longest request)
A database branch can give load tests data at production size, and their writes stay out of the production database. This example is simplified. A real load test would follow whole checkout journeys, not repeat one request.
What changes when a coding agent writes the code?
A coding agent usually checks its work with the tests its session can run, which use one user and a small test database. Code that is slow only at scale passes them. A query per cart item costs little when the test cart holds 2 items. A demo has one user too, so it can reveal many bugs in AI-generated code but not slowness under heavy traffic.
An agent asked to fix a failing load test can also edit the test, e.g. by lowering the number of virtual users. The test then passes because it measures less.
A team can keep the load profile and its targets in files the agent does not edit. A test that counts the database queries for a cart of 20 items, run in the agent's session, can catch a query per item before the change reaches staging.
Can performance, load, and stress tests be used together?
The three types work together. A timing with one user sets the baseline, a load test checks the expected peak, and a stress test finds the margin above it. Chaos engineering can join a load test by stopping one part of the system, to see how the rest copes at the peak.
Together they still leave gaps:
- Staging is not production. Results carry over only as far as staging's servers and data match production's. Load tests usually run in staging anyway, because a load test on production can slow the service for real customers. Some teams also run small load tests on production at a quiet hour, with a way to stop them at once.
- The browser is left out. A load tool that sends HTTP requests does not time the page itself. An end-to-end test in a real browser can time it for one user.
- Only the modeled traffic runs. Customers who behave differently from the load profile can still reach a slow path.
How do performance, load, and stress testing compare?
The three types differ on these points:
| Point | Performance testing | Load testing | Stress testing |
|---|---|---|---|
| What it checks | Speed and stability under a workload | Targets at the expected traffic | How and where the software fails |
| Load applied | Any, from one user upward | Low to the expected peak | At or past the peak, or fewer resources |
| Main results | Response time, throughput, errors, and resource use | Response times and errors at the peak | Breaking point, failure mode, and recovery |
| Where it usually runs | A CI/CD pipeline or staging | Staging, with data at production size | Staging, not production |
| How often | Short checks on each change | Before launches and traffic peaks | When the margin is unknown |
FAQs
What is spike testing?
Spike testing is a kind of performance testing that sends a sudden burst of traffic, e.g. the first minute of a sale. It then checks that the software recovers and returns to its normal state.
What is soak testing?
Soak testing, also called endurance testing, is a kind of performance testing that holds a steady load on software for hours. It finds problems that grow over time and stay hidden in a short test, e.g. a memory leak.
Which open-source tools run load tests?
Apache JMeter and ApacheBench are two established open-source tools that run load tests. ApacheBench repeats one HTTP request, which suits a quick check of one endpoint. JMeter runs virtual users through a load profile, which suits whole journeys, and other open-source tools define those users in code.
Should load tests run against production?
Load tests usually run against a staging environment, because heavy traffic on production can slow the service for real customers. Teams that test on production usually keep the test small, run it at a quiet hour, and have a way to stop it at once.