Skip to content

What is chaos engineering?

Chaos engineering is running controlled experiments that inject failures, e.g. a killed server, to learn how a system copes before a real outage happens.

Last updated , 8 min read

What is chaos engineering?

Chaos engineering is a reliability practice that breaks parts of a system in controlled experiments to find weaknesses before a real outage exposes them. It is also called chaos testing. The condition an experiment adds, e.g. a stopped server, is a fault, which need not be a defect in the code. Adding one on purpose is fault injection.

The International Software Testing Qualifications Board (ISTQB) defines chaos engineering as randomly injecting failures to gather information about how resilient a system is. It defines fault injection as creating adverse conditions to check whether a system detects them and still behaves reliably. Chaos Monkey, an open-source tool from Netflix, injects failures at random. It terminates virtual machines and containers in production.

The Principles of Chaos Engineering, written for large distributed systems, describe planned experiments instead of random failures. Each experiment tests a hypothesis about a measurable steady state. The Principles prefer experiments on production traffic, a form of testing in production, with a small blast radius.

Chaos experiments run error handling that ordinary tests seldom reach, e.g. a retry after a timeout. Code coverage often shows that code as never run. Adversarial testing and fuzzing attack a program through its input, while chaos engineering attacks what the program depends on.

How does a chaos experiment work?

A safe chaos experiment can run in six steps. Steps 3 and 6 add limits and cleanup to the four steps in the Principles:

  1. The team defines the steady state, a measurable output that shows normal behavior, e.g. the rate of completed checkouts.
  2. The team writes a hypothesis that the steady state will hold when one named fault happens, e.g. one app server stops.
  3. The team limits the blast radius, the users and systems the experiment can affect, e.g. by starting in a staging environment. It also sets a stop condition that ends the experiment early.
  4. A fault injection tool adds the fault to an experimental group, while a control group runs without it.
  5. Monitoring compares the steady state of the two groups, or of one group before and during the fault. If the stop condition triggers, the tool ends the experiment at once.
  6. The team removes the fault, restores the system, and records any difference as a weakness to fix.
The steps of a chaos experiment Steady state normal output Hypothesis it will hold Inject one fault small blast radius Compare with the control harm too high Stop condition end early, restore Record the result fix any weakness
The fault is the only change between the two groups, so a difference in the steady state points to a weakness. The stop condition limits the harm while the experiment runs.

A steady state that held raises confidence for that fault only. Teams often run their first experiments as a game day, a planned session in which a team injects faults at a set time and practices its incident response. The Principles recommend that experiments later run automatically and continuously.

Common faults fall into a few groups:

  • Stopped parts. A server or container stops, e.g. after docker stop.
  • Network faults. Calls to a dependency get added delay or no reply.
  • Error responses. A dependency returns errors or malformed answers.

What is an example of fault injection?

Here is an illustrative example. Acme Co. sells furniture online. A coding agent implemented the request "Let customers edit their delivery address during checkout." The change saves the address and then sends it to an address validation service. The agent's tests replace that service with a stub that answers at once, and they pass.

A developer at Acme runs an experiment in staging. A script there completes a checkout with 2 items every 10 seconds and edits the delivery address each time. The steady state is that each order keeps its 2 items. The hypothesis is that checkout still completes, with a notice that the address was not checked, when the address service stops answering.

The developer pauses the service's container, which suspends its processes, so requests to it get no reply:

docker pause acme-address

The next checkout waits for the service, and the agent's code times out after 5 seconds. Its error handler then restarts the checkout session. The address saved, but the cart emptied. The steady state breaks, so the hypothesis fails. The developer ends the experiment with docker unpause acme-address.

The agent changes the handler to keep the session and show the notice, and a rerun keeps the steady state. This example is simplified. A real experiment would also set a stop condition, e.g. ending it after 3 failed checkouts.

What changes when a coding agent writes the code?

A coding agent can write failure handling and its tests in the same session. Those tests often replace each dependency with a stub that fails in one clean way, e.g. by raising an error at once. A real dependency can also hang or answer late, and the error paths then run in a way that no test tried.

Retries show the gap. An agent can answer a failing call by adding a retry, and a test whose stub fails at once still passes. A delay fault shows the case that test missed. The first payment request succeeds late, the retry sends a second one, and the customer is charged twice. A retry is safe only when the operation has idempotency, which means running it twice gives the same result as running it once.

A practical adjustment is to write one small experiment for each dependency a change adds. The experiment stops or delays that dependency in an ephemeral environment and checks the steady state. For a command-line tool, a simple fault is an interrupt partway through a run. A tool that can handle Ctrl-C stops without leaving partly written files.

What are the limits of chaos engineering?

Chaos engineering shows how a system copes with the faults a team tried. Its main limits are:

  • Only the injected faults. An experiment says nothing about faults that nobody injected. No findings is not the same as complete coverage.
  • Staging is not production. Traffic and settings in staging differ from production, so a result there can miss how production behaves. An experiment in production can still harm customers, even with a small blast radius.
  • Timing. Some failures depend on the order of events, e.g. a race condition between a retry and a late reply, so one clean run can miss them.
  • Measurement. A steady state that counts only uptime can hold while results are wrong, e.g. orders that lose their items. A team needs monitoring of the right output before an experiment can show anything.
  • Cost. Each experiment needs preparation and a tested way to stop it early, by hand or through the tool.

How is chaos engineering different from load testing?

Load testing runs software under the traffic expected in production, to measure response times and find where it slows down. Chaos engineering usually keeps the traffic normal and asks what happens when one part of the system fails.

The two overlap. The Principles count a spike in traffic as one kind of event an experiment can add. A team can also inject a fault during a load test to see how the system copes at its peak.

FAQs

Should chaos experiments run outside production?

Chaos experiments should usually start outside production, e.g. in a staging environment, where a mistake cannot reach customers. The Principles of Chaos Engineering prefer production traffic, because staging differs from production. A team moves an experiment to production only with a small blast radius and a stop condition.

What is fault injection?

Fault injection is a technique that creates an adverse condition on purpose, e.g. a paused service, to check whether software detects it and still behaves reliably. Chaos engineering runs fault injection inside a planned experiment, with a hypothesis and a measured steady state.

Is chaos engineering only for large systems?

Chaos engineering also works on small systems, although its Principles were written for large distributed ones. A team with a small app and one outside service can still test what happens when that service stops answering, as the Acme checkout experiment did in staging.

What is a game day?

A game day is a planned session in which a team injects faults at a set time and practices its incident response. Teams often use game days for their first experiments, because people can watch the steady state and stop a fault early. Later experiments can run automatically and continuously.