Skip to content

What is fuzzing?

Fuzzing is feeding a program large numbers of random or malformed inputs to find the crashes, hangs, and input bugs that ordinary tests miss.

Last updated , 9 min read

What is fuzzing?

Fuzzing is a testing technique that runs a program on many random or malformed inputs to find crashes, hangs, and bugs in input handling. It is also called fuzz testing, and the tool that makes the inputs is a fuzzer. Fuzzing needs no expected result for each input, because a crash or a hang is a failure.

The International Software Testing Qualifications Board (ISTQB) defines fuzz testing as a technique that uses high volumes of random data to generate test inputs. In a 1990 study, Barton Miller and colleagues built a tool called fuzz that sent random characters to UNIX utilities, and the input crashed or hung many of them. Their tool used no model of how the programs should behave. Many fuzzers are guided instead, by code coverage or by the input format, so the purely random definition fits only part of the practice.

Ordinary tests send the inputs their author thought of. A fuzzer also sends inputs that nobody planned for, so code that a suite covers fully can still crash on an input that no test tried. Fuzzing is one technique of adversarial testing. Security teams also use it to find input bugs that attackers could exploit.

How does fuzzing work?

A fuzz harness, also called a fuzz target, is a small test harness that takes an input from the fuzzer and passes it to the code under test. A fuzzing run repeats five steps:

  1. The fuzzer takes an input from its corpus, a set of saved inputs that usually starts with valid seeds.
  2. The fuzzer changes the input, e.g. by deleting some of its bytes. This is mutation-based fuzzing. Generation-based fuzzers build inputs from a description of the format instead.
  3. The harness runs the code with the changed input.
  4. The fuzzer checks the run for a crash, a hang, or a failed check in the harness.
  5. If the run passes, the loop repeats. On a failure, the fuzzer saves the input, and some fuzzers reduce it to a smaller input that still fails. With the harness, that input forms a minimal reproducible example of the bug.

A crash or a hang is an implicit test oracle. The fuzzer detects a hang with a time limit on each input.

Good fuzz targets take input from outside the program:

  • Parsers and request handlers. Code that reads a file format or an API request turns untrusted bytes into structure.
  • Command-line arguments. For a command-line application, a harness can split the fuzzer's bytes into arguments and call the tool's argument parser directly. Calling code in the same process is faster than starting the program for each input.
  • Round trips. A pair of functions that undo each other, e.g. encode and decode, can be fuzzed with a check that decoding returns the original.

What is an example of fuzzing?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout."

The agent adds a Go function, ParseAddress, that splits "12 Elm Street" into a house number and a street. Its 2 tests pass, and go test -cover reports full statement coverage.

The developer adds a fuzz harness with one seed. Go's fuzzer changes the seed and passes each result to the function inside f.Fuzz:

func FuzzParseAddress(f *testing.F) {
	f.Add("12 Elm Street")
	f.Fuzz(func(t *testing.T, s string) {
		addr, err := ParseAddress(s)
		if err == nil && addr.Street == "" {
			t.Errorf("ParseAddress(%q) returned no street", s)
		}
	})
}

Besides crashes, the harness checks one rule. An accepted address must have a street. The developer starts the fuzzer, and the run fails, shown here without progress lines and the stack trace:

$ go test -fuzz=FuzzParseAddress
fuzz: minimizing 38-byte failing input file
--- FAIL: FuzzParseAddress (0.06s)
    --- FAIL: FuzzParseAddress (0.00s)
        testing.go:1927: panic: runtime error: index out of range [1] with length 1

    Failing input written to testdata/fuzz/FuzzParseAddress/771e938e4458e983
    To re-run:
    go test -run=FuzzParseAddress/771e938e4458e983
FAIL

Go saves the failing input after reducing it to one character:

$ cat testdata/fuzz/FuzzParseAddress/771e938e4458e983
go test fuzz v1
string("0")

ParseAddress splits the text at the first space and reads the second part, parts[1]. A house number with no street, e.g. 12, has no second part, so the code panics. The passing tests ran that line too, but only with an address that contains a space.

The agent changes the function to return an error when there is no street. Go keeps the saved input in testdata/fuzz, so a plain go test reruns it as a regression test. This example is simplified. A real project would check more rules, e.g. that a parsed address formats back to the original text.

What is coverage-guided fuzzing?

Coverage-guided fuzzing is a fuzzing method that records which code each input reaches and keeps the inputs that reach code no earlier input reached. The fuzzer then mutates those inputs further.

Coverage guidance adds two steps to the fuzzing loop. Before fuzzing starts, the compiler usually adds code to the program that records the branches each run takes. After each run, the fuzzer adds the input to the corpus if the run took a branch that no earlier run took.

The coverage-guided fuzzing loop Corpus saved inputs Mutate change one input Run the harness record branches fails Save the input and report it Reached more code? no, drop the input yes, keep it in the corpus
Coverage is the feedback. An input that reaches more code becomes a starting point for later mutations, and an input that fails leaves the loop as a report.

Random bytes rarely get past the first validation check. Coverage feedback keeps each input that gets one branch further, so later mutations start from there. Go's built-in fuzzing works this way, and so do libFuzzer and AFL++ for C and C++ programs.

What changes when a coding agent writes the code?

A coding agent can write the code and its fuzz harness in the same session. Whether the fuzzer finds anything then depends on seeds and checks that the same agent wrote.

In Dan Luu's 2026 experiment, agents that each wrote a Zstd decoder in Rust and were told to fuzz it built useful structured inputs in only 10 of 160 runs. About half of those found real bugs. In the other runs, the agents sent random bytes or random changes to valid inputs, which their own decoders mostly rejected as invalid. Agents told to use named techniques mostly did not beat the default prompt.

A practical adjustment is to review the harness as closely as the code. A reviewer can check that it starts from valid seeds, reaches code past the input validation, and checks at least one rule about results. An agent's fuzz run belongs in a sandbox, because it runs agent-written code unattended for a long time, and that code can write files or reach the network.

What are the limits of fuzzing?

A fuzz run that finds nothing shows only that the inputs it tried did not fail. Fuzzing has these limits:

  • A fuzzer finds only failures it can detect. Without a check in the harness, a wrong result that does not crash passes. C and C++ fuzzers usually run with sanitizers, e.g. AddressSanitizer, which turn silent memory errors into crashes.
  • Deep code is hard to reach. A strict check, e.g. a checksum, stops most changed inputs before the code behind it, and code that no input reached gets no check. A harness can recompute the checksum for each input, or the team can build the program with the check turned off for fuzzing. Static application security testing (SAST) can still flag code that no input reached.
  • Timing bugs rarely appear. A fuzzer varies the input, not the order in which threads run, so a fuzz run seldom shows a race condition.
  • A run has no natural end. By default, go test -fuzz runs until a failure or until someone stops it, so teams set a limit with -fuzztime.

How is fuzzing different from random testing?

ISTQB defines random testing as a technique that generates input values at random and treats the program as a black box. It keeps fuzz testing as a separate term, although the 1990 fuzz paper calls its own method random testing.

In practice, random testing often samples valid inputs and checks each result against an expected result. Fuzzing aims at malformed and unexpected input, and it usually treats a crash, a hang, or a failed check as the failure. Coverage-guided fuzzers also read the program's coverage, so they do not treat it as a black box.

Property-based testing also generates inputs, often at random. It checks each result against a rule that a person writes.

FAQs

Is fuzzing effective in practice?

Fuzzing is effective at finding crashes and hangs in code that reads outside input, and the original fuzz study crashed or hung many standard UNIX utilities. Results depend on the harness and its seeds, because random bytes that fail validation reach little of the code.

What can be fuzzed besides parsers?

Besides parsers, fuzzing suits code that takes input from outside the program, e.g. an API request handler. Command-line argument handling can be fuzzed too, and so can a pair of functions that undo each other, with a harness that checks the round trip.

Why fuzz code that already has high coverage?

Code with high coverage has run most of its lines with at least one input, not with every input. A fuzzer runs the same lines with many other inputs, which is how the Acme parser crashed on a line its passing tests had covered.

How do you fuzz command-line arguments?

Command-line arguments can be fuzzed through a harness that turns each fuzzer input into a list of arguments for the tool's argument parser. The harness runs in one process, which is faster than starting the program per input.