Skip to content

What is adversarial testing in software?

Adversarial testing is deliberately trying to break software with unexpected, hostile, or careless input and actions, instead of confirming the happy path.

Last updated , 8 min read

What is adversarial testing in software?

Adversarial testing is a testing approach that tries on purpose to make software fail with unplanned, hostile, or careless input and actions. The goal is to find what breaks, not to confirm the expected path. In machine learning, the same name means feeding a model inputs built to make it produce wrong or unsafe output.

Tests written alongside a feature usually check that it works when people use it as intended. Adversarial testing checks what happens when they do not. Careless users send odd input and actions by accident, and attackers send them on purpose.

A suite can reach high code coverage without ever sending the software input it was not built for. Coverage records which lines ran, not which inputs the tests tried. Adversarial testing aims at the inputs and orders of actions that the code's authors did not plan for.

How does adversarial testing work?

An adversarial test session is a loop of four steps:

  1. The tester picks one assumption in the code, e.g. that the delivery address is never empty.
  2. The tester sends input or an action that breaks that assumption.
  3. The tester checks the result for a failure signal, e.g. a server error page.
  4. The tester records the steps behind each failure and keeps them as a test that runs on later changes. If nothing failed, the tester moves to the next assumption.
The adversarial testing loop Pick an assumption address never empty Break it unplanned input Check the result failure signal? Keep as a test steps that repeat it no failure, next assumption
The check in the third step decides the path. A clean result closes one assumption only, so the loop starts again with the next one.

Testers use several techniques to break an assumption:

  • Bad input. Negative testing sends invalid, unexpected, or missing input and checks that the software handles it without crashing or saving bad data.
  • Repeated and interrupted actions. Testers repeat an action or stop it halfway, e.g. clicking "Save" twice.
  • Random actions. Monkey testing sends random clicks and keystrokes without a script, to see whether the software crashes.
  • Generated input. Fuzzing feeds a program large numbers of random or malformed inputs. Property-based testing generates inputs against a rule that must hold for each of them.
  • Tampered requests. Testers send requests straight to the server, past the checks in the page. A two-account test signs in as two users and checks that neither can see or change the other's data.
  • Guided exploration. In exploratory testing, a tester designs and runs tests at the same time, guided by what the software does. A bug bash puts many people on the software at once, often including people from outside the team that built it.

What is an example of adversarial testing?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The agent's own test enters a valid address, saves it, and passes. The developer then starts checkout with 2 items in the cart and tries to break the change:

  1. The developer saves an empty address. The page shows the message "Enter a delivery address." That is correct.
  2. The developer pastes an address of 500 characters. The form rejects it with a length message.
  3. The developer clicks "Save address" twice in quick succession. The address saved, but the cart emptied.
  4. The developer signs in as a second customer and tries to change the address on order A-1042, which belongs to the first customer. The page shows "Order not found."
  5. The developer writes the double click into a test. The test fails on the current code and passes once the agent fixes the save handler.

The test from step 5 is:

test("saving the address twice keeps the cart", async ({ page }) => {
  await addItemsToCart(page, 2);
  await page.goto("https://shop.example.com/checkout");
  await page.getByLabel("Delivery address").fill("12 Elm Street");
  await page.getByRole("button", { name: "Save address" }).dblclick();
  await expect(page.getByText("Address saved")).toBeVisible();
  await page.reload();
  await expect(page.locator(".cart-item")).toHaveCount(2);
});

After the reload, the last line checks the saved cart, which the agent's own test never did. This example is simplified. A real project would repeat the same attacks on the payment step.

What changes when a coding agent writes the code?

A prompt usually describes the input a feature should accept, e.g. a valid address. When the coding agent also writes the tests, they send that input and pass. The prompt rarely says what should happen with hostile or careless input. The code then often has no guard for it, and the tests have no case for it. Many bugs in AI-generated code sit in that gap, e.g. one customer reading another customer's orders.

The adversarial idea also works between agents. A 2026 paper set an agent that writes tests against an agent that injects bugs. The tests from that setup detected more real Defects4J faults than earlier large language model (LLM) methods and EvoSuite. Defects4J is a public set of real bugs from Java projects. EvoSuite generates tests with a search algorithm.

A fix can be as narrow as the failure report. Told that a double click empties the cart, a coding agent can disable the button after the first click. Two save requests sent straight to the server can still empty the cart. A practical adjustment is to retest each fix by breaking the same assumption another way, not only by repeating the reported steps.

What are the limits of adversarial testing?

Adversarial testing finds failures, but it cannot show that none are left. Its main limits are:

  • Sampling. A session tries a small sample of the possible inputs and actions. Each change can also open another way to break the software. No findings is not the same as complete coverage.
  • Oracles. A crash marks a failure on its own, but a wrong total that looks plausible passes unless someone knows the right answer. That source of the right answer is the test oracle.
  • Scope. Adversarial testing checks the software, not its tests. Mutation testing checks the tests instead, by injecting bugs and counting how many they catch.
  • Repeatability. Results depend on who runs the session. A session run by a person is hard to repeat the same way.
  • Safety. Hostile input can delete or corrupt data, so attacks belong on a disposable copy of the app with test data.

How is adversarial testing different from red teaming?

Red teaming has a team act as an attacker against a whole system, usually with a goal, e.g. reaching customer data. Adversarial testing usually attacks one feature or change, and it covers careless users as well as hostile ones.

Fuzzing is one adversarial testing technique. The three practices compare as follows:

PracticeWho attacksTargetTypical result
Adversarial testingA tester or a second agentOne feature or changeA failure with steps to repeat it
Red teamingA team acting as an attackerA whole system and its defensesA path to data or control
FuzzingA tool that generates inputA program's input handlingA crash or hang on one input

Red teams for LLM applications also attack the model with prompt injection, which puts instructions into text the model reads. The OWASP Gen AI Security Project ranks prompt injection first in its 2025 list of LLM risks.

FAQs

Should a team pay people to try to break its app before launch?

Paying people to try to break an app before launch can find failures that its builders miss, because outside testers use the app in ways nobody planned. The payment is worth more when each finding comes with steps that repeat it, so the team can keep it as a test.

Can a team assume users will not do silly things?

A team cannot assume that users will avoid silly actions, e.g. clicking a button twice. Users do these things by accident, and attackers do them on purpose. Adversarial testing tries such actions before customers do.

How is adversarial testing of an AI model different?

Adversarial testing of an AI model feeds the model inputs built to make it produce wrong or unsafe output, e.g. hidden instructions in content the model reads. Adversarial testing of software attacks a program's input and workflows instead.

Is adversarial testing a one-time event?

Adversarial testing is not a one-time event, because each change can open another way to break the software. Teams repeat it as the software changes and keep each failure it finds as a test that runs on later changes.