What is a test oracle?
A test oracle is the source of the expected result that a test compares with the actual result. The oracle can be a fixed value, e.g. an order total of $240.00, or a rule that every result must follow. When the actual result and the oracle disagree, the test fails.
The International Software Testing Qualifications Board (ISTQB) defines a test oracle as "a source to determine an expected result." In a 2015 survey, Barr and colleagues use the term more widely, for any check that tells correct behavior from incorrect behavior. Under that wider meaning, a crash is an implicit oracle, because it marks a failure without any expected value.
A unit test usually states its oracle as an assertion. A suite can reach high code coverage with weak oracles, because coverage counts the lines that ran, not the results that were checked.
How does a test oracle work?
An automated test uses its oracle in four steps:
- The test runs the software with a known input.
- The test records the actual result, e.g. the value a function returns.
- The test compares the actual result with the oracle.
- The test passes if the two agree and fails if they do not.
The expected result can come from four types of oracle:
- Specified oracle. A person writes the expected value or rule in advance, usually from the requirements. Property-based testing uses rules that must hold for every generated input.
- Derived oracle. The expected result comes from something that already exists, e.g. an earlier version of the program. Differential testing and golden file testing work this way.
- Implicit oracle. A general failure signal marks the result as wrong, e.g. a crash. A nonzero exit code from a command-line tool is a common implicit oracle.
- Human oracle. A person looks at the result and judges it. This check is slow and hard to repeat the same way, and LLM-as-a-judge setups use a language model in the person's place.
What is an example of a test oracle?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent changes the checkout code and writes this test:
test("customer can change the delivery address", async () => {
const order = await startCheckout({ items: 2 });
await order.setAddress("12 Elm Street");
expect(order.address).toBe("12 Elm Street");
});
The test passes. Its only explicit oracle is the expected address. The address saved, but the cart emptied. The test never checks the cart, so it cannot detect the failure.
A stronger oracle comes from the request itself. Changing the address should not change what the customer is buying. That rule adds one line to the test:
expect(order.items).toHaveLength(2);
With that line, the test fails on the broken change. The code under test is the same, and only the oracle changed. This example is simplified. A real checkout would need more oracles, e.g. one for the order total.
What changes when a coding agent writes the code?
When a coding agent writes both the code and its tests, the agent generates both from the same prompt and context, a limit of AI self-verification. If the code misses part of the request, the test often misses the same part and passes. A 2026 study of 86,156 changes to test files by five coding agents found that 80.2% had weak or no explicit oracle signals. Many of those tests confirmed that the code ran while checking little or nothing about its output.
An agent that is scored on passing tests can also change the oracle itself. On ImpossibleBench, where passing a task requires breaking its specification, GPT-5 "passed" 54% of one set of impossible SWE-bench tasks, by editing the tests or gaming them. This is one form of reward hacking.
The same paper found that read-only tests stopped the test edits without lowering scores on the original tasks, but they did not stop other shortcuts. A safer setup keeps at least one oracle in a file that the agent can read but not edit.
What is the test oracle problem?
The test oracle problem is the difficulty of deciding whether a test passed or failed for a given input and state. It is hardest when nobody knows the correct answer in advance, e.g. the ranking a search engine returns. The problem sets four limits on every test suite:
- An oracle checks only what it names. The first checkout test checked the address and missed the cart. Mutation testing measures how many injected bugs a suite's oracles catch.
- A wrong oracle can hide a wrong program. When the expected result is itself wrong, e.g. a golden file recorded from buggy output, code that repeats the bug passes.
- An implicit oracle catches only plain failures. A command that prints a wrong total and exits with code 0 passes it.
- Some results have no exact expected value. Teams then use partial oracles, which check some properties of a result when the full expected result is unknown. Metamorphic testing is one example, because it checks how the output should change when the input changes in a known way.
How is a test oracle different from an assertion?
An assertion is a line of test code that checks a condition and fails the test when the condition is false. The oracle is where the expected condition comes from. In expect(order.items).toHaveLength(2), the assertion is the whole statement, including the toHaveLength(2) matcher that does the check. The oracle is the rule that changing the address keeps the cart intact.
A test can contain many assertions and still have a weak oracle, e.g. when each assertion rebuilds its expected value with the code's own formula. That pattern is called a tautological test, and it is one reason AI-written tests can pass on broken code.
How can a team get an oracle the coding agent did not write?
Write some checks from the request before the agent starts. Then run the finished software through end-to-end testing, which adds implicit oracles that the agent's tests did not include.
RunStory independently runs the software against your change and returns evidence to the coding agent. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. After your agent makes the change, RunStory repeats the failing workflow to check that the problem is resolved. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
What is an automated test oracle?
An automated test oracle is an oracle that a program applies without a person, e.g. an assertion that compares two totals. Specified, derived, and implicit oracles can be applied this way. A human oracle cannot, because a person judges each result.
What is a partial oracle?
A partial oracle checks some properties of a result when the full expected result is unknown. A test that checks that an order total is never negative uses a partial oracle. It catches some wrong totals and misses others.
How do you test code when you do not know the correct answer?
When the exact answer is unknown, a test can check a relation instead of a value. Metamorphic testing checks how the output should change when the input changes in a known way. Property-based testing checks a rule that must hold for every generated input.
Do passing tests prove that code is correct?
Passing tests do not prove that code is correct. Passing tests show that the results matched the test oracles on the inputs that ran. If an oracle is weak or missing, the tests can pass while the code is still broken.