Skip to content

What are catching tests (catching JiTTests)?

Catching tests are tests generated for a specific change and meant to fail if that change introduced a bug, rather than to guard the code for years.

Last updated , 8 min read

What are catching tests (catching JiTTests)?

A catching test is a generated test aimed at one proposed code change and built to fail if that change introduced a bug. Meta calls them Catching JiTTests, short for catching just-in-time tests, because a tool writes them before the change merges. The name does not refer to just-in-time (JIT) compilers or to scheduling when tests run.

After review, a catching test is usually thrown away instead of joining the test suite. In Meta's research, a catching test is contrasted with a hardening test, which passes when it is written and stays in the codebase to guard against later bugs.

A catching test is also not meant to raise code coverage. Meta's version checks whether one change broke behavior that worked before.

How are catching tests generated?

In the method Meta's researchers describe, a tool works through six steps for each proposed change:

  1. The tool reads the proposed change, with its title and summary.
  2. A language model infers what the change is meant to do and lists ways it could go wrong.
  3. The tool turns each risk into a mutant, a copy of the old code with that bug planted in it.
  4. The model writes a test for each mutant that passes on the old code and fails on the mutant.
  5. The tool runs each test on the changed code, and a test that fails there becomes a candidate catch.
  6. Assessors filter out likely false alarms, and a person picks which of the rest to send to the change's author.
How a catching test reaches the author of a change Proposed change code and summary Mutant test must fail Generated test aimed at one risk Old code test must pass Changed code test fails Assessors drop false alarms Author of the change was this intended?
A failure counts only when the same test passed on the old code. The assessors, a person's check, and one question to the author separate real bugs from false alarms.

The mutants do a different job than in mutation testing. Mutation testing uses mutants to judge whether existing tests fail on wrong code. Here they are targets for writing tests that then run on the real change. The old code supplies the expected result, so a failure shows only that behavior changed. Whether that is a bug depends on what the change was meant to do, which is a test oracle question.

What is an example of a catching test?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Save the delivery address to the customer's account during checkout." The agent writes the change and its own tests, and all of them pass.

A test generation tool reads the change and infers that the rest of the order should stay as it was. One risk it lists is that saving the address rebuilds the order without its items. Acme's old checkout already had an address step, so a test can run it on both versions. The tool plants that bug in the old code and generates this test, which passes on the old code and fails on the mutant:

test("changing the address keeps the cart", async () => {
  const order = await startCheckout({ items: 2 });
  await order.setAddress("12 Elm Street");
  expect(order.items).toHaveLength(2);
});

The tool then runs the test on the agent's change, where it fails:

expect(received).toHaveLength(expected)

Expected length: 2
Received length: 0
Received array:  []

The assessors find no sign of a broken test, so the tool asks the developer whether the order should drop from 2 items to 0. The developer answers no. The address saved, but the cart emptied.

The developer passes the failing test to the agent, which fixes the save function. The test now passes on the change, and the developer discards it. This example is simplified. A real tool would generate many tests per change and drop most of them before a person sees them.

What changes when a coding agent writes the code?

A coding agent writes its own tests against the changed code, so they tend to agree with it, bugs included. A catching test takes its expected result from the old code, so it can fail on a change where the agent's tests give a false pass.

The inferred intent comes partly from the change summary, which a coding agent often writes. A team can give the tool the original request as well.

A confirmed catch gives the agent a runnable form of reproduction steps, often close to a minimal reproducible example. A team can send it in a bug report and keep the test outside the files the agent may edit.

Without a dedicated tool, a separate agent session can write tests that pass on the old code and run them on the changed code. That setup reports more false alarms, because no assessors filter its failures.

How do you filter false alarms?

A catching test fails on any change in the behavior it checks, so many of its failures are not bugs. A false positive here is a failure that reports a bug the change does not have. Common sources of false alarms include these three:

  • Intended change. The change was meant to alter the behavior, so a test that expects the old behavior fails.
  • Broken test. The generated test has a bug of its own or checks an internal detail, e.g. an exact call count.
  • Flaky run. A flaky test passes and fails on the same code, e.g. when a server it needs is sometimes down.

Meta's system uses two kinds of assessor. Assessors built on fixed rules match known patterns in the change, the test code, and the run's trace. Assessors built on language models read the test, its failure, the change, and the inferred intent, and rate how likely a real bug is.

In Meta's first deployment, a person picked which remaining failures went to the author. The author sees a plain question, whether the change in behavior was intended, and opens the generated test only if the answer is no.

A reported failure is still often a false alarm. In a 2026 Meta study, of 41 candidate catches reported to engineers, 8 were real bugs and 4 of them were serious.

Generating throwaway tests for each change has these trade-offs:

  • No upkeep. The tests never join the suite, so nobody updates them when the code changes.
  • No lasting guard. A discarded test cannot catch the same bug if it returns.
  • Review time. Each report takes a person's time, and frequent false alarms teach people to ignore reports.
  • Narrow reach. The tests check only the risks the model listed, so reviewers in code review still read the rest of the change. No findings is not the same as complete coverage.
  • Old behavior only. In Meta's method, a catching test must pass on the old code, so it cannot check a feature the change adds.
  • Compute cost. Each change needs many model calls and test runs, so Meta ran generation only on changes a risk model flagged.

How are catching tests different from hardening tests?

A hardening test stays in the suite, where it runs in regression testing to catch later breakage. A catching test is written to fail on a bad change and is usually discarded after review. A catching test that passes on the final change can be kept as a hardening test. In Meta's framework, a hardening test also counts as a catching test when a later change breaks the behavior it checks.

The two kinds differ on these points:

PointHardening testCatching test
When it is writtenAny time, usually with the codeAfter a change is proposed, before it merges
Result when writtenPassesPasses on the old code, fails on a bad change
Where it livesIn the test suiteOutside the codebase, then usually discarded
What it protectsCurrent behavior against later changesThe codebase against one change

A tool can generate either kind. Meta's ACH system generated mutants aimed at privacy bugs and then wrote tests to catch them. Engineers accepted 73% of those tests, which were hardening tests written against the current code. Both kinds differ from test impact analysis, which picks the existing tests to run for a change and writes no tests.

FAQs

Who reviews a catching test that fails?

A catching test that fails is reviewed by people after automated assessors have removed likely false alarms. In Meta's first deployment, a person picked which failures to send to the author of the change. The author answers one plain question, whether the change in behavior was intended, and reads the test only if the answer is no.

Are catching tests kept in the suite?

Catching tests are usually not kept in the suite. They are generated for one change and discarded after review, which leaves nothing to maintain. A catching test that passes on the final change can be kept as a hardening test.

How are catching tests different from mutation testing?

Catching test generation at Meta and mutation testing both use mutants, which are copies of code with planted bugs. Mutation testing uses mutants to judge whether existing tests fail on wrong code. Catching test generation uses them as targets for writing tests, then runs those tests on the real change.

Do catching tests replace code review?

Catching tests do not replace code review. They check only the risks that a model listed for one change, so reviewers still read the rest of the change. A confirmed catch adds a failing test that the author can pass to a coding agent.