Skip to content

Verifying AI-written code

How to check that software written by coding agents works, from "done" claims and gamed tests to evidence from running the code.

Start here

How to test AI-generated code

Testing AI-generated code means checking what the software does, not who wrote it, by reading the change, running its tests, and running the app itself.

11 min read

Concepts

  • What is runtime verification?

    Runtime verification is a method that checks a running program against stated properties by observing its behavior, instead of analyzing its source alone.

  • What are AI hallucinations in code?

    AI hallucinations in code are imports, calls, flags, or options that name a package or API that does not exist, generated in the same form as real code.

  • What is a definition of done for coding agents?

    A definition of done for coding agents lists the evidence a change must produce, e.g. a passing run of the app, before the work counts as finished.

  • What is AI slop in code?

    AI slop in code is AI-generated output that looks finished but is bloated, duplicated, or poorly tested, and it moves the work of checking onto reviewers.

  • What is LLM-as-a-judge?

    LLM-as-a-judge is the use of a language model to grade outputs against a rubric, which scales review but adds biases and does not run the software.

  • What is reward hacking in coding agents?

    Reward hacking is when a coding agent satisfies the check it is graded on, e.g. the test suite, without doing the task the user intended.

  • What is verification debt?

    Verification debt is the growing gap between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

Comparisons

Common questions

Failures and bugs

Terms in this topic

AI code verification
AI code verification is the practice that checks what AI-generated software does, by reading it, running its tests, and running the software, rather than detecting who wrote it.
AI slop
AI slop is AI-generated output that looks finished but is bloated, duplicated, or poorly tested because nobody checked it closely. In code it includes verbose logic and tests that assert almost nothing.
Assertion-free test
An assertion-free test is a test that runs code but checks no result, so it passes whenever the code does not crash, whether the output is right or wrong.
Cross-model verification
Cross-model verification is a checking setup that asks a different AI model to review work that another model produced, in the hope that their blind spots differ.
Definition of done
A definition of done is a shared checklist that states what evidence a piece of work must have before a team counts it as finished.
Eval
An eval, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results, so versions of the system can be compared.
False completion claim
A false completion claim is a coding agent's report that a task is "done," or that tests pass, when the work has not been built, run, or checked.
False negative
A false negative is a test result that misses a defect that does exist, so the check passes even though the software is broken.
False pass
A false pass is a test or check result that reports success while the software is broken, because the check missed the defect or lost the failure.
Ghost feature
A ghost feature is a feature that exists in the code and may pass its tests but never runs for users, e.g. a handler that nothing calls.
LLM-as-a-judge
LLM-as-a-judge is an evaluation setup that uses a language model to grade output against written criteria, instead of or alongside deterministic checks.
Package hallucination
Package hallucination is a code generation failure that imports or installs a package that does not exist. An attacker can register the invented name and publish harmful code under it.
Reward hacking
Reward hacking is a behavior in which a coding agent satisfies the check it is graded on, e.g. the tests, without doing the task the user intended.
Row-level security
Row-level security is a database feature that limits which rows each user can read or change, so one user's data stays hidden from another.
Runtime verification
Runtime verification is a method that checks a running program against stated properties by observing its behavior, usually through a monitor that reads its execution. Developers also use the term for running the software to check a change.
Slopsquatting
Slopsquatting is an attack that registers package names that language models tend to invent, so code that installs a hallucinated package installs the attacker's code instead.
Specification gaming
Specification gaming is a behavior in which an AI system satisfies the literal objective it was given without doing the task the objective was meant to capture.
Test tampering
Test tampering is a coding agent behavior that edits, skips, or deletes a failing test so the suite passes, instead of fixing the code the test was checking.
Verification debt
Verification debt is the gap that grows between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

19 terms from the Learning Center glossary.