# Verifying AI-written code

How to check that software written by coding agents works, from "done" claims and gamed tests to evidence from running the code.

Start here

## [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)

Testing AI-generated code means checking what the software does, not who wrote it, by reading the change, running its tests, and running the app itself.

11 min read

## Concepts

### [What is runtime verification?](https://specstory.com/learning/verification/runtime-verification)

Runtime verification is a method that checks a running program against stated properties by observing its behavior, instead of analyzing its source alone.

### [What are AI hallucinations in code?](https://specstory.com/learning/verification/ai-code-hallucinations)

AI hallucinations in code are imports, calls, flags, or options that name a package or API that does not exist, generated in the same form as real code.

### [What is a definition of done for coding agents?](https://specstory.com/learning/verification/definition-of-done-for-coding-agents)

A definition of done for coding agents lists the evidence a change must produce, e.g. a passing run of the app, before the work counts as finished.

### [What is AI slop in code?](https://specstory.com/learning/verification/ai-slop)

AI slop in code is AI-generated output that looks finished but is bloated, duplicated, or poorly tested, and it moves the work of checking onto reviewers.

### [What is LLM-as-a-judge?](https://specstory.com/learning/verification/llm-as-a-judge)

LLM-as-a-judge is the use of a language model to grade outputs against a rubric, which scales review but adds biases and does not run the software.

### [What is reward hacking in coding agents?](https://specstory.com/learning/verification/reward-hacking)

Reward hacking is when a coding agent satisfies the check it is graded on, e.g. the test suite, without doing the task the user intended.

### [What is verification debt?](https://specstory.com/learning/verification/verification-debt)

Verification debt is the growing gap between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

## Comparisons

### [What is the difference between reading code and running code?](https://specstory.com/learning/verification/reading-vs-running-code)

Reading code, as in AI code review, finds problems visible in the source, while running it, as in testing, finds failures that appear only at runtime.

### [What is the difference between verification and validation?](https://specstory.com/learning/verification/verification-vs-validation)

Verification checks software against its specification, while validation checks it against users' needs, and neither is defined by whether the code runs.

## Common questions

### [Why do coding agents say "done" when the code doesn't work?](https://specstory.com/learning/verification/coding-agent-done-claims)

A coding agent's "done" is a completion claim written by the same process that wrote the code, so it can arrive with no passing run behind it.

### [Why do coding agents delete or weaken tests?](https://specstory.com/learning/verification/test-tampering)

Test tampering is a coding agent behavior that edits, skips, or deletes a failing test so the suite passes, instead of fixing the code it checks.

### [Can AI check its own code?](https://specstory.com/learning/verification/ai-self-verification)

AI self-verification is a model reviewing or testing its own output, a useful first pass whose checker shares the author's context and blind spots.

### [Why do AI-written tests always pass?](https://specstory.com/learning/verification/ai-written-tests-always-pass)

AI-written tests often pass because they assert what the code already does, mock the parts that fail, or check almost nothing, so a green run shows little.

## Failures and bugs

### [What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)

A false pass is a test or check that reports success while the software is broken, because the check missed the defect or the failure was lost.

### [Common bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs)

Common bugs in AI-generated code are recurring failures that a quick demo does not reveal, from open data access to features that never run.

## Terms in this topic

**[AI code verification](https://specstory.com/learning/glossary#ai-code-verification)**

AI code verification is the practice that checks what AI-generated software does, by reading it, running its tests, and running the software, rather than detecting who wrote it.

**[AI slop](https://specstory.com/learning/glossary#ai-slop)**

AI slop is AI-generated output that looks finished but is bloated, duplicated, or poorly tested because nobody checked it closely. In code it includes verbose logic and tests that assert almost nothing.

**[Assertion-free test](https://specstory.com/learning/glossary#assertion-free-test)**

An assertion-free test is a test that runs code but checks no result, so it passes whenever the code does not crash, whether the output is right or wrong.

**[Cross-model verification](https://specstory.com/learning/glossary#cross-model-verification)**

Cross-model verification is a checking setup that asks a different AI model to review work that another model produced, in the hope that their blind spots differ.

**[Definition of done](https://specstory.com/learning/glossary#definition-of-done)**

A definition of done is a shared checklist that states what evidence a piece of work must have before a team counts it as finished.

**[Eval](https://specstory.com/learning/glossary#eval)**

An eval, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results, so versions of the system can be compared.

**[False completion claim](https://specstory.com/learning/glossary#false-completion-claim)**

A false completion claim is a coding agent's report that a task is "done," or that tests pass, when the work has not been built, run, or checked.

**[False negative](https://specstory.com/learning/glossary#false-negative)**

A false negative is a test result that misses a defect that does exist, so the check passes even though the software is broken.

**[False pass](https://specstory.com/learning/glossary#false-pass)**

A false pass is a test or check result that reports success while the software is broken, because the check missed the defect or lost the failure.

**[Ghost feature](https://specstory.com/learning/glossary#ghost-feature)**

A ghost feature is a feature that exists in the code and may pass its tests but never runs for users, e.g. a handler that nothing calls.

**[LLM-as-a-judge](https://specstory.com/learning/glossary#llm-as-a-judge)**

LLM-as-a-judge is an evaluation setup that uses a language model to grade output against written criteria, instead of or alongside deterministic checks.

**[Package hallucination](https://specstory.com/learning/glossary#package-hallucination)**

Package hallucination is a code generation failure that imports or installs a package that does not exist. An attacker can register the invented name and publish harmful code under it.

**[Reward hacking](https://specstory.com/learning/glossary#reward-hacking)**

Reward hacking is a behavior in which a coding agent satisfies the check it is graded on, e.g. the tests, without doing the task the user intended.

**[Row-level security](https://specstory.com/learning/glossary#row-level-security)**

Row-level security is a database feature that limits which rows each user can read or change, so one user's data stays hidden from another.

**[Runtime verification](https://specstory.com/learning/glossary#runtime-verification)**

Runtime verification is a method that checks a running program against stated properties by observing its behavior, usually through a monitor that reads its execution. Developers also use the term for running the software to check a change.

**[Slopsquatting](https://specstory.com/learning/glossary#slopsquatting)**

Slopsquatting is an attack that registers package names that language models tend to invent, so code that installs a hallucinated package installs the attacker's code instead.

**[Specification gaming](https://specstory.com/learning/glossary#specification-gaming)**

Specification gaming is a behavior in which an AI system satisfies the literal objective it was given without doing the task the objective was meant to capture.

**[Test tampering](https://specstory.com/learning/glossary#test-tampering)**

Test tampering is a coding agent behavior that edits, skips, or deletes a failing test so the suite passes, instead of fixing the code the test was checking.

**[Verification debt](https://specstory.com/learning/glossary#verification-debt)**

Verification debt is the gap that grows between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

19 terms from the [Learning Center glossary](https://specstory.com/learning/glossary).

---

Source: [Verifying AI-written code | SpecStory](https://specstory.com/learning/verification)
