# Glossary

Short definitions of the terms used across the Learning Center. Each links to its full article when one exists.

190 terms

## A

**Acceptance criteria**

Acceptance criteria are the conditions that a change must meet for the person who asked for it to accept it. They are written so that a person or a test can check each one.

**Acceptance testing**

Acceptance testing is a level of testing that checks whether a system meets its acceptance criteria and its users' needs. It lets the system's users and owners decide whether to accept it.

[Read the article: What is acceptance testing (UAT)?](https://specstory.com/learning/testing/acceptance-testing)

**Accessibility tree**

The accessibility tree is the browser's structured view of a page's elements, names, and roles, which assistive technology and many browser agents read the page through.

**Adversarial testing**

Adversarial testing is a testing approach that deliberately tries to break software with unexpected, hostile, or careless inputs and actions, instead of confirming the expected path.

[Read the article: What is adversarial testing in software?](https://specstory.com/learning/test-quality/adversarial-testing)

**Agent harness**

An agent harness is the software around a language model that gives it tools, context, memory, and a loop, which turns the model into a working agent.

[Read the article: What is an agent harness?](https://specstory.com/learning/ai-coding/agent-harness)

**Agent hook**

An agent hook is a script that a coding agent's harness runs at a fixed point in its loop. A Stop hook runs when the agent tries to end its turn.

[Read the article: What are agent hooks (Stop hooks)?](https://specstory.com/learning/ci-cd/agent-hooks)

**Agent loop**

The agent loop is the cycle in which a model requests an action, the harness runs the tool, and the model reads the result. It ends when the model requests no tool, or when a step limit, a time limit, or a person stops it.

[Read the article: What is a coding agent?](https://specstory.com/learning/ai-coding/coding-agent)

**Agent trajectory**

An agent trajectory is the ordered sequence of an agent's actions, tool calls, and observations while it works on a task, which is what a transcript records.

[Read the article: What is AI coding session history?](https://specstory.com/learning/ai-coding/ai-coding-session-history)

**Agentic engineering**

Agentic engineering is a way of building software that directs coding agents while keeping engineering practices, from written specs and tests to code review and verification.

[Read the article: What is agentic engineering?](https://specstory.com/learning/ai-coding/agentic-engineering)

**Agentic testing**

Agentic testing is software testing that AI agents carry out themselves, choosing, running, and judging checks on the software, rather than only writing test scripts for later.

[Read the article: What is agentic testing?](https://specstory.com/learning/testing/agentic-testing)

**AGENTS.md**

AGENTS.md is a Markdown file in a repository that gives coding agents project instructions, e.g. how to build, test, and change the code.

[Read the article: What is AGENTS.md?](https://specstory.com/learning/ai-coding/agents-md)

**AI agent**

An AI agent is a program that uses a language model in a loop to choose actions, call tools, and read the results until it finishes a task.

**AI code attribution**

AI code attribution is the practice that records which code an AI tool wrote and under which prompt, e.g. with commit trailers.

**AI code review**

AI code review is a review practice that uses a language model to read a pull request or diff and comment on possible bugs, style problems, and risks.

[Read the article: What is AI code review?](https://specstory.com/learning/code-review/ai-code-review)

**AI code verification**

AI code verification is the practice that checks what AI-generated software does, by reading it, running its tests, and running the software, rather than detecting who wrote it.

[Read the article: How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)

**AI slop**

AI slop is AI-generated output that looks finished but is bloated, duplicated, or poorly tested because nobody checked it closely. In code it includes verbose logic and tests that assert almost nothing.

[Read the article: What is AI slop in code?](https://specstory.com/learning/verification/ai-slop)

**Approval fatigue**

Approval fatigue is a failure in which people approve a coding agent's permission prompts without reading them, because the prompts are so frequent that they stop protecting anything.

[Read the article: What is approval fatigue in coding agents?](https://specstory.com/learning/environments/approval-fatigue)

**Arrange-act-assert**

Arrange-act-assert is a test layout that sets up the inputs first, then runs the code under test, and then checks the result, each in its own block.

**Assertion**

An assertion is a statement in a test that checks a condition and fails the test when the condition is false.

**Assertion-free test**

An assertion-free test is a test that runs code but checks no result, so it passes whenever the code does not crash, whether the output is right or wrong.

[Read the article: Why do AI-written tests always pass?](https://specstory.com/learning/verification/ai-written-tests-always-pass)

## B

**Background coding agent**

A background coding agent is a coding agent that works on a task without a person watching and returns a branch or pull request when it finishes.

**Behavior-driven development (BDD)**

Behavior-driven development is a practice that describes how software should behave as plain-language Given-When-Then scenarios that the team agrees on and tools can run as tests.

**Blameless postmortem**

A blameless postmortem is an incident review that works out how a failure happened and how to prevent a repeat, without blaming the individuals involved.

[Read the article: What is root cause analysis (RCA)?](https://specstory.com/learning/debugging/root-cause-analysis)

**Blast radius**

The blast radius of a change is the set of features, users, and systems that could break because of it. That set is often much wider than the files in the diff.

**Branch coverage**

Branch coverage is a code coverage measure that counts which outcomes of each decision point, e.g. both sides of an if statement, the tests have run.

[Read the article: What is code coverage?](https://specstory.com/learning/test-quality/code-coverage)

**Browser test framework**

A browser test framework is an open-source library that drives a real browser from test code, so a test can load pages, act on them, and check the results.

[Read the article: What is a browser test framework?](https://specstory.com/learning/cli-and-web/browser-test-frameworks)

**Bug bash**

A bug bash is a time-boxed session that gathers people from across a team to use the software together and report as many defects as they can find.

## C

**Catching test**

A catching test is a test that is generated for one code change and designed to fail if that change introduced a bug, then usually thrown away after review.

[Read the article: What are catching tests (catching JiTTests)?](https://specstory.com/learning/test-quality/catching-tests)

**Characterization test**

A characterization test is a test that records what existing code does now, right or wrong, so later changes can be checked against that recorded behavior.

[Read the article: What is golden file testing (golden master)?](https://specstory.com/learning/test-quality/golden-file-testing)

**Chrome DevTools Protocol**

The Chrome DevTools Protocol is an interface that lets a program drive and inspect a Chromium-based browser, and many automation tools and browser agents are built on it.

**CI pipeline**

A CI pipeline is the ordered set of automated steps, e.g. build, test, and package, that runs each time code is pushed or merged.

[Read the article: What is CI/CD?](https://specstory.com/learning/ci-cd/ci-cd)

**CI/CD**

CI/CD is a practice that merges code often and builds and tests each change in an automated pipeline, so the software is kept ready to release.

[Read the article: What is CI/CD?](https://specstory.com/learning/ci-cd/ci-cd)

**Code coverage**

Code coverage is a test measure that reports the percentage of a program's code that runs during its tests, which shows what was executed but not what was checked.

[Read the article: What is code coverage?](https://specstory.com/learning/test-quality/code-coverage)

**Code review**

Code review is a practice that has people or tools read a proposed code change before it merges, to catch defects, share knowledge, and keep the codebase consistent.

[Read the article: What is code review?](https://specstory.com/learning/code-review/code-review)

**Coding agent**

A coding agent is an AI system that edits files and runs commands in a loop to complete a programming task with little step-by-step direction.

[Read the article: What is a coding agent?](https://specstory.com/learning/ai-coding/coding-agent)

**Coding agent transcript**

A coding agent transcript is a saved record that holds a coding session's prompts, replies, tool calls, and results. A collection of saved transcripts is called AI coding session history.

[Read the article: What is AI coding session history?](https://specstory.com/learning/ai-coding/ai-coding-session-history)

**Command-line interface (CLI)**

A command-line interface is a way of using a program that takes typed commands, flags, and input in a terminal or script, and returns output, errors, and an exit code.

[Read the article: How to test a command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications)

**Commit**

A commit is a saved snapshot of changes in a Git repository, recorded with a message, an author, and a link to the commit before it.

**Container**

A container is a packaged process that runs with its own filesystem and limits but shares the host's kernel, which makes it lighter and less isolated than a virtual machine.

[Read the article: What is the difference between containers, gVisor, and microVMs?](https://specstory.com/learning/environments/containers-vs-gvisor-vs-microvms)

**Context window**

A context window is the amount of text, measured in tokens, that a language model can take into account at once, including instructions, files, and history.

[Read the article: What is context engineering?](https://specstory.com/learning/ai-coding/context-engineering)

**Continuous delivery**

Continuous delivery is a practice that keeps every change that passes the pipeline ready to release, so a person can release it to production at any time with one decision.

**Continuous deployment**

Continuous deployment is a practice that releases every change that passes the automated pipeline to production, with no manual approval step between the merge and the release.

**Continuous integration (CI)**

Continuous integration is a practice that merges every developer's changes into the main branch often, with an automated build and test run on each merge.

[Read the article: What is continuous integration (CI)?](https://specstory.com/learning/ci-cd/continuous-integration)

**Coverage-guided fuzzing**

Coverage-guided fuzzing is a fuzzing method that keeps the inputs which reach new code paths and mutates them further, so the fuzzer explores a program efficiently.

[Read the article: What is fuzzing?](https://specstory.com/learning/test-quality/fuzzing)

**Cross-model verification**

Cross-model verification is a checking setup that asks a different AI model to review work that another model produced, in the hope that their blind spots differ.

[Read the article: Can AI check its own code?](https://specstory.com/learning/verification/ai-self-verification)

## D

**DAST**

Dynamic application security testing (DAST) is a security practice that probes a running application from the outside for vulnerabilities. It is separate from functional testing.

**Debugging**

Debugging is the process that finds why software does the wrong thing, by reproducing the failure, narrowing down its cause, changing the code, and confirming the fix.

[Read the article: What is debugging?](https://specstory.com/learning/debugging/debugging)

**Definition of done**

A definition of done is a shared checklist that states what evidence a piece of work must have before a team counts it as finished.

[Read the article: What is a definition of done for coding agents?](https://specstory.com/learning/verification/definition-of-done-for-coding-agents)

**Design for testability**

Design for testability is a design approach that makes software easier to test, e.g. by giving a CLI clear exit codes and machine-readable output.

**Dev container**

A dev container is a container defined in the repository that provides a ready-made development environment with its tools installed.

[Read the article: What is the difference between Docker, dev containers, and VMs?](https://specstory.com/learning/environments/docker-vs-devcontainer-vs-vm)

**Dynamic analysis**

Dynamic analysis is a method that observes a program while it runs, e.g. through tests, to find bugs that only appear at runtime.

[Read the article: What is the difference between static and dynamic analysis?](https://specstory.com/learning/code-review/static-vs-dynamic-analysis)

## E

**Egress control**

Egress control is a network control that limits which outside hosts a sandbox or agent can reach, usually with an allowlist or a proxy.

**End-to-end test**

An end-to-end test is a test that follows a whole user journey through a running system, from the interface through services and data, the way a real user would.

[Read the article: What is end-to-end testing?](https://specstory.com/learning/testing/end-to-end-testing)

**Environment parity**

Environment parity is the degree to which development, test, staging, and production environments match, so a check in one predicts behavior in another.

**Ephemeral environment**

An ephemeral environment is a short-lived copy of an application that is created for one change or test run and destroyed afterwards. A preview environment is one kind.

[Read the article: What is an ephemeral environment?](https://specstory.com/learning/environments/ephemeral-environments)

**Equivalent mutant**

An equivalent mutant is a code change made during mutation testing that behaves exactly the same as the original program, so no test can ever detect or kill it.

**Eval**

An eval, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results, so versions of the system can be compared.

[Read the article: What is LLM-as-a-judge?](https://specstory.com/learning/verification/llm-as-a-judge)

**Exit code**

An exit code is the number that a program returns when it finishes, where 0 usually means success and any other value signals an error that scripts can check.

[Read the article: What are exit codes?](https://specstory.com/learning/cli-and-web/exit-codes)

**Exploratory testing**

Exploratory testing is a testing approach that designs and runs tests at the same time, guided by what the tester learns about the software while using it.

[Read the article: What is exploratory testing, and can an AI agent do it?](https://specstory.com/learning/testing/exploratory-testing)

## F

**Fail-open**

Fail-open is a behavior that lets an action through when a check errors or times out, while fail-closed code denies it. Fail-open authentication is a common security bug.

**False completion claim**

A false completion claim is a coding agent's report that a task is "done," or that tests pass, when the work has not been built, run, or checked.

[Read the article: Why do coding agents say "done" when the code doesn't work?](https://specstory.com/learning/verification/coding-agent-done-claims)

**False negative**

A false negative is a test result that misses a defect that does exist, so the check passes even though the software is broken.

[Read the article: What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)

**False pass**

A false pass is a test or check result that reports success while the software is broken, because the check missed the defect or lost the failure.

[Read the article: What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)

**False positive**

A false positive is a check result that reports a problem that does not exist, which wastes time and teaches people to ignore real findings.

**Fault seeding**

Fault seeding is a technique that deliberately inserts known bugs into code to see how many of them a test suite or review process detects.

[Read the article: What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

**Flaky test**

A flaky test is a test that both passes and fails on the same code. A single result from it cannot be trusted as evidence of a bug or of a fix.

[Read the article: What is a flaky test?](https://specstory.com/learning/debugging/flaky-tests)

**Formal verification**

Formal verification is a method that proves mathematically that a program or model meets a specification, instead of testing it on sample inputs.

**Fuzzing**

Fuzzing is a testing technique that feeds a program large numbers of random or malformed inputs to find crashes, hangs, and bugs in input handling.

[Read the article: What is fuzzing?](https://specstory.com/learning/test-quality/fuzzing)

## G

**Ghost feature**

A ghost feature is a feature that exists in the code and may pass its tests but never runs for users, e.g. a handler that nothing calls.

[Read the article: What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)

**Git bisect**

Git bisect is a Git command that runs a binary search through a repository's history to find the first commit where a behavior changed.

[Read the article: What is git bisect?](https://specstory.com/learning/debugging/git-bisect)

**Git worktree**

A Git worktree is an extra working directory attached to the same repository, so several branches can be checked out at once, e.g. one per parallel agent.

**Given-when-then**

Given-when-then is a scenario format that states a starting state, an action, and an expected outcome in plain language, used in behavior-driven development and acceptance criteria.

**Golden file**

A golden file is a saved, approved copy of a program's output that a test compares new output against, failing when anything in the output changes.

[Read the article: What is golden file testing (golden master)?](https://specstory.com/learning/test-quality/golden-file-testing)

**Goodhart's law**

Goodhart's law is the observation that a measure which becomes a target stops being a good measure, because people and systems optimize the number instead of the goal.

## H

**Happy path**

The happy path is the expected route through a feature, with valid input and no errors, that most demos and many generated tests follow.

**Headless browser**

A headless browser is a web browser that runs without a visible window and is controlled by code, so tests and agents can load pages and act on them.

[Read the article: What is a headless browser?](https://specstory.com/learning/cli-and-web/headless-browser)

**Heisenbug**

A heisenbug is a bug that disappears or changes when someone tries to observe it, e.g. when a debugger or extra logging changes the timing.

[Read the article: What is a race condition, and how do you reproduce one?](https://specstory.com/learning/debugging/race-condition)

**Hermetic test**

A hermetic test is a test that depends only on what it declares and sets up itself, with no shared state, network, or leftover files. It gives the same result anywhere.

[Read the article: What is test isolation (hermetic tests)?](https://specstory.com/learning/environments/test-isolation)

**Hook**

A hook is a script that a tool runs automatically at a set point in its workflow, e.g. Git running a check before it records a commit.

[Read the article: What are Git hooks?](https://specstory.com/learning/ci-cd/git-hooks)

**Human-in-the-loop**

Human-in-the-loop is a way of running an AI system in which a person approves, corrects, or reviews its work at set points before the work takes effect.

[Read the article: What is human-in-the-loop for coding agents?](https://specstory.com/learning/ai-coding/human-in-the-loop)

**Hydration error**

A hydration error is a web app failure that happens when the HTML a server rendered does not match what the browser's JavaScript renders on load.

[Read the article: What is a hydration error?](https://specstory.com/learning/cli-and-web/hydration-error)

## I

**Idempotency**

Idempotency is a property of an operation that gives the same result whether it runs once or several times, e.g. a payment request that is safe to retry.

**Implicit oracle**

An implicit oracle is a general failure signal that marks a result as wrong without any expected value, e.g. a crash.

[Read the article: What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)

**Indirect prompt injection**

Indirect prompt injection is a prompt injection that arrives through content an agent reads during a task, e.g. a web page, instead of from the user.

[Read the article: What is prompt injection in coding agents?](https://specstory.com/learning/environments/prompt-injection-in-coding-agents)

**Integration test**

An integration test is a test that checks that separate parts of a system work together, e.g. a service and its database, rather than each part alone.

[Read the article: What is integration testing?](https://specstory.com/learning/testing/integration-testing)

**Invariant**

An invariant is a condition that must stay true for every valid state or input of a program, e.g. an order total that is never negative.

## J

**Jailbreak**

A jailbreak is a prompt that gets a language model to break its own safety rules. Prompt injection differs because it attacks an application built on the model, by mixing untrusted text into its instructions.

## L

**Large language model (LLM)**

A large language model is a model that is trained on large amounts of text to generate text, including code, from a prompt. Coding agents use one to choose what to do next.

**Least privilege**

Least privilege is a security principle that gives a person, program, or agent only the access its task needs, for no longer than it needs it.

[Read the article: What is least privilege for AI agents?](https://specstory.com/learning/environments/least-privilege-for-ai-agents)

**Lethal trifecta**

The lethal trifecta is a risk pattern in which an agent has access to private data, reads untrusted content, and can send data out, which lets prompt injection leak data.

[Read the article: What is prompt injection in coding agents?](https://specstory.com/learning/environments/prompt-injection-in-coding-agents)

**Linting**

Linting is a static check that compares source code against style and correctness rules with a tool called a linter, without running the code.

[Read the article: What is linting?](https://specstory.com/learning/code-review/linting)

**LLM-as-a-judge**

LLM-as-a-judge is an evaluation setup that uses a language model to grade output against written criteria, instead of or alongside deterministic checks.

[Read the article: What is LLM-as-a-judge?](https://specstory.com/learning/verification/llm-as-a-judge)

**Load testing**

Load testing is a performance test that runs software under the traffic expected in production, to measure response times and find where it slows down.

## M

**MCP server**

An MCP server is a program that exposes tools or data to AI agents over the Model Context Protocol, so an agent can call them in a standard way.

[Read the article: What is the Model Context Protocol (MCP)?](https://specstory.com/learning/ai-coding/model-context-protocol)

**Merge queue**

A merge queue is a system that tests each approved pull request against the current main branch and the changes ahead of it, and merges only those that pass.

**Metamorphic testing**

Metamorphic testing is a technique that checks how a program's output should change when its input changes in a known way, so code can be tested without an exact expected output.

[Read the article: What is metamorphic testing?](https://specstory.com/learning/test-quality/metamorphic-testing)

**MicroVM**

A microVM is a small virtual machine that starts quickly and gives each workload its own kernel, which isolates it more strongly than a container.

**Minimal reproducible example**

A minimal reproducible example is the smallest complete code, input, and setup that still shows a bug, so anyone can run it and see the same failure.

[Read the article: What is a minimal reproducible example?](https://specstory.com/learning/debugging/minimal-reproducible-example)

**Model Context Protocol (MCP)**

The Model Context Protocol is an open standard that lets AI applications connect to tools and data through MCP servers, so an agent can call them in a uniform way.

[Read the article: What is the Model Context Protocol (MCP)?](https://specstory.com/learning/ai-coding/model-context-protocol)

**Monkey testing**

Monkey testing is a testing technique that feeds software random clicks, keystrokes, or inputs without a script, to see whether it crashes or misbehaves.

**Mutant**

A mutant is a copy of a program with one small deliberate change, e.g. a flipped comparison, that mutation testing uses to check whether the tests fail.

[Read the article: What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

**Mutation score**

Mutation score is a test-strength measure that reports the share of valid mutants a test suite detects, which shows how well the tests check behavior.

[Read the article: What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

**Mutation testing**

Mutation testing is a technique that checks a test suite by injecting small deliberate bugs, called mutants, and counting how many of them the tests catch.

[Read the article: What is mutation testing?](https://specstory.com/learning/test-quality/mutation-testing)

## N

**Negative testing**

Negative testing is a testing technique that checks how software handles invalid, unexpected, or missing input, instead of only the input it was designed for.

**Nondeterminism**

Nondeterminism is a property of a program or test that can produce different results from the same code and input, e.g. because of timing.

## O

**Observed vs. expected**

Observed vs. expected is the part of a bug report that states what the software did next to what it should have done, as facts rather than guesses about causes.

[Read the article: What are reproduction steps in a bug report?](https://specstory.com/learning/debugging/reproduction-steps)

**Out-of-scope edit**

An out-of-scope edit is a change that a coding agent makes outside what its task asked for, e.g. an unrequested refactor.

**OWASP Top 10**

The OWASP Top 10 is a list that the Open Worldwide Application Security Project publishes of the most critical security risk categories for web applications.

## P

**Package hallucination**

Package hallucination is a code generation failure that imports or installs a package that does not exist. An attacker can register the invented name and publish harmful code under it.

[Read the article: What are AI hallucinations in code?](https://specstory.com/learning/verification/ai-code-hallucinations)

**Partial oracle**

A partial oracle is a test oracle that checks some properties of a result when the full expected result is unknown, e.g. that an order total is never negative.

[Read the article: What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)

**Pass on retry**

A pass on retry is a test result that fails first and then passes when the same test runs again on the same code. It marks the test as flaky or the bug as intermittent.

[Read the article: How to tell a flaky test from a real bug](https://specstory.com/learning/debugging/flaky-test-or-real-bug)

**Playwright MCP**

Playwright MCP is an open-source server that lets an AI agent control a real browser through the Model Context Protocol, using structured page snapshots instead of screenshots.

[Read the article: What is Playwright MCP?](https://specstory.com/learning/cli-and-web/playwright-mcp)

**Pre-commit hook**

A pre-commit hook is a script that Git runs before it records a commit. It can lint, format, or test the staged changes and stop the commit if a check fails.

[Read the article: What is a pre-commit hook?](https://specstory.com/learning/ci-cd/pre-commit-hook)

**Pre-push hook**

A pre-push hook is a script that Git runs before it sends commits to a remote repository, which can run slower checks than a pre-commit hook and stop the push.

[Read the article: What is the difference between pre-commit, pre-push, and CI checks?](https://specstory.com/learning/ci-cd/pre-commit-vs-pre-push-vs-ci)

**Prompt injection**

Prompt injection is an attack that hides instructions in content an AI model reads, e.g. a web page, so the model follows them instead of its user.

[Read the article: What is prompt injection in coding agents?](https://specstory.com/learning/environments/prompt-injection-in-coding-agents)

**Property-based testing**

Property-based testing is a technique that states rules that should hold for every input, then checks them against many generated inputs and shrinks any failing input.

[Read the article: What is property-based testing?](https://specstory.com/learning/test-quality/property-based-testing)

**Pull request**

A pull request is a request that asks to merge a set of commits into another branch, where people and tools review and check the change first.

## Q

**Quality assurance (QA)**

Quality assurance is a set of activities that aims to prevent defects by improving how a team builds and tests software. Testing, a form of quality control, checks the product itself.

[Read the article: What is QA testing?](https://specstory.com/learning/testing/qa-testing)

**Quality gate**

A quality gate is a point in a delivery pipeline where a change is measured against set criteria before it goes further. Some gates block the change, and others only report.

[Read the article: What is a quality gate in CI/CD?](https://specstory.com/learning/ci-cd/quality-gate)

## R

**Race condition**

A race condition is a bug that makes a result depend on the timing or order of events running at the same time, e.g. two requests changing the same record.

[Read the article: What is a race condition, and how do you reproduce one?](https://specstory.com/learning/debugging/race-condition)

**Red teaming**

Red teaming is a practice that has a team act as an attacker or hostile user to find weaknesses in a system before real attackers or users do.

**Red-green-refactor**

Red-green-refactor is the test-driven development cycle that writes a failing test, then the code that makes it pass, then cleans up the code while the test stays green.

[Read the article: What is test-driven development (TDD)?](https://specstory.com/learning/testing/test-driven-development)

**Refactoring**

Refactoring is a practice that changes the structure of existing code without changing its behavior, e.g. to remove duplication, so the code is easier to read and change.

**Regression testing**

Regression testing is a type of testing that checks that a change has not broken behavior that worked before, including parts of the software the change did not touch.

[Read the article: What is regression testing?](https://specstory.com/learning/testing/regression-testing)

**Reproduction script**

A reproduction script is a runnable script that replays the steps that trigger a bug, so a person or an agent can reproduce the failure without guessing.

[Read the article: What are reproduction steps in a bug report?](https://specstory.com/learning/debugging/reproduction-steps)

**Reproduction steps**

Reproduction steps are the exact actions, inputs, and conditions that make a bug happen again, written so that a person or a coding agent can follow them.

[Read the article: What are reproduction steps in a bug report?](https://specstory.com/learning/debugging/reproduction-steps)

**Retesting**

Retesting is a check that repeats the exact steps that exposed a defect on the fixed code, to confirm that the failure no longer happens. It is often called fix verification.

[Read the article: How to verify a bug fix](https://specstory.com/learning/debugging/bug-fix-verification)

**Review bottleneck**

The review bottleneck is the queue that forms when code is produced faster than people can review it, so pull requests wait or get approved with less scrutiny.

[Read the article: What is the code review bottleneck?](https://specstory.com/learning/code-review/review-bottleneck)

**Reward hacking**

Reward hacking is a behavior in which a coding agent satisfies the check it is graded on, e.g. the tests, without doing the task the user intended.

[Read the article: What is reward hacking in coding agents?](https://specstory.com/learning/verification/reward-hacking)

**Root cause analysis (RCA)**

Root cause analysis is a method that traces a failure past its symptoms to the underlying cause, so the fix removes the reason the failure happened.

[Read the article: What is root cause analysis (RCA)?](https://specstory.com/learning/debugging/root-cause-analysis)

**Row-level security**

Row-level security is a database feature that limits which rows each user can read or change, so one user's data stays hidden from another.

[Read the article: Common bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs)

**Rubber-stamp review**

A rubber-stamp review is a code review that approves a change without checking it closely, often because the change is large or the queue is long.

**Rules file backdoor**

A rules file backdoor is an attack that hides instructions in a coding agent's instruction or rules file. The agent then writes harmful code while appearing to follow the project's rules.

**Runtime verification**

Runtime verification is a method that checks a running program against stated properties by observing its behavior, usually through a monitor that reads its execution. Developers also use the term for running the software to check a change.

[Read the article: What is runtime verification?](https://specstory.com/learning/verification/runtime-verification)

## S

**Sandbox**

A sandbox is an isolated environment that limits which files and network hosts the code inside it can reach, so a bad command can damage only what the sandbox allows.

[Read the article: What is a sandbox for AI agents?](https://specstory.com/learning/environments/ai-sandbox)

**Sanity testing**

Sanity testing is a quick, narrow check that a specific fix or change behaves sensibly before deeper testing, often used interchangeably with smoke testing.

[Read the article: What is smoke testing?](https://specstory.com/learning/testing/smoke-testing)

**SAST**

Static application security testing (SAST) is a security practice that scans source code for vulnerabilities without running it. It is separate from functional testing.

**Secret scanning**

Secret scanning is a check that searches code and history for credentials, e.g. API keys, so they can be removed and rotated before someone misuses them.

**Self-healing test**

A self-healing test is a test that updates its own locators or steps when the app changes, which cuts upkeep but can also hide real breakage.

**Shift left**

Shift left is an approach that moves testing and other quality checks earlier in development, when a problem is smaller and cheaper to fix.

[Read the article: What is shift-left testing?](https://specstory.com/learning/ci-cd/shift-left-testing)

**Shrinking**

Shrinking is a step in property-based testing that reduces a failing generated input to the smallest input it can find that still fails, so the bug is easier to understand.

[Read the article: What is property-based testing?](https://specstory.com/learning/test-quality/property-based-testing)

**SIGPIPE**

SIGPIPE is the signal that a program receives when it writes to a pipe whose reader has closed. By default it ends the program, and shells report exit code 141.

[Read the article: What is a broken pipe error (SIGPIPE)?](https://specstory.com/learning/cli-and-web/broken-pipe)

**Skipped test**

A skipped test is a test that is marked not to run, which many tools still report as part of a passing suite, so it can hide a missing check.

**Slopsquatting**

Slopsquatting is an attack that registers package names that language models tend to invent, so code that installs a hallucinated package installs the attacker's code instead.

[Read the article: What are AI hallucinations in code?](https://specstory.com/learning/verification/ai-code-hallucinations)

**Smoke test**

A smoke test is a quick, shallow test that a build starts and its most basic functions work, run before deeper testing is worth the time.

[Read the article: What is smoke testing?](https://specstory.com/learning/testing/smoke-testing)

**Snapshot test**

A snapshot test is a test that saves a program's output once and fails when later output differs, until someone reviews and approves the new output.

[Read the article: What is snapshot testing?](https://specstory.com/learning/testing/snapshot-testing)

**Software regression**

A software regression is a defect in which something that worked before a change stops working after it, often in a part of the code the change did not touch.

[Read the article: What is regression testing?](https://specstory.com/learning/testing/regression-testing)

**Software specification**

A software specification is a written description that states what a program should do, its inputs and outputs, and its constraints, so people and tests can check the result.

**Software testing**

Software testing is the practice that runs or examines a program to find differences between what it does and what it should do, before users find them.

[Read the article: What is software testing?](https://specstory.com/learning/testing/software-testing)

**Spec drift**

Spec drift is the gap that opens when code and its specification change separately, so the written spec no longer describes what the software does.

**Spec-driven development**

Spec-driven development is a way of working that writes a specification of what software should do before a coding agent writes it. The spec is then used to guide and check the work.

[Read the article: What is spec-driven development?](https://specstory.com/learning/ai-coding/spec-driven-development)

**Specification gaming**

Specification gaming is a behavior in which an AI system satisfies the literal objective it was given without doing the task the objective was meant to capture.

[Read the article: What is reward hacking in coding agents?](https://specstory.com/learning/verification/reward-hacking)

**Stack trace**

A stack trace is a report that lists the chain of function calls active when an error happened, from the failing line back through the calls that led to it.

**Staging environment**

A staging environment is a production-like copy of an application that teams use for final testing before release, running the same build with similar settings but no real users.

[Read the article: What is a staging environment?](https://specstory.com/learning/environments/staging-environment)

**Standard streams**

The standard streams are the three channels that every process gets by default, which are standard input for data in, standard output for results, and standard error for diagnostics.

[Read the article: What is the difference between stdout and stderr?](https://specstory.com/learning/cli-and-web/stdout-vs-stderr)

**Static analysis**

Static analysis is a method that examines source code without running it, using rules and type information to flag likely bugs and security flaws.

[Read the article: What is static code analysis?](https://specstory.com/learning/code-review/static-analysis)

**Status code**

A status code is a number that a web server returns with each HTTP response to report the result, e.g. 404 for not found.

**stderr**

stderr is the standard error stream that a program writes errors and diagnostics to, so they stay separate from the results on standard output.

[Read the article: What is the difference between stdout and stderr?](https://specstory.com/learning/cli-and-web/stdout-vs-stderr)

**stdout**

stdout is the standard output stream that a program writes its normal results to, so another program, a file, or a terminal can read them.

[Read the article: What is the difference between stdout and stderr?](https://specstory.com/learning/cli-and-web/stdout-vs-stderr)

**Subagent**

A subagent is a coding agent that another agent starts to handle one part of a task in its own context, then returns a result to the parent.

**System testing**

System testing is a level of testing that checks the complete, integrated system against its requirements, after integration testing and before acceptance testing.

**System under test**

The system under test is the part of the software that a test exercises and checks, as opposed to the harness around it, including its test doubles.

[Read the article: What is a test harness?](https://specstory.com/learning/testing/test-harness)

## T

**Tautological test**

A tautological test is a test that cannot fail because its expected value comes from the same code, mock, or constant as the result it checks. It passes whether the code is right or wrong.

[Read the article: What is a tautological test?](https://specstory.com/learning/test-quality/tautological-test)

**Technical debt**

Technical debt is the future cost of shortcuts in code or design, which a team pays as slower changes and more bugs until someone cleans the shortcuts up.

[Read the article: What is technical debt, and do coding agents add to it?](https://specstory.com/learning/code-review/technical-debt)

**Test automation**

Test automation is the use of software to run tests and compare actual results with expected results, so the same checks can repeat without a person doing each step.

[Read the article: What is test automation?](https://specstory.com/learning/testing/test-automation)

**Test charter**

A test charter is a short mission statement for an exploratory testing session that names the area to explore, the risks to look for, and a time box.

[Read the article: What is exploratory testing, and can an AI agent do it?](https://specstory.com/learning/testing/exploratory-testing)

**Test data management**

Test data management is the practice that creates, refreshes, and cleans up the data tests need, so each run starts from a known state without real personal data.

**Test double**

A test double is a stand-in for a real component that a test uses in its place. Stubs return canned answers, and mocks also check that they were called as expected.

[Read the article: What is the difference between mocks, stubs, and fakes?](https://specstory.com/learning/testing/mocks-vs-stubs)

**Test fixture**

A test fixture is the fixed setup that a test needs before it runs, e.g. seeded data, together with the steps that clean it up.

**Test harness**

A test harness is the set of code and tools that runs tests against a system, including drivers that start it, test doubles for its dependencies, and checks on the results.

[Read the article: What is a test harness?](https://specstory.com/learning/testing/test-harness)

**Test impact analysis**

Test impact analysis is a technique that maps which tests exercise which code, so a change runs only the tests it could affect instead of the whole suite.

[Read the article: What is test impact analysis?](https://specstory.com/learning/ci-cd/test-impact-analysis)

**Test isolation**

Test isolation is a property of a test that runs with its own state and data, so its result does not depend on other tests, their order, or leftover files.

[Read the article: What is test isolation (hermetic tests)?](https://specstory.com/learning/environments/test-isolation)

**Test oracle**

A test oracle is the source of truth that a test compares a result against, e.g. a specification that states the correct output for a given input.

[Read the article: What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)

**Test oracle problem**

The test oracle problem is the difficulty of deciding whether a test passed or failed for a given input and state, which is hardest when nobody knows the correct answer.

[Read the article: What is a test oracle?](https://specstory.com/learning/test-quality/test-oracle)

**Test pyramid**

The test pyramid is a model of a test suite that has many fast unit tests at the base and fewer integration tests in the middle. A few end-to-end tests sit at the top.

[Read the article: What is the test pyramid, and does it hold for AI-written code?](https://specstory.com/learning/testing/test-pyramid)

**Test quarantine**

Test quarantine is a policy that moves a flaky test out of the blocking suite while it is fixed. The test still runs and reports but no longer fails the build.

[Read the article: What is a flaky test?](https://specstory.com/learning/debugging/flaky-tests)

**Test suite**

A test suite is a collection of tests that a project runs together to check its code, e.g. every unit test in a repository.

**Test tampering**

Test tampering is a coding agent behavior that edits, skips, or deletes a failing test so the suite passes, instead of fixing the code the test was checking.

[Read the article: Why do coding agents delete or weaken tests?](https://specstory.com/learning/verification/test-tampering)

**Test-driven development (TDD)**

Test-driven development is a practice that writes a failing test before the code, then only enough code to pass it, then refactors, in small repeated cycles.

[Read the article: What is test-driven development (TDD)?](https://specstory.com/learning/testing/test-driven-development)

**Test-order dependence**

Test-order dependence is a flaw that makes a test pass or fail depending on which tests ran before it, usually because they share state.

[Read the article: What is test isolation (hermetic tests)?](https://specstory.com/learning/environments/test-isolation)

**Testing in production**

Testing in production is a practice that checks software after release with real traffic, e.g. through monitoring, alongside testing before release.

**Tool calling**

Tool calling is a model capability that lets a language model request a named function with arguments, which the surrounding harness runs and returns results from.

**Tool-call spoofing**

Tool-call spoofing is an agent failure in which a transcript shows a tool call or tool output that did not happen, so the record no longer matches what ran.

**TTY**

A TTY is a terminal device that a program can detect, and many programs change their output, color, or prompts when they write to a pipe instead of a TTY.

[Read the article: Why does CLI output change when piped?](https://specstory.com/learning/cli-and-web/tty-vs-pipe)

**Two-account test**

A two-account test is a behavior check that signs in as two different users and confirms that neither can see or change the other's data.

**Type checking**

Type checking is a static check that verifies values are used according to their declared types, e.g. that a function expecting a number never receives a string.

## U

**Unit test**

A unit test is a test that checks the smallest testable part of a program, e.g. one function, in isolation from the rest of the system.

[Read the article: What is unit testing?](https://specstory.com/learning/testing/unit-testing)

**User journey test**

A user journey test is a test that follows one complete task a user performs, e.g. signing up, from start to finish through the running software, and checks each step.

[Read the article: What is a user journey test?](https://specstory.com/learning/testing/user-journey-testing)

## V

**Verification debt**

Verification debt is the gap that grows between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.

[Read the article: What is verification debt?](https://specstory.com/learning/verification/verification-debt)

**Vibe coding**

Vibe coding is a way of building software in which a person describes the result to an AI and accepts the code without reading it, judging it only by using the app.

[Read the article: What is vibe coding?](https://specstory.com/learning/ai-coding/vibe-coding)

---

Source: [Software testing and AI coding glossary | SpecStory](https://specstory.com/learning/glossary)
