Glossary
Short definitions of the terms used across the Learning Center. Each links to its full article when one exists.
A
- Acceptance criteria
Acceptance criteria are the conditions that a change must meet for the person who asked for it to accept it. They are written so that a person or a test can check each one.
- Acceptance testing
Acceptance testing is a level of testing that checks whether a system meets its acceptance criteria and its users' needs. It lets the system's users and owners decide whether to accept it.
Read the article: What is acceptance testing (UAT)?- Accessibility tree
The accessibility tree is the browser's structured view of a page's elements, names, and roles, which assistive technology and many browser agents read the page through.
- Adversarial testing
Adversarial testing is a testing approach that deliberately tries to break software with unexpected, hostile, or careless inputs and actions, instead of confirming the expected path.
Read the article: What is adversarial testing in software?- Agent harness
An agent harness is the software around a language model that gives it tools, context, memory, and a loop, which turns the model into a working agent.
Read the article: What is an agent harness?- Agent hook
An agent hook is a script that a coding agent's harness runs at a fixed point in its loop. A Stop hook runs when the agent tries to end its turn.
Read the article: What are agent hooks (Stop hooks)?- Agent loop
The agent loop is the cycle in which a model requests an action, the harness runs the tool, and the model reads the result. It ends when the model requests no tool, or when a step limit, a time limit, or a person stops it.
Read the article: What is a coding agent?- Agent trajectory
An agent trajectory is the ordered sequence of an agent's actions, tool calls, and observations while it works on a task, which is what a transcript records.
Read the article: What is AI coding session history?- Agentic engineering
Agentic engineering is a way of building software that directs coding agents while keeping engineering practices, from written specs and tests to code review and verification.
Read the article: What is agentic engineering?- Agentic testing
Agentic testing is software testing that AI agents carry out themselves, choosing, running, and judging checks on the software, rather than only writing test scripts for later.
Read the article: What is agentic testing?- AGENTS.md
AGENTS.md is a Markdown file in a repository that gives coding agents project instructions, e.g. how to build, test, and change the code.
Read the article: What is AGENTS.md?- AI agent
An AI agent is a program that uses a language model in a loop to choose actions, call tools, and read the results until it finishes a task.
- AI code attribution
AI code attribution is the practice that records which code an AI tool wrote and under which prompt, e.g. with commit trailers.
- AI code review
AI code review is a review practice that uses a language model to read a pull request or diff and comment on possible bugs, style problems, and risks.
Read the article: What is AI code review?- AI code verification
AI code verification is the practice that checks what AI-generated software does, by reading it, running its tests, and running the software, rather than detecting who wrote it.
Read the article: How to test AI-generated code- AI slop
AI slop is AI-generated output that looks finished but is bloated, duplicated, or poorly tested because nobody checked it closely. In code it includes verbose logic and tests that assert almost nothing.
Read the article: What is AI slop in code?- Approval fatigue
Approval fatigue is a failure in which people approve a coding agent's permission prompts without reading them, because the prompts are so frequent that they stop protecting anything.
Read the article: What is approval fatigue in coding agents?- Arrange-act-assert
Arrange-act-assert is a test layout that sets up the inputs first, then runs the code under test, and then checks the result, each in its own block.
- Assertion
An assertion is a statement in a test that checks a condition and fails the test when the condition is false.
- Assertion-free test
An assertion-free test is a test that runs code but checks no result, so it passes whenever the code does not crash, whether the output is right or wrong.
Read the article: Why do AI-written tests always pass?
B
- Background coding agent
A background coding agent is a coding agent that works on a task without a person watching and returns a branch or pull request when it finishes.
- Behavior-driven development (BDD)
Behavior-driven development is a practice that describes how software should behave as plain-language Given-When-Then scenarios that the team agrees on and tools can run as tests.
- Blameless postmortem
A blameless postmortem is an incident review that works out how a failure happened and how to prevent a repeat, without blaming the individuals involved.
Read the article: What is root cause analysis (RCA)?- Blast radius
The blast radius of a change is the set of features, users, and systems that could break because of it. That set is often much wider than the files in the diff.
- Branch coverage
Branch coverage is a code coverage measure that counts which outcomes of each decision point, e.g. both sides of an if statement, the tests have run.
Read the article: What is code coverage?- Browser test framework
A browser test framework is an open-source library that drives a real browser from test code, so a test can load pages, act on them, and check the results.
Read the article: What is a browser test framework?- Bug bash
A bug bash is a time-boxed session that gathers people from across a team to use the software together and report as many defects as they can find.
C
- Catching test
A catching test is a test that is generated for one code change and designed to fail if that change introduced a bug, then usually thrown away after review.
Read the article: What are catching tests (catching JiTTests)?- Characterization test
A characterization test is a test that records what existing code does now, right or wrong, so later changes can be checked against that recorded behavior.
Read the article: What is golden file testing (golden master)?- Chrome DevTools Protocol
The Chrome DevTools Protocol is an interface that lets a program drive and inspect a Chromium-based browser, and many automation tools and browser agents are built on it.
- CI pipeline
A CI pipeline is the ordered set of automated steps, e.g. build, test, and package, that runs each time code is pushed or merged.
Read the article: What is CI/CD?- CI/CD
CI/CD is a practice that merges code often and builds and tests each change in an automated pipeline, so the software is kept ready to release.
Read the article: What is CI/CD?- Code coverage
Code coverage is a test measure that reports the percentage of a program's code that runs during its tests, which shows what was executed but not what was checked.
Read the article: What is code coverage?- Code review
Code review is a practice that has people or tools read a proposed code change before it merges, to catch defects, share knowledge, and keep the codebase consistent.
Read the article: What is code review?- Coding agent
A coding agent is an AI system that edits files and runs commands in a loop to complete a programming task with little step-by-step direction.
Read the article: What is a coding agent?- Coding agent transcript
A coding agent transcript is a saved record that holds a coding session's prompts, replies, tool calls, and results. A collection of saved transcripts is called AI coding session history.
Read the article: What is AI coding session history?- Command-line interface (CLI)
A command-line interface is a way of using a program that takes typed commands, flags, and input in a terminal or script, and returns output, errors, and an exit code.
Read the article: How to test a command-line application- Commit
A commit is a saved snapshot of changes in a Git repository, recorded with a message, an author, and a link to the commit before it.
- Container
A container is a packaged process that runs with its own filesystem and limits but shares the host's kernel, which makes it lighter and less isolated than a virtual machine.
Read the article: What is the difference between containers, gVisor, and microVMs?- Context window
A context window is the amount of text, measured in tokens, that a language model can take into account at once, including instructions, files, and history.
Read the article: What is context engineering?- Continuous delivery
Continuous delivery is a practice that keeps every change that passes the pipeline ready to release, so a person can release it to production at any time with one decision.
- Continuous deployment
Continuous deployment is a practice that releases every change that passes the automated pipeline to production, with no manual approval step between the merge and the release.
- Continuous integration (CI)
Continuous integration is a practice that merges every developer's changes into the main branch often, with an automated build and test run on each merge.
Read the article: What is continuous integration (CI)?- Coverage-guided fuzzing
Coverage-guided fuzzing is a fuzzing method that keeps the inputs which reach new code paths and mutates them further, so the fuzzer explores a program efficiently.
Read the article: What is fuzzing?- Cross-model verification
Cross-model verification is a checking setup that asks a different AI model to review work that another model produced, in the hope that their blind spots differ.
Read the article: Can AI check its own code?
D
- DAST
Dynamic application security testing (DAST) is a security practice that probes a running application from the outside for vulnerabilities. It is separate from functional testing.
- Debugging
Debugging is the process that finds why software does the wrong thing, by reproducing the failure, narrowing down its cause, changing the code, and confirming the fix.
Read the article: What is debugging?- Definition of done
A definition of done is a shared checklist that states what evidence a piece of work must have before a team counts it as finished.
Read the article: What is a definition of done for coding agents?- Design for testability
Design for testability is a design approach that makes software easier to test, e.g. by giving a CLI clear exit codes and machine-readable output.
- Dev container
A dev container is a container defined in the repository that provides a ready-made development environment with its tools installed.
Read the article: What is the difference between Docker, dev containers, and VMs?- Dynamic analysis
Dynamic analysis is a method that observes a program while it runs, e.g. through tests, to find bugs that only appear at runtime.
Read the article: What is the difference between static and dynamic analysis?
E
- Egress control
Egress control is a network control that limits which outside hosts a sandbox or agent can reach, usually with an allowlist or a proxy.
- End-to-end test
An end-to-end test is a test that follows a whole user journey through a running system, from the interface through services and data, the way a real user would.
Read the article: What is end-to-end testing?- Environment parity
Environment parity is the degree to which development, test, staging, and production environments match, so a check in one predicts behavior in another.
- Ephemeral environment
An ephemeral environment is a short-lived copy of an application that is created for one change or test run and destroyed afterwards. A preview environment is one kind.
Read the article: What is an ephemeral environment?- Equivalent mutant
An equivalent mutant is a code change made during mutation testing that behaves exactly the same as the original program, so no test can ever detect or kill it.
- Eval
An eval, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results, so versions of the system can be compared.
Read the article: What is LLM-as-a-judge?- Exit code
An exit code is the number that a program returns when it finishes, where 0 usually means success and any other value signals an error that scripts can check.
Read the article: What are exit codes?- Exploratory testing
Exploratory testing is a testing approach that designs and runs tests at the same time, guided by what the tester learns about the software while using it.
Read the article: What is exploratory testing, and can an AI agent do it?
F
- Fail-open
Fail-open is a behavior that lets an action through when a check errors or times out, while fail-closed code denies it. Fail-open authentication is a common security bug.
- False completion claim
A false completion claim is a coding agent's report that a task is "done," or that tests pass, when the work has not been built, run, or checked.
Read the article: Why do coding agents say "done" when the code doesn't work?- False negative
A false negative is a test result that misses a defect that does exist, so the check passes even though the software is broken.
Read the article: What is a false pass (false green test)?- False pass
A false pass is a test or check result that reports success while the software is broken, because the check missed the defect or lost the failure.
Read the article: What is a false pass (false green test)?- False positive
A false positive is a check result that reports a problem that does not exist, which wastes time and teaches people to ignore real findings.
- Fault seeding
Fault seeding is a technique that deliberately inserts known bugs into code to see how many of them a test suite or review process detects.
Read the article: What is mutation testing?- Flaky test
A flaky test is a test that both passes and fails on the same code. A single result from it cannot be trusted as evidence of a bug or of a fix.
Read the article: What is a flaky test?- Formal verification
Formal verification is a method that proves mathematically that a program or model meets a specification, instead of testing it on sample inputs.
- Fuzzing
Fuzzing is a testing technique that feeds a program large numbers of random or malformed inputs to find crashes, hangs, and bugs in input handling.
Read the article: What is fuzzing?
G
- Ghost feature
A ghost feature is a feature that exists in the code and may pass its tests but never runs for users, e.g. a handler that nothing calls.
Read the article: What is a false pass (false green test)?- Git bisect
Git bisect is a Git command that runs a binary search through a repository's history to find the first commit where a behavior changed.
Read the article: What is git bisect?- Git worktree
A Git worktree is an extra working directory attached to the same repository, so several branches can be checked out at once, e.g. one per parallel agent.
- Given-when-then
Given-when-then is a scenario format that states a starting state, an action, and an expected outcome in plain language, used in behavior-driven development and acceptance criteria.
- Golden file
A golden file is a saved, approved copy of a program's output that a test compares new output against, failing when anything in the output changes.
Read the article: What is golden file testing (golden master)?- Goodhart's law
Goodhart's law is the observation that a measure which becomes a target stops being a good measure, because people and systems optimize the number instead of the goal.
H
- Happy path
The happy path is the expected route through a feature, with valid input and no errors, that most demos and many generated tests follow.
- Headless browser
A headless browser is a web browser that runs without a visible window and is controlled by code, so tests and agents can load pages and act on them.
Read the article: What is a headless browser?- Heisenbug
A heisenbug is a bug that disappears or changes when someone tries to observe it, e.g. when a debugger or extra logging changes the timing.
Read the article: What is a race condition, and how do you reproduce one?- Hermetic test
A hermetic test is a test that depends only on what it declares and sets up itself, with no shared state, network, or leftover files. It gives the same result anywhere.
Read the article: What is test isolation (hermetic tests)?- Hook
A hook is a script that a tool runs automatically at a set point in its workflow, e.g. Git running a check before it records a commit.
Read the article: What are Git hooks?- Human-in-the-loop
Human-in-the-loop is a way of running an AI system in which a person approves, corrects, or reviews its work at set points before the work takes effect.
Read the article: What is human-in-the-loop for coding agents?- Hydration error
A hydration error is a web app failure that happens when the HTML a server rendered does not match what the browser's JavaScript renders on load.
Read the article: What is a hydration error?
I
- Idempotency
Idempotency is a property of an operation that gives the same result whether it runs once or several times, e.g. a payment request that is safe to retry.
- Implicit oracle
An implicit oracle is a general failure signal that marks a result as wrong without any expected value, e.g. a crash.
Read the article: What is a test oracle?- Indirect prompt injection
Indirect prompt injection is a prompt injection that arrives through content an agent reads during a task, e.g. a web page, instead of from the user.
Read the article: What is prompt injection in coding agents?- Integration test
An integration test is a test that checks that separate parts of a system work together, e.g. a service and its database, rather than each part alone.
Read the article: What is integration testing?- Invariant
An invariant is a condition that must stay true for every valid state or input of a program, e.g. an order total that is never negative.
J
- Jailbreak
A jailbreak is a prompt that gets a language model to break its own safety rules. Prompt injection differs because it attacks an application built on the model, by mixing untrusted text into its instructions.
L
- Large language model (LLM)
A large language model is a model that is trained on large amounts of text to generate text, including code, from a prompt. Coding agents use one to choose what to do next.
- Least privilege
Least privilege is a security principle that gives a person, program, or agent only the access its task needs, for no longer than it needs it.
Read the article: What is least privilege for AI agents?- Lethal trifecta
The lethal trifecta is a risk pattern in which an agent has access to private data, reads untrusted content, and can send data out, which lets prompt injection leak data.
Read the article: What is prompt injection in coding agents?- Linting
Linting is a static check that compares source code against style and correctness rules with a tool called a linter, without running the code.
Read the article: What is linting?- LLM-as-a-judge
LLM-as-a-judge is an evaluation setup that uses a language model to grade output against written criteria, instead of or alongside deterministic checks.
Read the article: What is LLM-as-a-judge?- Load testing
Load testing is a performance test that runs software under the traffic expected in production, to measure response times and find where it slows down.
M
- MCP server
An MCP server is a program that exposes tools or data to AI agents over the Model Context Protocol, so an agent can call them in a standard way.
Read the article: What is the Model Context Protocol (MCP)?- Merge queue
A merge queue is a system that tests each approved pull request against the current main branch and the changes ahead of it, and merges only those that pass.
- Metamorphic testing
Metamorphic testing is a technique that checks how a program's output should change when its input changes in a known way, so code can be tested without an exact expected output.
Read the article: What is metamorphic testing?- MicroVM
A microVM is a small virtual machine that starts quickly and gives each workload its own kernel, which isolates it more strongly than a container.
- Minimal reproducible example
A minimal reproducible example is the smallest complete code, input, and setup that still shows a bug, so anyone can run it and see the same failure.
Read the article: What is a minimal reproducible example?- Model Context Protocol (MCP)
The Model Context Protocol is an open standard that lets AI applications connect to tools and data through MCP servers, so an agent can call them in a uniform way.
Read the article: What is the Model Context Protocol (MCP)?- Monkey testing
Monkey testing is a testing technique that feeds software random clicks, keystrokes, or inputs without a script, to see whether it crashes or misbehaves.
- Mutant
A mutant is a copy of a program with one small deliberate change, e.g. a flipped comparison, that mutation testing uses to check whether the tests fail.
Read the article: What is mutation testing?- Mutation score
Mutation score is a test-strength measure that reports the share of valid mutants a test suite detects, which shows how well the tests check behavior.
Read the article: What is mutation testing?- Mutation testing
Mutation testing is a technique that checks a test suite by injecting small deliberate bugs, called mutants, and counting how many of them the tests catch.
Read the article: What is mutation testing?
N
- Negative testing
Negative testing is a testing technique that checks how software handles invalid, unexpected, or missing input, instead of only the input it was designed for.
- Nondeterminism
Nondeterminism is a property of a program or test that can produce different results from the same code and input, e.g. because of timing.
O
- Observed vs. expected
Observed vs. expected is the part of a bug report that states what the software did next to what it should have done, as facts rather than guesses about causes.
Read the article: What are reproduction steps in a bug report?- Out-of-scope edit
An out-of-scope edit is a change that a coding agent makes outside what its task asked for, e.g. an unrequested refactor.
- OWASP Top 10
The OWASP Top 10 is a list that the Open Worldwide Application Security Project publishes of the most critical security risk categories for web applications.
P
- Package hallucination
Package hallucination is a code generation failure that imports or installs a package that does not exist. An attacker can register the invented name and publish harmful code under it.
Read the article: What are AI hallucinations in code?- Partial oracle
A partial oracle is a test oracle that checks some properties of a result when the full expected result is unknown, e.g. that an order total is never negative.
Read the article: What is a test oracle?- Pass on retry
A pass on retry is a test result that fails first and then passes when the same test runs again on the same code. It marks the test as flaky or the bug as intermittent.
Read the article: How to tell a flaky test from a real bug- Playwright MCP
Playwright MCP is an open-source server that lets an AI agent control a real browser through the Model Context Protocol, using structured page snapshots instead of screenshots.
Read the article: What is Playwright MCP?- Pre-commit hook
A pre-commit hook is a script that Git runs before it records a commit. It can lint, format, or test the staged changes and stop the commit if a check fails.
Read the article: What is a pre-commit hook?- Pre-push hook
A pre-push hook is a script that Git runs before it sends commits to a remote repository, which can run slower checks than a pre-commit hook and stop the push.
Read the article: What is the difference between pre-commit, pre-push, and CI checks?- Prompt injection
Prompt injection is an attack that hides instructions in content an AI model reads, e.g. a web page, so the model follows them instead of its user.
Read the article: What is prompt injection in coding agents?- Property-based testing
Property-based testing is a technique that states rules that should hold for every input, then checks them against many generated inputs and shrinks any failing input.
Read the article: What is property-based testing?- Pull request
A pull request is a request that asks to merge a set of commits into another branch, where people and tools review and check the change first.
Q
- Quality assurance (QA)
Quality assurance is a set of activities that aims to prevent defects by improving how a team builds and tests software. Testing, a form of quality control, checks the product itself.
Read the article: What is QA testing?- Quality gate
A quality gate is a point in a delivery pipeline where a change is measured against set criteria before it goes further. Some gates block the change, and others only report.
Read the article: What is a quality gate in CI/CD?
R
- Race condition
A race condition is a bug that makes a result depend on the timing or order of events running at the same time, e.g. two requests changing the same record.
Read the article: What is a race condition, and how do you reproduce one?- Red teaming
Red teaming is a practice that has a team act as an attacker or hostile user to find weaknesses in a system before real attackers or users do.
- Red-green-refactor
Red-green-refactor is the test-driven development cycle that writes a failing test, then the code that makes it pass, then cleans up the code while the test stays green.
Read the article: What is test-driven development (TDD)?- Refactoring
Refactoring is a practice that changes the structure of existing code without changing its behavior, e.g. to remove duplication, so the code is easier to read and change.
- Regression testing
Regression testing is a type of testing that checks that a change has not broken behavior that worked before, including parts of the software the change did not touch.
Read the article: What is regression testing?- Reproduction script
A reproduction script is a runnable script that replays the steps that trigger a bug, so a person or an agent can reproduce the failure without guessing.
Read the article: What are reproduction steps in a bug report?- Reproduction steps
Reproduction steps are the exact actions, inputs, and conditions that make a bug happen again, written so that a person or a coding agent can follow them.
Read the article: What are reproduction steps in a bug report?- Retesting
Retesting is a check that repeats the exact steps that exposed a defect on the fixed code, to confirm that the failure no longer happens. It is often called fix verification.
Read the article: How to verify a bug fix- Review bottleneck
The review bottleneck is the queue that forms when code is produced faster than people can review it, so pull requests wait or get approved with less scrutiny.
Read the article: What is the code review bottleneck?- Reward hacking
Reward hacking is a behavior in which a coding agent satisfies the check it is graded on, e.g. the tests, without doing the task the user intended.
Read the article: What is reward hacking in coding agents?- Root cause analysis (RCA)
Root cause analysis is a method that traces a failure past its symptoms to the underlying cause, so the fix removes the reason the failure happened.
Read the article: What is root cause analysis (RCA)?- Row-level security
Row-level security is a database feature that limits which rows each user can read or change, so one user's data stays hidden from another.
Read the article: Common bugs in AI-generated code- Rubber-stamp review
A rubber-stamp review is a code review that approves a change without checking it closely, often because the change is large or the queue is long.
- Rules file backdoor
A rules file backdoor is an attack that hides instructions in a coding agent's instruction or rules file. The agent then writes harmful code while appearing to follow the project's rules.
- Runtime verification
Runtime verification is a method that checks a running program against stated properties by observing its behavior, usually through a monitor that reads its execution. Developers also use the term for running the software to check a change.
Read the article: What is runtime verification?
S
- Sandbox
A sandbox is an isolated environment that limits which files and network hosts the code inside it can reach, so a bad command can damage only what the sandbox allows.
Read the article: What is a sandbox for AI agents?- Sanity testing
Sanity testing is a quick, narrow check that a specific fix or change behaves sensibly before deeper testing, often used interchangeably with smoke testing.
Read the article: What is smoke testing?- SAST
Static application security testing (SAST) is a security practice that scans source code for vulnerabilities without running it. It is separate from functional testing.
- Secret scanning
Secret scanning is a check that searches code and history for credentials, e.g. API keys, so they can be removed and rotated before someone misuses them.
- Self-healing test
A self-healing test is a test that updates its own locators or steps when the app changes, which cuts upkeep but can also hide real breakage.
- Shift left
Shift left is an approach that moves testing and other quality checks earlier in development, when a problem is smaller and cheaper to fix.
Read the article: What is shift-left testing?- Shrinking
Shrinking is a step in property-based testing that reduces a failing generated input to the smallest input it can find that still fails, so the bug is easier to understand.
Read the article: What is property-based testing?- SIGPIPE
SIGPIPE is the signal that a program receives when it writes to a pipe whose reader has closed. By default it ends the program, and shells report exit code 141.
Read the article: What is a broken pipe error (SIGPIPE)?- Skipped test
A skipped test is a test that is marked not to run, which many tools still report as part of a passing suite, so it can hide a missing check.
- Slopsquatting
Slopsquatting is an attack that registers package names that language models tend to invent, so code that installs a hallucinated package installs the attacker's code instead.
Read the article: What are AI hallucinations in code?- Smoke test
A smoke test is a quick, shallow test that a build starts and its most basic functions work, run before deeper testing is worth the time.
Read the article: What is smoke testing?- Snapshot test
A snapshot test is a test that saves a program's output once and fails when later output differs, until someone reviews and approves the new output.
Read the article: What is snapshot testing?- Software regression
A software regression is a defect in which something that worked before a change stops working after it, often in a part of the code the change did not touch.
Read the article: What is regression testing?- Software specification
A software specification is a written description that states what a program should do, its inputs and outputs, and its constraints, so people and tests can check the result.
- Software testing
Software testing is the practice that runs or examines a program to find differences between what it does and what it should do, before users find them.
Read the article: What is software testing?- Spec drift
Spec drift is the gap that opens when code and its specification change separately, so the written spec no longer describes what the software does.
- Spec-driven development
Spec-driven development is a way of working that writes a specification of what software should do before a coding agent writes it. The spec is then used to guide and check the work.
Read the article: What is spec-driven development?- Specification gaming
Specification gaming is a behavior in which an AI system satisfies the literal objective it was given without doing the task the objective was meant to capture.
Read the article: What is reward hacking in coding agents?- Stack trace
A stack trace is a report that lists the chain of function calls active when an error happened, from the failing line back through the calls that led to it.
- Staging environment
A staging environment is a production-like copy of an application that teams use for final testing before release, running the same build with similar settings but no real users.
Read the article: What is a staging environment?- Standard streams
The standard streams are the three channels that every process gets by default, which are standard input for data in, standard output for results, and standard error for diagnostics.
Read the article: What is the difference between stdout and stderr?- Static analysis
Static analysis is a method that examines source code without running it, using rules and type information to flag likely bugs and security flaws.
Read the article: What is static code analysis?- Status code
A status code is a number that a web server returns with each HTTP response to report the result, e.g. 404 for not found.
- stderr
stderr is the standard error stream that a program writes errors and diagnostics to, so they stay separate from the results on standard output.
Read the article: What is the difference between stdout and stderr?- stdout
stdout is the standard output stream that a program writes its normal results to, so another program, a file, or a terminal can read them.
Read the article: What is the difference between stdout and stderr?- Subagent
A subagent is a coding agent that another agent starts to handle one part of a task in its own context, then returns a result to the parent.
- System testing
System testing is a level of testing that checks the complete, integrated system against its requirements, after integration testing and before acceptance testing.
- System under test
The system under test is the part of the software that a test exercises and checks, as opposed to the harness around it, including its test doubles.
Read the article: What is a test harness?
T
- Tautological test
A tautological test is a test that cannot fail because its expected value comes from the same code, mock, or constant as the result it checks. It passes whether the code is right or wrong.
Read the article: What is a tautological test?- Technical debt
Technical debt is the future cost of shortcuts in code or design, which a team pays as slower changes and more bugs until someone cleans the shortcuts up.
Read the article: What is technical debt, and do coding agents add to it?- Test automation
Test automation is the use of software to run tests and compare actual results with expected results, so the same checks can repeat without a person doing each step.
Read the article: What is test automation?- Test charter
A test charter is a short mission statement for an exploratory testing session that names the area to explore, the risks to look for, and a time box.
Read the article: What is exploratory testing, and can an AI agent do it?- Test data management
Test data management is the practice that creates, refreshes, and cleans up the data tests need, so each run starts from a known state without real personal data.
- Test double
A test double is a stand-in for a real component that a test uses in its place. Stubs return canned answers, and mocks also check that they were called as expected.
Read the article: What is the difference between mocks, stubs, and fakes?- Test fixture
A test fixture is the fixed setup that a test needs before it runs, e.g. seeded data, together with the steps that clean it up.
- Test harness
A test harness is the set of code and tools that runs tests against a system, including drivers that start it, test doubles for its dependencies, and checks on the results.
Read the article: What is a test harness?- Test impact analysis
Test impact analysis is a technique that maps which tests exercise which code, so a change runs only the tests it could affect instead of the whole suite.
Read the article: What is test impact analysis?- Test isolation
Test isolation is a property of a test that runs with its own state and data, so its result does not depend on other tests, their order, or leftover files.
Read the article: What is test isolation (hermetic tests)?- Test oracle
A test oracle is the source of truth that a test compares a result against, e.g. a specification that states the correct output for a given input.
Read the article: What is a test oracle?- Test oracle problem
The test oracle problem is the difficulty of deciding whether a test passed or failed for a given input and state, which is hardest when nobody knows the correct answer.
Read the article: What is a test oracle?- Test pyramid
The test pyramid is a model of a test suite that has many fast unit tests at the base and fewer integration tests in the middle. A few end-to-end tests sit at the top.
Read the article: What is the test pyramid, and does it hold for AI-written code?- Test quarantine
Test quarantine is a policy that moves a flaky test out of the blocking suite while it is fixed. The test still runs and reports but no longer fails the build.
Read the article: What is a flaky test?- Test suite
A test suite is a collection of tests that a project runs together to check its code, e.g. every unit test in a repository.
- Test tampering
Test tampering is a coding agent behavior that edits, skips, or deletes a failing test so the suite passes, instead of fixing the code the test was checking.
Read the article: Why do coding agents delete or weaken tests?- Test-driven development (TDD)
Test-driven development is a practice that writes a failing test before the code, then only enough code to pass it, then refactors, in small repeated cycles.
Read the article: What is test-driven development (TDD)?- Test-order dependence
Test-order dependence is a flaw that makes a test pass or fail depending on which tests ran before it, usually because they share state.
Read the article: What is test isolation (hermetic tests)?- Testing in production
Testing in production is a practice that checks software after release with real traffic, e.g. through monitoring, alongside testing before release.
- Tool calling
Tool calling is a model capability that lets a language model request a named function with arguments, which the surrounding harness runs and returns results from.
- Tool-call spoofing
Tool-call spoofing is an agent failure in which a transcript shows a tool call or tool output that did not happen, so the record no longer matches what ran.
- TTY
A TTY is a terminal device that a program can detect, and many programs change their output, color, or prompts when they write to a pipe instead of a TTY.
Read the article: Why does CLI output change when piped?- Two-account test
A two-account test is a behavior check that signs in as two different users and confirms that neither can see or change the other's data.
- Type checking
Type checking is a static check that verifies values are used according to their declared types, e.g. that a function expecting a number never receives a string.
U
- Unit test
A unit test is a test that checks the smallest testable part of a program, e.g. one function, in isolation from the rest of the system.
Read the article: What is unit testing?- User journey test
A user journey test is a test that follows one complete task a user performs, e.g. signing up, from start to finish through the running software, and checks each step.
Read the article: What is a user journey test?
V
- Verification debt
Verification debt is the gap that grows between how much AI-written code a team produces and how much of it anyone has checked by running or testing it.
Read the article: What is verification debt?- Vibe coding
Vibe coding is a way of building software in which a person describes the result to an AI and accepts the code without reading it, judging it only by using the app.
Read the article: What is vibe coding?