Key points
- Run the built program as a separate process, the way people and scripts call it.
- Check the exit code, stdout, stderr, and files, because each one is part of the contract.
- Test piped output and terminal output separately, because many programs behave differently in each.
How do you test a command-line application?
The usual way to test a command-line application is to run the built program with known input and check its exit code and output. A command-line interface (CLI) is a way of using a program that takes typed commands, flags, and input, and returns output, errors, and an exit code. Those results form the contract that scripts rely on.
Many CLI bugs sit at the edge of the program, where it parses arguments, reads input, writes output, and exits. A unit test that calls a function directly skips that edge. So a useful suite also runs the built program for each behavior that callers depend on, e.g. an error when an input file is missing.
The method works in any language, because the test controls a process, not the code inside it. A common mix is many fast tests of the program's functions and fewer tests that run the whole program.
What should a CLI test check?
A CLI's contract is everything a caller can observe and depend on. A test suite checks these parts of it:
- Exit code. The program returns 0 on success and a nonzero code on failure, so a script can branch on the result.
- Standard output. The result, and only the result, goes to stdout, where a pipe or a file can capture it.
- Standard error. Error and progress messages go to stderr, so they reach a person without corrupting piped output.
- Arguments. Documented flags parse as described, and an unknown flag fails with a usage message and a nonzero exit code.
- Terminal and pipe behavior. Colors and prompts appear only when a stream is a terminal, and the program stops quietly when a pipe closes early.
- Side effects. Files land where the flags say, and an interrupted run leaves no partial files behind.
- Signals and time. The program stops soon after Ctrl+C and finishes or fails within a set time instead of hanging.
The Command Line Interface Guidelines describe most of these conventions and state that "Exit codes are how scripts determine whether a program succeeded or failed."
How does black-box CLI testing work?
A black-box test runs the CLI the way a script does:
- The test builds the program or finds the built binary.
- The test creates a temporary directory and sets the environment variables the run needs, e.g.
NO_COLOR=1. - The test starts the program as a child process with its arguments, writes any input to its stdin, and closes stdin.
- The test waits for the process to exit and stops it after a timeout.
- The test records the exit code, stdout, and stderr separately.
- The test compares each result with its expected value and checks any files the program wrote.
In black-box testing, the test uses only what a caller can observe. The test never imports the program's code, so the two can use different languages, e.g. a pytest suite that runs a Go binary. A black-box CLI test is an end-to-end test for a program whose interface is text.
Most languages have a library for this:
| Language | Library | How a test checks the result |
|---|---|---|
| Go | testscript | exec runs the program, and stdout and stderr match its output to patterns. |
| Rust | assert_cmd | Command::cargo_bin finds the binary, and assert() checks exit code and output. |
| Python | subprocess with pytest | subprocess.run with capture_output=True returns returncode, stdout, and stderr. |
| Shell | Bats | run sets $status and puts both streams in $output, or stderr in $stderr with --separate-stderr. |
| Node.js | node:child_process with Jest or Vitest | spawnSync returns the exit status, stdout, and stderr. |
What are the types of CLI tests?
A CLI test suite can draw on five kinds of test.
In-process tests
These tests call the program's entry point inside the test runner, with arguments, input, and output streams passed in as parameters. In Go, os.Exit ends the process at once and skips deferred functions, so a testable main only calls os.Exit(run(os.Args[1:], os.Stdin, os.Stdout, os.Stderr)). When run parses flags with its own flag.FlagSet set to flag.ContinueOnError, tests can call it with any argument list and buffers, a form of design for testability that needs no mocks. In Python, Click's CliRunner calls a command in the test process, and pytest's capture fixtures collect what other code writes to stdout and stderr.
Process tests
Process tests start the built binary as a child process and check what it returns. They are slower, but they exercise the real argument parsing, the real streams, and the exit code that a shell receives. A few per command can catch the wiring bugs that faster tests miss.
Golden file tests
Golden file testing compares a command's full output with an approved copy saved in a file and fails on any difference. It suits long output, e.g. help text, where one assertion per line would be brittle. A snapshot test library manages the saved copies and rewrites them on request, so each rewrite needs a review.
Terminal tests
Terminal tests run the program under a pseudoterminal (PTY), which gives the program a terminal device instead of a pipe. A test that runs through a pipe checks only the pipe behavior, so prompts and colors need a PTY test, e.g. through Python's pty module.
Smoke tests
A smoke test runs the installed program once or twice to check that it starts, e.g. with acme --version. It catches a broken build or a package with a missing file before deeper tests run.
What is an example of a CLI test?
Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, acme, exports orders. A developer at Acme asks a coding agent to "Add an --output flag that writes the export to a file." The agent adds the flag and a write_export function. Its unit test calls write_export with a temporary path, checks the file, and passes.
The developer then runs the built command the way a nightly job would:
acme export --format csv --output orders.csv
echo $?
ls orders.csv
order_id,total,delivery_address
A-1042,$240.00,12 Elm Street
0
ls: cannot access 'orders.csv': No such file or directory
The flag parses, but the command never passes its value to write_export. The rows go to stdout, the exit code is 0, and no file exists. A black-box test in Bats runs the built command and checks each result:
@test "export --output writes the file and prints nothing" {
out="$BATS_TEST_TMPDIR/orders.csv"
run acme export --format csv --output "$out"
[ "$status" -eq 0 ]
[ -z "$output" ]
[ "$(head -n 1 "$out")" = "order_id,total,delivery_address" ]
}
The test fails on the empty output check. The agent passes the flag's value to write_export, and the test passes. This example is simplified. A real CLI would need more checks, e.g. one that an interrupted export leaves no partial file.
What changes when a coding agent writes the code?
A coding agent that writes a command can stop at unit tests that call its own functions. Those tests pass when the logic is right. They skip the built binary, so they can miss a flag that never reaches the code or a zero exit code after an error. The agent's output then says the task is "done" while a script that calls the command fails.
An agent can also change what counts as correct. When a golden file test fails, the agent can rerun it with the flag that rewrites the saved output, so the broken output becomes the expected result.
A practical adjustment is a short suite of black-box tests that run the built binary. The suite lives in files that the agent's permission settings block it from editing, e.g. a deny rule for tests/cli/, and a person reviews each golden file change. A failing test's command, input, and output are reproduction steps the agent can rerun after its fix. Testing any AI-generated code follows the same rule of checking what the software does when it runs.
What are the limits of CLI tests?
CLI tests have four limits that a team can plan for:
- They check only the cases they run. A command with many flags has more combinations than a suite can run, so teams test the combinations that callers use most.
- The environment differs. Tests run in a clean temporary directory with set variables, so a bug that appears only with a user's own config file can pass the suite.
- Golden files keep accidental details. A golden file records whatever the output was, e.g. the order of keys in JSON output, so a harmless change can fail the test.
- Failures show symptoms. A black-box test reports what a caller observed, not which line of code caused it.
A passing suite shows that the tested cases worked. No findings is not the same as complete coverage.
How is testing a CLI different from testing a web app?
A CLI test talks to one process through arguments, text streams, and an exit code. A web app test drives a browser, usually a headless browser in continuous integration, and checks a page that renders over time and depends on the network. That makes web app tests slower and more prone to timing failures. Both are end-to-end checks of running software, and some web apps also ship a CLI, e.g. for setup, so a team can need both.
What does CLI and web app testing cover?
The other CLI and web app articles each cover one check or one failure:
- Exit codes report whether a finished program succeeded.
- Stdout and stderr carry a program's results and its errors.
- TTY detection can change a program's output in a pipe.
- A broken pipe error comes from writing into a pipe with no reader.
- Shell scripts are tested with ShellCheck and Bats.
- Terminal UIs are tested in a pseudoterminal against a screen snapshot.
- A CLI that a coding agent built is verified from a fresh install.
- Headless browsers run with no window, under a program's control.
- Browser test frameworks let test code drive a real browser.
- Playwright MCP lets an agent operate a browser through tools.
- Hydration errors happen when server HTML differs from the browser's render.
How do you run CLI checks on every agent change?
After each agent change, run the built binary with the arguments your users and scripts use, and keep the command, input, output, and exit code as evidence.
A CLI can pass its unit tests and still fail when it runs as a real process. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps and tests them in isolated sandboxes on Linux virtual machines. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result.
FAQs
What is the best way to test command-line tools?
The best way to test command-line tools is to run the built program as a separate process with known input, then check its exit code, stdout, stderr, and files. Tests that call the entry point directly add fast checks of the logic.
How do you test a CLI without mocks everywhere?
A CLI can be tested without mocks everywhere when its entry point takes arguments, input, and output streams as parameters. Tests pass in buffers and read what the program wrote, and a few black-box tests run the built binary.
How do you test code that reads stdin?
Code that reads stdin can be tested by passing the input stream in as a parameter, so a test can hand it a buffer that holds the input. A black-box test writes the input to the process's stdin instead.
How do you test code that calls os.Exit?
Code that calls os.Exit is hard to test directly, because os.Exit ends the whole test process and skips deferred functions. A common fix is to call os.Exit only in main, with a run function that returns the exit code for tests to check.
Is there a framework for end-to-end tests of a CLI?
Most languages have a library for end-to-end tests of a CLI, e.g. testscript for Go. Any test runner can also start the program as a child process and check its exit code, stdout, and stderr, so a dedicated framework is optional.