Skip to content

How to verify a CLI that a coding agent built

Verifying a CLI that a coding agent built means installing it fresh, running the commands a user would, and checking exit codes, output, and files.

Last updated , 9 min read

Key points

  • Install the CLI from a fresh clone on a clean machine, not in the agent's working copy.
  • Run each command that the request, the README, and the help text describe, exactly as written.
  • Check the exit code, both output streams, and files of each command, then close a pipe and interrupt a run.

How do you verify a CLI that a coding agent built?

To verify a command-line interface (CLI) that a coding agent built, install it from a fresh clone on a clean machine. Then run the commands a user would, and check each exit code, output stream, and file on its own. Finally, pipe the output to a reader that exits early, and interrupt a long run.

A coding agent usually tests its work in the shell where it wrote the code, among packages left from earlier attempts. Its tools also often read output through a pipe, not a terminal. So its tests can pass while the installed tool fails on another machine or in a terminal.

Running the software is the rule for testing any AI-generated code, and the checks work in any language.

What do you need before you start?

A verification run needs five things:

  • A clean machine. A fresh virtual machine is best. At minimum, use an empty temporary directory and a fresh package environment, e.g. a Python virtual environment.
  • The agent's branch. A fresh clone checks out only the files committed to the branch, so it exposes a file that the agent never committed.
  • The tool's promises. Each example command in the request, the README, and the --help output is a check to run.
  • An empty home directory. Pointing HOME at an empty directory hides the tester's own config files, which is one part of test isolation.
  • A terminal and a pipe. A CLI can change its output when it writes to a TTY, so checks need both.

How do you verify a CLI step by step?

Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, acme, exports orders. A developer at Acme asks a coding agent to "Add JSON output to acme export." The agent adds a json value to --format, a test of its to_json function, and a README example. Its output says the task is "done."

1. Install from a fresh clone

The developer clones the agent's branch into an empty directory, installs it, and runs the JSON export:

tmp=$(mktemp -d)
git clone --quiet --branch json-export ~/src/acme-cli "$tmp/src"
python3 -m venv "$tmp/venv"
"$tmp/venv/bin/pip" install --quiet "$tmp/src"
mkdir "$tmp/home"
HOME="$tmp/home" "$tmp/venv/bin/acme" export --format json

The install succeeds, but the command fails:

Traceback (most recent call last):
  ...
ModuleNotFoundError: No module named 'orjson'

The agent's environment had orjson from an earlier task, and pyproject.toml never lists it. The developer installs it by hand so the other checks can run.

2. List the commands a user would run

The list comes from the request, the README, and acme export --help, plus one invalid value, one pipe, one terminal run, and one environment variable:

acme export --format csv
acme export --format json
acme export --format json --pretty
acme export --format json --output orders.json
acme export --format xml
acme export --format json | head -n 1
script -qec "acme export --format json" /dev/null
ACME_FORMAT=json acme export --format csv

Each line gets its expected result before it runs.

3. Run each command as written

Each line runs with empty input and a time limit. The developer captures stdout and stderr separately and checks the exit code. The JSON export passes. The README example fails:

$ acme export --format json --pretty
usage: acme [-h] {export} ...
acme: error: unrecognized arguments: --pretty
$ echo $?
2

The agent documented --pretty but never added it to the parser. The run with the invalid xml value exits with code 2 as expected. Under the Linux script command, which gives the export a terminal, the JSON output matches the piped output.

4. Close the pipe early

The pipe check prints the first order, then a crash on stderr:

$ acme export --format json | head -n 1
{"order_id":"A-1042","total":"$240.00"}
Traceback (most recent call last):
  ...
BrokenPipeError: [Errno 32] Broken pipe

head exits after one line, and the export is larger than the pipe's buffer, so its next write fails with a broken pipe. The agent's test called to_json directly, so it never wrote to a closed pipe.

5. Set one environment variable at a time

With HOME pointing at an empty directory, ACME_FORMAT=json acme export --format csv prints CSV, so the flag overrides the variable. The Command Line Interface Guidelines rank flags first, then environment variables, then config files.

6. Interrupt a long run

In a terminal, the developer starts a large export with --output orders.json and presses Ctrl+C. The command stops, but orders.json holds part of the orders and looks complete to the next script.

7. Send the failures back and rerun the list

The developer sends the agent the 4 failures, each with its command, exit code, and stderr text. After the fixes, the whole list runs again from a fresh clone, the same way a team would verify a bug fix.

This example is simplified. A real CLI would need a longer list, e.g. one line per subcommand.

What do agents get wrong in CLIs?

Many agent mistakes in CLIs come from where the agent ran its checks. The steps above target most of them:

  • Undeclared or invented dependencies. The code imports a package that the agent's environment already had, or one that does not exist. A fresh install and one run expose both.
  • Flags that exist only in the docs. The README shows a flag that the parser lacks, or the code calls another tool with an option it lacks. These hallucinations surface only when the command runs.
  • Failures that exit 0. An error handler prints a message and returns normally, so scripts read the zero exit code as success.
  • Output shaped for one mode. Color codes can leak into piped output, or a prompt can hang a script. Code that runs only in a terminal, e.g. a progress bar, goes unchecked, because agent tools often read a pipe.
  • Unhandled pipes and interrupts. Writing to a closed pipe can print a crash, and Ctrl+C can leave a partial file behind.
  • Golden files that follow the bug. When an output check fails, the agent can rewrite the saved output to match. A golden file or snapshot test then records the bug as expected.

In a study of 576,000 code samples from 16 models, at least 5.2% of package references from commercial models, and 21.7% from open-source models, named nonexistent packages. A clean install catches a missing package, but not one that an attacker registered under the invented name, so run the install on a machine without secrets.

What are common mistakes?

These mistakes let a broken CLI pass verification:

  • Checking in the agent's working copy. The working copy can hold untracked files, and the agent's environment can hold packages that users lack.
  • Trusting the summary. A message that says "done" is text the agent generated, not a record of a run.
  • Checking one result per command. A command can print correct output and exit with the wrong code, so check each result separately.

How do you check that it worked?

The verification is complete when these statements are true:

  • A fresh clone installs and runs in a clean environment with an empty home directory.
  • Each command from the request, the README, and the help text runs as written.
  • Each failure path exits with a nonzero code and writes its message to stderr, and stdout holds only results.
  • The tool stops without a crash when a reader closes the pipe, and Ctrl+C leaves no partial files.
  • The fixed failures pass on a fresh install, and the commands that passed before still pass.

Keep the list as a shell script that checks each line's exit code and output, so it runs against the next agent change. Undo one fix and confirm that the script fails. When other scripts parse a command's output, add a golden file for that command, and have a person review each update. These checks show that the listed commands worked on one clean machine. No findings is not the same as complete coverage.

How does RunStory help with CLIs a coding agent built?

A CLI can pass the tests its agent wrote and still fail on a clean machine. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

How do you test a CLI on a clean machine?

A CLI is tested on a clean machine by cloning its branch into a fresh environment, e.g. a virtual machine, and installing it as a user would. Then point the HOME variable at an empty directory and run the documented commands.

Can an agent invent a flag that does not exist?

A coding agent can invent a flag that does not exist, in README examples or in calls to other tools. Running each documented command exactly as written and checking its result exposes it, whether the parser rejects the flag or ignores it.

How do you test a CLI that other scripts depend on?

A CLI that other scripts depend on needs checks of what those scripts read. Test the exit code of each failure path, keep results on stdout and messages on stderr, and save a reviewed golden file of the parsed output.

How do you test a CLI's config files and environment variables?

A CLI's config files and environment variables are tested from an empty home directory, adding one setting at a time. Each check confirms the documented precedence, usually a flag over an environment variable and an environment variable over a config file.