Key points
- Pass most answers through the input stream, and use a pseudoterminal for prompts that read or check the terminal.
- Wait for the exact text of each prompt before sending its answer, and give each wait a deadline.
- Check the output, files, and exit code after the answers, and test each flag that skips a prompt.
How do you test interactive CLI prompts?
To test an interactive prompt in a command-line interface (CLI), send it scripted answers and check what the program does next. Most tests pass the answers in as the input stream. A few run the program in a pseudoterminal and wait for each question, with a deadline, before answering. Then check the output, files, and exit code.
An interactive prompt is a question that a CLI prints before it waits for a typed answer, e.g. Delete 3 orders? [y/N]. A test can feed that answer in four ways:
| Method | How answers arrive | What it misses |
|---|---|---|
| Injected reader | The test passes a buffer as the input stream | The terminal and the built program |
| Pipe to stdin | The test writes every answer to stdin at once | Prompts that read or check the terminal |
| Expect driver | The test answers each prompt in a pseudoterminal as it appears | Screens that redraw |
| Skip flag | A flag supplies the answer, so no prompt runs | The prompt itself |
An injected reader is a form of design for testability, e.g. Go's fmt.Fscan reading an io.Reader from the test.
What do you need before you start?
Gather these before writing the first test:
- A prompt script. List the exact text of each prompt in order, with its answer and the expected result.
- An expect driver. pexpect runs a program in a pseudoterminal from Python, on Linux and macOS but not Windows. Other languages have libraries that do the same, e.g.
node-ptyfor Node.js. - A deadline for each prompt. The steps use 5 seconds, which leaves room for a slow continuous integration (CI) runner.
- An empty home directory. Point
HOMEat a temporary directory, so the login test writes its files there instead of in the developer's own settings.
How do you test interactive CLI prompts step by step?
Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Add an acme login command that asks for an email and an API token." The agent uses Python's input() for both answers. Its passing test replaces input() with a stub that returns two strings, so the developer adds an expect test.
1. Start the command in a pseudoterminal
import os
import pexpect
def test_login_prompts(tmp_path):
env = {**os.environ, "HOME": str(tmp_path)}
child = pexpect.spawn("acme login", env=env, encoding="utf-8", timeout=5)
child.expect_exact("Email: ")
spawn gives the command a pseudoterminal, so a TTY check in the program returns true. expect_exact waits up to 5 seconds for the exact prompt text.
2. Answer each prompt after it appears
child.sendline("dev@example.com")
child.expect_exact("API token: ")
child.sendline("TEST_TOKEN")
child.expect_exact("Logged in as dev@example.com")
assert "TEST_TOKEN" not in child.before
child.before holds the output between the previous match and this one, so the last line checks that the terminal did not echo the token.
3. Fix the echo that the test found
The test fails:
> assert "TEST_TOKEN" not in child.before
E AssertionError: assert 'TEST_TOKEN' not in 'TEST_TOKEN\r\n'
The terminal echoed the token, because input() leaves echo on, and the agent's stub never used a terminal. For the token, the agent switches to Python's getpass, which turns echo off before printing the prompt, and the test passes.
4. Check the exit code and the files
child.expect(pexpect.EOF)
child.close()
assert child.exitstatus == 0
saved = (tmp_path / ".acme" / "credentials").read_text()
assert "TEST_TOKEN" in saved
close stores the exit code, and the last line checks that the login saved the token.
5. Test the runs without prompts
A script has no person to answer, so the developer asks for --email and --token-stdin flags. The token uses stdin, because a flag value can show up in the process list and shell history. Two scripted runs check both paths:
tmp_home="$(mktemp -d)"
printf 'TEST_TOKEN\n' | HOME="$tmp_home" acme login \
--email dev@example.com --token-stdin
echo "exit=$?"
HOME="$tmp_home" acme login < /dev/null
echo "exit=$?"
Logged in as dev@example.com
exit=0
acme: stdin is not a terminal; pass --email and --token-stdin
exit=2
The second run exits at once with a nonzero code and names the flags on stderr. This example is simplified. A real login also needs a test for a rejected token.
How does expect testing work?
Expect testing drives an interactive program through a pseudoterminal, alternating between waiting for output and sending input. This name comes from Expect, a Tcl tool for scripting interactive programs. An expect test runs in five steps:
- The driver starts the program with one end of a pseudoterminal pair as its terminal.
- The program prints a prompt, and the driver reads it from the other end.
- The driver's
expectcall returns when the prompt text matches, or raises an error at the deadline or when the program exits. - The driver's
sendlinecall sends the answer as typed input, and the terminal echoes it back when echo is on. - The driver waits for the output to end and reads the exit code.
An answer sent before its prompt can be lost, because Python's getpass clears pending input when it turns echo off. An early answer can also be echoed before echo turns off, or reach the wrong question after a change adds a prompt. A test that waits fails on the prompt text instead.
A plain pipe can also hide the prompt. On a pipe, many programs buffer stdout in blocks, so a prompt with no newline stays in the buffer while the program waits. In a pseudoterminal, the prompt arrives as it does on screen.
Pseudoterminal output includes echoed answers and ends lines with \r\n, so match prompt text, not whole lines. A menu driven by arrow keys needs the screen checks of a terminal UI test.
What changes when a coding agent writes the code?
A coding agent runs commands through a shell tool, which usually gives them no terminal and no person to answer. At a prompt, an empty stdin returns end of file at once, so Python's input() raises EOFError. A stdin pipe that stays open with nothing written makes the read wait until the tool's timeout stops the command.
The agent often has no way to type into a running prompt either. Its tests may then replace input() with a stub that never checks echo, the prompt text, or the wait. A task that describes only a person at the keyboard can also produce a command with no flags, which then stalls the agent's own runs.
A practical adjustment is to name a flag for each prompt in the task and give the agent an expect test with a short deadline. An unanswered prompt then fails with the last output, e.g. buffer (last 100 chars): 'dev@example.com\r\nAPI token: ', which names the stalled question.
What are common mistakes?
These mistakes make prompt tests fail on working code or pass on broken code:
- Fixed sleeps. A pause before each answer can be too short on a busy CI runner and cause a flaky test. Wait for the prompt text instead.
- Accidental regular expressions. The
expectmethod in pexpect reads a string as a regular expression, soDelete 3 orders? [y/N]never matches its own prompt. Useexpect_exact. - Piped passwords. Password prompts often read the terminal directly. Python's
getpassopens/dev/ttywhenever the process has a controlling terminal, so a piped test waits for the keyboard on a laptop but reads stdin in CI. - A reader for each prompt. A Go program that makes a
bufio.Scannerfor each prompt works at a terminal, which returns one line per read. From a pipe, the first scanner takes both answers. - Only the expected answers. Also send an empty line for the default, an invalid answer, and end of file with
sendeof(), the Ctrl-D that a person types.
How do you check that it worked?
A prompt test works when it fails on broken code and passes on fixed code. Check it in three ways:
- Break it on purpose. Put
input()back in place ofgetpass, and confirm that the echo check fails. - Drop an answer. Remove the
sendlinecall for the email, and confirm that the test fails at its deadline and names the text it waited for,API token:. - Run it in CI. The job has no terminal, but pexpect opens its own, so the test should pass there too.
The expect tests are integration tests of the prompt code and the terminal. A passing test shows that the scripted answers worked, not that every answer a person might type works.
FAQs
How do you test a password prompt?
A password prompt is tested in a pseudoterminal, because password prompts often read the terminal directly and turn echo off. The test waits for the prompt, sends a test value, and checks that the terminal does not echo it.
Should every prompt have a flag that skips it?
Every prompt should usually have a flag or argument that gives the same answer, so scripts, CI jobs, and coding agents can run the command unattended. A secret goes through stdin instead of a flag, and the prompt still needs its own test.
How long should a test wait for a prompt?
A test should wait for each prompt up to a deadline, not for a fixed time. A passing test moves on as soon as the prompt appears, so a generous deadline, e.g. 5 seconds, costs time only when a prompt never appears.
Can a coding agent answer an interactive prompt?
A coding agent often cannot answer an interactive prompt through its shell tool, because many shell tools give the command no terminal and no way to send input. The agent can run an expect test that answers the prompts, or pass the flags that skip them.