Skip to content

How to test shell scripts

Testing a shell script means running it with controlled input and environment, then checking its output, exit code, and the files it changes.

Last updated , 9 min read

Key points

  • Lint the script with ShellCheck first, because it flags quoting and portability bugs without running anything.
  • Replace each outside command with a stub, so the test controls its output and exit code.
  • Test the run where a command fails, not only the run where everything works.

How do you test shell scripts?

To test a shell script, lint its source with ShellCheck, then run it from Bats tests that replace its outside commands with stubs. Each test controls the input and environment, then checks the output on stdout and stderr, the exit code, and the files the script wrote. Run both checks on every change.

A shell script is a command-line application, so it is tested as a process. Four kinds of check cover it:

  • Static checks. bash -n reports syntax errors, and ShellCheck flags risky patterns, without running the script.
  • Black-box tests. Bats runs the whole script with controlled arguments, input, and environment.
  • Unit tests. A unit test sources the script and calls one function at a time.
  • Traced runs. bash -x script.sh prints each command as it runs.

What do you need before you start?

Gather these first:

  • The target shell. Note the shebang, e.g. #!/usr/bin/env bash, and the Bash version that runs the script.
  • ShellCheck and Bats. Install both, with Bats 1.7 or later, e.g. brew install shellcheck bats-core.
  • A list of outside commands. Each command with effects beyond the test needs a stub.
  • A scratch directory for each test. Bats gives each test its own $BATS_TEST_TMPDIR, which supports test isolation.

How do you test shell scripts step by step?

Here is an illustrative example. Acme Co. sells furniture online, and its command-line tool, acme, exports orders. A developer at Acme asks a coding agent to "Write a script that saves each night's order export as a compressed file." The agent writes export.sh and runs it once without an error:

#!/usr/bin/env bash
set -eu
out_dir=$1
acme export --format csv | gzip > $out_dir/orders.csv.gz
echo "saved $out_dir/orders.csv.gz"

1. Lint the script with ShellCheck

The developer runs shellcheck -f gcc export.sh, which prints one line per finding:

export.sh:4:35: note: Double quote to prevent globbing and word splitting. [SC2086]

The unquoted $out_dir breaks on a path with a space. By default ShellCheck does not flag the pipe, which hides a worse bug.

2. Add Bats and a test file

Bats, the Bash Automated Testing System, is an open-source test framework for Bash 3.2 and later. A .bats file holds @test blocks, and each block runs as a Bash function with set -e. A line that exits nonzero fails the test. Other open-source shell test frameworks exist, e.g. ShellSpec and shUnit2. Bats adds three features that tests use:

  • The run helper. run executes a command in a subshell and saves its exit code in $status and its stdout and stderr together in $output. With --separate-stderr, stderr goes to $stderr instead.
  • Setup and teardown. Functions named setup and teardown run before and after each test.
  • Reports. Bats prints the Test Anything Protocol (TAP) when its output is not a terminal.

A short script usually needs only tests of the whole script. Because a unit test sources the file, the script must run its main code only behind a guard, e.g. if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then main "$@"; fi.

3. Stub the outside commands

A stub is a small fake program that returns fixed output and a chosen exit code. The setup function puts a stub folder first on PATH, so the script finds the stub before the real acme:

setup() {
  mkdir "$BATS_TEST_TMPDIR/bin"
  PATH="$BATS_TEST_TMPDIR/bin:$BATS_TEST_DIRNAME/..:$PATH"
}

@test "fails when acme export fails" {
  cat > "$BATS_TEST_TMPDIR/bin/acme" <<'STUB'
#!/bin/sh
echo "acme: database is locked" >&2
exit 3
STUB
  chmod +x "$BATS_TEST_TMPDIR/bin/acme"
  run export.sh "$BATS_TEST_TMPDIR"
  [ "$status" -eq 3 ]
}
How a PATH stub replaces a command in a test Bats test run export.sh export.sh calls acme Stub acme first match on PATH, exits 3 never reached Real acme later on PATH, never runs
The stub sets what the script receives from acme. The test then checks how the script handled it.

A shell function named acme also works, because Bash looks up functions before PATH. The function reaches a child Bash process only after export -f acme. Neither method replaces a command called by its full path.

4. Run the tests and read the failure

A second test stubs a successful export. The developer runs bats --tap --print-output-on-failure test/:

1..2
not ok 1 fails when acme export fails
# (in test file test/export.bats, line 14)
#   `[ "$status" -eq 3 ]' failed
# Last output:
# acme: database is locked
# saved /tmp/bats-run-2v39ZH/test/1/orders.csv.gz
ok 2 writes the export to a compressed file

The export failed, but the script printed its success line and exited with 0. Without pipefail, a pipeline returns the exit code of its last command, here gzip, which succeeded on empty input. The second test passed on the same broken script.

5. Fix the script and rerun both checks

The fix quotes the path and adds pipefail:

 #!/usr/bin/env bash
-set -eu
+set -euo pipefail
 out_dir=$1
-acme export --format csv | gzip > $out_dir/orders.csv.gz
+acme export --format csv | gzip > "$out_dir/orders.csv.gz"
 echo "saved $out_dir/orders.csv.gz"

The set line now turns on three Bash options:

  • -e (errexit). The shell exits when a command fails, except in an if, while, or until test, before the last command of a pipeline or an && or || list, or after !.
  • -u (nounset). An unset variable is an error, so a missing argument stops the script.
  • -o pipefail. A pipeline returns the code of the last command in it that failed, or 0 if none failed.

Both tests pass, and ShellCheck exits with 0.

6. Run the checks on every change

Add shellcheck export.sh and bats test/ to a pre-commit hook and to a continuous integration (CI) job, e.g. in GitHub Actions. Either command exits nonzero on a finding or a failed test, which fails the hook or job. This example is simplified. A real script would also need a test that a failed export leaves no partial archive.

What does a shell linter catch?

A shell linter is a static analysis tool that flags risky patterns in a script without running it. ShellCheck checks sh, Bash, dash, ksh, and BusyBox scripts. It usually picks the dialect from the shebang line. Its findings include these groups:

  • Quoting. An unquoted variable, e.g. rm $file, splits on spaces and expands glob characters (SC2086).
  • Unchecked failures. A failed cd with no check leaves the script in the wrong directory (SC2164).
  • Portability. Bash syntax in a #!/bin/sh script, e.g. [[ ]], fails where /bin/sh is dash (SC3010).

Each finding has a severity of error, warning, info, or style, and -S warning hides the last two. A # shellcheck disable=SC2086 comment turns one check off for the next command. The option --enable=check-extra-masked-returns turns on SC2312, which flags the Acme pipeline while pipefail is off.

A linter cannot check what acme returns, so a script still needs tests.

What changes when a coding agent writes the code?

A coding agent can check a shell script it wrote, e.g. an install script, by running it once. That run takes the success path, so the Acme script passed with a pipe that could hide a failed export.

An agent can also run a script as bash install.sh, which ignores the shebang and the executable bit. A #!/bin/sh script with Bash syntax then passes, but fails where /bin/sh is dash, e.g. on Debian. A file from an agent's edit tool is usually not executable, so ./install.sh exits with 126.

A practical adjustment is to require, per outside command, a test in which its stub fails. Bats runs the script by name, so the success test also checks the shebang and the executable bit. Review each # shellcheck disable comment that a diff adds. The steps to verify a CLI that an agent built cover a whole tool.

What are common mistakes?

These mistakes hide a failure or report a false one:

  • Negating a command. Bash's -e option does not stop at a command negated with !, so ! grep -q error "$log" fails a Bats test only on its last line. Use run ! with bats_require_minimum_version 1.5.0, or add || false.
  • Piping into run. In run acme export | grep A-1042, the pipe applies to the output of run, which is empty, not of acme. Wrap the pipeline in bash -c.
  • Testing only on macOS. On the Bash 3.2 in macOS, a failed [[ ]] stops a Bats test only on its last line. BSD tools also differ from GNU tools, e.g. sed -i.
  • Closing a pipe early. With pipefail, acme export | head -n 1 can exit with 141 when head closes the pipe early, which is a broken pipe. Read to the end instead, e.g. with sed -n 1p.

How do you check that it worked?

The tests work when they fail on a broken script. These checks confirm it:

  • Remove pipefail again and confirm that the first test fails.
  • Run ShellCheck on each shell file, including scripts with no .sh suffix.
  • Run the suite on the target system and Bash version.
  • If a test uses a golden file, change one line of it and confirm that the test fails.

A passing suite shows only that the tested paths worked.

FAQs

What is Bats?

Bats, the Bash Automated Testing System, is an open-source test framework for shell scripts. Each test is a Bash function that fails when a line in it exits nonzero. Its run helper records a command's exit code and output.

How do you mock a command in a shell test?

A command in a shell test is mocked with a stub, a small script of the same name that returns fixed output and a chosen exit code. The test puts the stub's folder first on PATH, so the script runs the stub.

What does set -euo pipefail do?

The line set -euo pipefail turns on three Bash options. A failed command stops the script, an unset variable is an error, and a pipeline fails when any command in it fails. The first option has exceptions, e.g. a command after an exclamation mark.

Should shell scripts have unit tests?

Shell scripts with many functions can have unit tests that source the script and call one function at a time. A short script usually needs only tests that run the whole script as a process. Bats runs both kinds.