Key points
- A workflow file in the .github/workflows folder checks out the code, installs dependencies, and runs the tests.
- GitHub fails a step whose command exits with a nonzero code, and by default a failed step fails the job.
- A coding agent can wait for the pull request's checks and read the failed log before it stops.
How do you run tests in GitHub Actions?
To run tests in GitHub Actions, add a YAML workflow file to the repository's .github/workflows folder. The file tells GitHub to check out the code, install its dependencies, and run the test command on each push or pull request. The job fails when that command returns a nonzero exit code, and GitHub shows the result on the pull request.
GitHub Actions is the continuous integration (CI) and automation service built into GitHub. A workflow holds jobs, and each job runs its steps in order on a machine called a runner. A test workflow is often the first stage of a CI/CD pipeline. The GitHub Actions documentation covers each key.
A green check means only that no step failed the job. It does not show which tests ran or what they checked.
What do you need before you start?
A test workflow needs four things in place:
- One test command. A single command, e.g.
npm test, runs the whole suite with no prompts. - A reliable exit code. The command exits with 0 when every test passes and with a nonzero code when any test fails, as a command-line application should.
- A lockfile. A committed lockfile, e.g.
package-lock.json, pins dependency versions, so the runner installs what the developer tested. - Tests that need no secrets. Unit tests that need no passwords or API keys also run on pull requests from forks, which by default receive no stored secrets.
Running the CI checks locally first catches a missing setup step on the laptop.
How do you run tests in GitHub Actions step by step?
Here is an illustrative example. Acme Co. sells furniture online, and its web store is a Node.js project whose unit tests run with npm test. A developer at Acme adds a workflow so that each pull request runs those tests.
1. Run the tests from a clean copy
In a fresh clone, the developer runs npm ci, which installs the exact versions in package-lock.json. The tests then pass and exit with 0:
$ npm test
Test Suites: 6 passed, 6 total
Tests: 48 passed, 48 total
$ echo $?
0
2. Add the workflow file
The developer creates .github/workflows/test.yml:
name: test
on:
pull_request:
push:
branches: [main]
jobs:
unit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v7
with:
node-version: 24
- run: npm ci
- run: npm test
The on key runs the workflow when a pull request is opened, reopened, or gets more commits, and for each push to main. For a pull request, actions/checkout checks out a merge commit that GitHub builds from the branch and its base branch. GitHub's Node.js guide follows the same pattern. A Python project swaps in actions/setup-python and runs pytest instead.
3. Limit the token and the run time
The developer adds two settings:
permissions:
contents: read
jobs:
unit:
runs-on: ubuntu-latest
timeout-minutes: 15
The permissions key limits the workflow's GITHUB_TOKEN to reading the repository, so a script in the job cannot push code with it. The timeout-minutes line stops a hung test after 15 minutes, not the default 360.
4. Push a branch and open a pull request
The developer pushes a branch and opens a pull request with GitHub's command-line tool, gh:
git switch -c add-test-workflow
git add .github/workflows/test.yml
git commit -m "Run unit tests in GitHub Actions"
git push -u origin add-test-workflow
gh pr create --fill
The pull request shows a check for the unit job, and its log reports the same 48 passed tests.
5. Read a failing run
Later, the developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's pull request fails one unit test:
FAIL src/checkout/address.test.js
● changing the address keeps the cart
expect(received).toBe(expected) // Object.is equality
Expected: 2
Received: 0
Tests: 1 failed, 47 passed, 48 total
Error: Process completed with exit code 1.
The address saved, but the cart emptied. Jest exits with 1, so the step, the job, and the check fail. The agent fixes the cart code, and the next run passes. This example is simplified. A real project would also run its linter and end-to-end tests.
How do you keep a test workflow fast?
These settings cut the wait for a result:
- Cache dependency downloads. Adding
cache: npmto theactions/setup-nodestep keeps npm's download cache between runs. The action does not cachenode_modules, sonpm cistill runs. - Cancel outdated runs. A
concurrencygroup per workflow and branch, withcancel-in-progress: true, cancels a running workflow when a newer push to that branch starts one. - Run quick checks first. A short job runs the linter and a smoke test, and slower jobs list it under
needs, so they start only if it passes. A bad push then fails sooner, but a good push waits for both jobs in turn. - Run versions in parallel. A
strategy.matrixkey creates one job per value, e.g.node-version: [22, 24], and the setup step readsmatrix.node-version. By default, a failed job cancels the rest of the matrix, andfail-fast: falselets each version finish. - Split a slow suite. A matrix can also give each runner one shard of the suite, e.g.
npx jest --shard=1/3, which works only when the tests have test isolation.
What changes when a coding agent writes the code?
A coding agent's session can end with "done" before its pull request's checks finish. An agent that runs gh pr checks --watch waits for the checks and gets a nonzero exit code if one fails. It can then read the failure with gh run view RUN_ID --log-failed and fix the code in the same session.
Some coding agents run inside a GitHub Actions job and push with that job's GITHUB_TOKEN. GitHub's token documentation states that events from that token start no workflow runs, with a few exceptions. A push to a pull request is one, but its run waits until a person with write access approves it. Until then, the agent's last commit has no test result.
A practical adjustment is to end the agent's instructions with gh pr checks --watch. For an agent that pushes from a workflow, a GitHub App token or a personal access token skips that approval.
What are common mistakes?
A nonzero exit code from a step's shell fails the step and the job. GitHub then skips the later steps unless a step's condition says otherwise, e.g. an actions/upload-artifact step marked if: failure() that keeps the test report.
The first three mistakes below hide a failure, and the last one leaks a secret:
- Piping the test output. The workflow syntax reference states that a Linux
runstep with noshellkey usesbash -ewithoutpipefail. Sonpm test | tee test.logreturns the exit code oftee, which is 0. Settingshell: bashaddspipefail. - Turning failures off. A
continue-on-error: trueline on the test step, or|| trueafter the command, lets the job pass when tests fail. - Rerunning until green. The "Re-run failed jobs" option on the run's page, or
gh run rerun RUN_ID --failed, runs the failed jobs again on the same commit. A pass with no change points to a flaky test or an outside cause, not a fix. - Writing a secret into the file. Anyone who can read the repository can read a token in a workflow file, the kind of leak that secret scanning searches for. Store it as a repository secret instead. GitHub masks a stored secret in the log, but it can miss a changed copy. By default, runs from forks receive no stored secrets.
How do you check that it worked?
The workflow works when these statements are true:
- Each pull request shows the test check.
- The test step's log shows as many tests as the suite holds, e.g. 48, not 0.
- A run that finds no tests fails, which rules out a false pass from an empty suite. Jest fails such a run unless
--passWithNoTestsis set. - A test broken on purpose on a throwaway branch turns the check red, and the log names that test.
- Two runs of the same commit give the same result.
A green run then shows that the workflow runs the suite and reports its failures. It does not show that the suite tests what the change was meant to do.
FAQs
Why does a job pass when a test step fails?
A job passes after a test step fails when a setting or the shell hides the failure. The usual causes are a step set to continue on error, a fallback to true after the test command, and output piped to another command without pipefail. The step's log still lists the failed tests.
How do you run tests on several versions at once?
Tests run on several versions at once through a matrix in the job's strategy. The matrix creates one parallel job per value, e.g. one per Node.js version. By default, a failed job cancels the rest, and turning off fail fast lets each version finish.
How do you keep secrets out of a test workflow?
Secrets stay out of a test workflow when the unit tests need none and any other secret is stored as a repository secret, not written in the file. By default, pull requests from forks receive no stored secrets, so tests that need a key cannot run there.
How do you rerun only the failed jobs in a workflow?
A rerun of only the failed jobs starts from the "Re-run failed jobs" option on the run's page, or from the GitHub command-line tool with its failed flag. The rerun tests the same commit, so a pass afterward points to a flaky test, not a fix.