# What is the difference between reading code and running code?

Reading code, as in AI code review, finds problems visible in the source, while running it, as in testing, finds failures that appear only at runtime.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Compare what reading and running code can observe
-   Explain which bug classes need the software to run
-   Identify how to combine review with runtime checks

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [What is runtime verification?](https://specstory.com/learning/verification/runtime-verification)
-   [What is a false pass (false green test)?](https://specstory.com/learning/verification/false-pass)
-   [Why do coding agents say "done" when the code doesn't work?](https://specstory.com/learning/verification/coding-agent-done-claims)

## What is the difference between reading code and running code?

Reading code examines a change's text, and running code starts the software and observes what it does. [Code review](https://specstory.com/learning/code-review/code-review), including [AI code review](https://specstory.com/learning/code-review/ai-code-review), reads the code, and testing usually runs it. Reading finds problems visible in the source, e.g. a missing error check. Running finds problems that appear only when the software runs, e.g. a checkout that loses the cart.

The two methods find different bugs, so neither replaces the other. [Static analysis](https://specstory.com/learning/code-review/static-analysis) tools, e.g. a linter, also read the code. The same split is often called [static and dynamic analysis](https://specstory.com/learning/code-review/static-vs-dynamic-analysis), or static and dynamic testing.

The [Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) from the National Institute of Standards and Technology (NIST) treats reviewing code (PW.7) and testing executable code (PW.8) as separate practices. A [coding agent](https://specstory.com/learning/ai-coding/coding-agent) can [say "done"](https://specstory.com/learning/verification/coding-agent-done-claims) about a change that nobody has read or run, so teams that [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing) need both.

## What does reading code catch?

Reading code means examining the source, the diff, and nearby files without executing anything. For tools, the International Software Testing Qualifications Board (ISTQB) [defines static analysis](https://glossary.istqb.org/en_US/term/static-analysis) as the automated evaluation of a component or system without executing it.

A careful reading often catches these problems:

-   **Logic errors.** A reversed condition is visible in the changed lines.
-   **Missing error handling.** An empty `catch` block sits in the diff.
-   **Code on paths that no run takes.** A reader can check an error branch that no test triggers.
-   **Risky patterns.** A reviewer can flag a database query built from user input.

AI code review applies the same method with a language model. Some review tools also run the project's tests in an isolated environment. That step is a run, but only of what those tests check. Other tools read logs from earlier runs of the deployed code, which some vendors call runtime context. That is still reading, because the change itself has not run.

A reading ends in a prediction. It can find some runtime bugs, e.g. a value that can be null where a function needs a number. It cannot show a failure that depends on something outside the text, e.g. database records.

## What does running code catch?

Running code means building the software, starting it, giving it input, and comparing what it does with what it should do. ISTQB [defines dynamic testing](https://glossary.istqb.org/en_US/term/dynamic-testing) as a test approach that executes the item under test. Running catches failures that depend on data, configuration, or timing. A run can also leave a record of the steps, the input, and the actual result.

Running takes four common forms:

-   **Tests.** A test suite runs parts of the code with chosen input and checks the results.
-   **End-to-end runs.** [End-to-end testing](https://specstory.com/learning/testing/end-to-end-testing) drives the whole app through a user workflow.
-   **Exploratory use.** In [exploratory testing](https://specstory.com/learning/testing/exploratory-testing), a person or an agent uses the software without a fixed script.
-   **Monitored runs.** [Runtime verification](https://specstory.com/learning/verification/runtime-verification) checks a running program against properties written in advance.

A test suite can pass while the app fails to start, if no test starts the whole app. A run is safest in a [sandbox](https://specstory.com/learning/environments/ai-sandbox), where a failing change cannot touch real data.

## When should a team read code or run it?

Read every change. Also run it whenever it touches a workflow that users or other services depend on, because that behavior depends on data, configuration, and timing that the diff does not show.

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's change goes through both kinds of check:

1.  An AI reviewer reads the diff and flags an empty `catch` block around the call that saves the address. A failed save would show the customer no error.
2.  The agent adds an error message, and the developer approves the change.
3.  A test runner starts the store in a sandbox, adds 2 items to the cart, and changes the address to "12 Elm Street" at checkout.
4.  The next page shows an empty cart. The address saved, but the cart emptied.

Each check found a bug that the other would have missed. The run never made the save fail. The save started a fresh checkout session, and the code that loads a session's cart was outside the diff.

This example is simplified. A real project would run more workflows, e.g. a save that fails.

## Which bugs show up only at runtime?

These bugs depend on things outside the source text, so they usually need a run:

-   **Failures across steps.** A bug that needs a sequence of actions, e.g. the emptied cart, appears only when the workflow runs in order.
-   **Timing bugs.** A [race condition](https://specstory.com/learning/debugging/race-condition) depends on the order of events that run at the same time, and that order can change between runs.
-   **Configuration and startup failures.** A missing environment variable can stop an app from starting even when its code reads correctly.
-   **Integration failures.** A change can break another service that expects the old response format.
-   **Browser rendering failures.** A [hydration error](https://specstory.com/learning/cli-and-web/hydration-error) appears when the server's HTML does not match what the browser's JavaScript renders on load.

Diagram: What reading and running each examine

Reading works from the text of the code. A run adds the data, configuration, and timing that runtime bugs depend on.

Reading can raise a suspicion about these bugs, e.g. two requests that change the same record. A run can show that the failure happens, but a run that misses it does not show that it cannot happen.

## Can reading and running code be used together?

Reading and running code are usually used together. Reading can start before the code builds, and it can point to an exact line. Running starts once there is a build, often while a reviewer is still reading the diff.

Each check has a cost. In [a 2024 industrial study](https://arxiv.org/abs/2412.18531), developers resolved 73.8% of an AI reviewer's comments, but pull requests took longer to close (from about 6 to about 8 hours on average). A run needs a working environment and test data.

Using both still leaves gaps:

-   **A comment is a suspicion until a run shows the failure.** Some reported problems never happen, and a run can confirm the ones that do.
-   **A run covers only the workflows it tries.** A run with gaps does not establish that the whole app works.
-   **A passing check can still be wrong.** A test that compares the wrong value passes on broken software, which is a [false pass](https://specstory.com/learning/verification/false-pass).

## How do reading code and running code compare?

The two methods differ on these points:

| Point | Reading code | Running code |
| --- | --- | --- |
| What it examines | The source, the diff, and related files | What the built software does with specific input |
| What it costs | Time to read and act on each comment | A working environment and test data |
| Paths it covers | Any path in the text, including untested ones | Only the paths that the run reaches |
| Bugs it finds well | Logic errors, missing error handling, and risky patterns | Failures across steps, timing, configuration, and startup |
| What it produces | Comments that predict behavior | Steps, input, and the actual result |
| Main blind spot | Behavior that depends on data, timing, or configuration | Code on paths that no run takes |

Both can catch a logic error that is visible in the source and also fails at runtime.

## How do you run a change before review?

Run the workflows that a change touches in a separate copy of the app before a person reviews the diff, and keep the steps, input, and output as evidence. A review can approve a change, and its tests can pass, while the workflow a customer uses is broken.

RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Can AI code review replace testing?

AI code review cannot replace testing, because a review ends in comments that predict behavior, and a test run records what the software did. The two find different bugs.

### Is running the code the same as running the tests?

Running the tests is one way to run the code, and a narrow one. A test suite runs only the code that its tests call and checks only what they assert. Starting the whole app and using it can find failures that no test covers.

### What does "runtime" mean in code review tools?

In code review tools, "runtime" has more than one meaning. Some tools use runtime context, meaning logs from earlier runs of the deployed code, while the change itself never runs. Other tools run the project's tests in an isolated environment, and some find runtime bugs by reading.

### Why did review miss a bug that one run would have shown?

Review misses a bug when its cause is not in the text that the reviewer read. The cause can sit in code outside the diff, in the data, or in the timing of events, and one run with specific input brings those parts together.

---

Source: [AI code review vs. testing | What each catches | SpecStory](https://specstory.com/learning/verification/reading-vs-running-code)
