Skip to content

What is the difference between reading code and running code?

Reading code, as in AI code review, finds problems visible in the source, while running it, as in testing, finds failures that appear only at runtime.

Last updated , 8 min read

What is the difference between reading code and running code?

Reading code examines a change's text, and running code starts the software and observes what it does. Code review, including AI code review, reads the code, and testing usually runs it. Reading finds problems visible in the source, e.g. a missing error check. Running finds problems that appear only when the software runs, e.g. a checkout that loses the cart.

The two methods find different bugs, so neither replaces the other. Static analysis tools, e.g. a linter, also read the code. The same split is often called static and dynamic analysis, or static and dynamic testing.

The Secure Software Development Framework from the National Institute of Standards and Technology (NIST) treats reviewing code (PW.7) and testing executable code (PW.8) as separate practices. A coding agent can say "done" about a change that nobody has read or run, so teams that test AI-generated code need both.

What does reading code catch?

Reading code means examining the source, the diff, and nearby files without executing anything. For tools, the International Software Testing Qualifications Board (ISTQB) defines static analysis as the automated evaluation of a component or system without executing it.

A careful reading often catches these problems:

  • Logic errors. A reversed condition is visible in the changed lines.
  • Missing error handling. An empty catch block sits in the diff.
  • Code on paths that no run takes. A reader can check an error branch that no test triggers.
  • Risky patterns. A reviewer can flag a database query built from user input.

AI code review applies the same method with a language model. Some review tools also run the project's tests in an isolated environment. That step is a run, but only of what those tests check. Other tools read logs from earlier runs of the deployed code, which some vendors call runtime context. That is still reading, because the change itself has not run.

A reading ends in a prediction. It can find some runtime bugs, e.g. a value that can be null where a function needs a number. It cannot show a failure that depends on something outside the text, e.g. database records.

What does running code catch?

Running code means building the software, starting it, giving it input, and comparing what it does with what it should do. ISTQB defines dynamic testing as a test approach that executes the item under test. Running catches failures that depend on data, configuration, or timing. A run can also leave a record of the steps, the input, and the actual result.

Running takes four common forms:

  • Tests. A test suite runs parts of the code with chosen input and checks the results.
  • End-to-end runs. End-to-end testing drives the whole app through a user workflow.
  • Exploratory use. In exploratory testing, a person or an agent uses the software without a fixed script.
  • Monitored runs. Runtime verification checks a running program against properties written in advance.

A test suite can pass while the app fails to start, if no test starts the whole app. A run is safest in a sandbox, where a failing change cannot touch real data.

When should a team read code or run it?

Read every change. Also run it whenever it touches a workflow that users or other services depend on, because that behavior depends on data, configuration, and timing that the diff does not show.

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent's change goes through both kinds of check:

  1. An AI reviewer reads the diff and flags an empty catch block around the call that saves the address. A failed save would show the customer no error.
  2. The agent adds an error message, and the developer approves the change.
  3. A test runner starts the store in a sandbox, adds 2 items to the cart, and changes the address to "12 Elm Street" at checkout.
  4. The next page shows an empty cart. The address saved, but the cart emptied.

Each check found a bug that the other would have missed. The run never made the save fail. The save started a fresh checkout session, and the code that loads a session's cart was outside the diff.

This example is simplified. A real project would run more workflows, e.g. a save that fails.

Which bugs show up only at runtime?

These bugs depend on things outside the source text, so they usually need a run:

  • Failures across steps. A bug that needs a sequence of actions, e.g. the emptied cart, appears only when the workflow runs in order.
  • Timing bugs. A race condition depends on the order of events that run at the same time, and that order can change between runs.
  • Configuration and startup failures. A missing environment variable can stop an app from starting even when its code reads correctly.
  • Integration failures. A change can break another service that expects the old response format.
  • Browser rendering failures. A hydration error appears when the server's HTML does not match what the browser's JavaScript renders on load.
What reading and running each examine Change Reading Read the text diff and related files Comments a prediction of behavior Running Run the software build, start, and use it Actual result steps, input, and output data, configuration, timing, other services
Reading works from the text of the code. A run adds the data, configuration, and timing that runtime bugs depend on.

Reading can raise a suspicion about these bugs, e.g. two requests that change the same record. A run can show that the failure happens, but a run that misses it does not show that it cannot happen.

Can reading and running code be used together?

Reading and running code are usually used together. Reading can start before the code builds, and it can point to an exact line. Running starts once there is a build, often while a reviewer is still reading the diff.

Each check has a cost. In a 2024 industrial study, developers resolved 73.8% of an AI reviewer's comments, but pull requests took longer to close (from about 6 to about 8 hours on average). A run needs a working environment and test data.

Using both still leaves gaps:

  • A comment is a suspicion until a run shows the failure. Some reported problems never happen, and a run can confirm the ones that do.
  • A run covers only the workflows it tries. A run with gaps does not establish that the whole app works.
  • A passing check can still be wrong. A test that compares the wrong value passes on broken software, which is a false pass.

How do reading code and running code compare?

The two methods differ on these points:

PointReading codeRunning code
What it examinesThe source, the diff, and related filesWhat the built software does with specific input
What it costsTime to read and act on each commentA working environment and test data
Paths it coversAny path in the text, including untested onesOnly the paths that the run reaches
Bugs it finds wellLogic errors, missing error handling, and risky patternsFailures across steps, timing, configuration, and startup
What it producesComments that predict behaviorSteps, input, and the actual result
Main blind spotBehavior that depends on data, timing, or configurationCode on paths that no run takes

Both can catch a logic error that is visible in the source and also fails at runtime.

How do you run a change before review?

Run the workflows that a change touches in a separate copy of the app before a person reviews the diff, and keep the steps, input, and output as evidence. A review can approve a change, and its tests can pass, while the workflow a customer uses is broken.

RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Can AI code review replace testing?

AI code review cannot replace testing, because a review ends in comments that predict behavior, and a test run records what the software did. The two find different bugs.

Is running the code the same as running the tests?

Running the tests is one way to run the code, and a narrow one. A test suite runs only the code that its tests call and checks only what they assert. Starting the whole app and using it can find failures that no test covers.

What does "runtime" mean in code review tools?

In code review tools, "runtime" has more than one meaning. Some tools use runtime context, meaning logs from earlier runs of the deployed code, while the change itself never runs. Other tools run the project's tests in an isolated environment, and some find runtime bugs by reading.

Why did review miss a bug that one run would have shown?

Review misses a bug when its cause is not in the text that the reviewer read. The cause can sit in code outside the diff, in the data, or in the timing of events, and one run with specific input brings those parts together.