What does "runtime" mean in AI code review?
In AI code review, "runtime" and "sandbox" are labels for four different things, and only some of them run the changed software. Runtime-aware review can mean data from earlier production runs, a sandbox for the tool's own analysis, a sandbox that tests suggested fixes, or a run of the app itself.
Runtime context is information about how code behaves when it runs, e.g. an execution trace from production, that a review tool reads without running the change. A reviewer that reads it stays on the static side of static and dynamic analysis, because the change has not run. The International Software Testing Qualifications Board (ISTQB) defines dynamic testing as a test approach that executes the item under test.
The Secure Software Development Framework from the National Institute of Standards and Technology (NIST) treats reviewing code (PW.7) and testing executable code (PW.8) as separate practices. It also suggests running tests in a sandboxed environment. A review tool with a sandbox can sit on either side of that line, depending on what runs inside it. In code review of work from a coding agent, the label alone does not show whether anyone ran the change before it merges.
How do review tools use runtime information?
Each meaning runs something different, or nothing at all:
- Production telemetry. The tool reads records from earlier runs of the deployed code, e.g. an execution trace, the list of steps that one request took through the program. The records can show that a changed function already fails in production. None of them come from the change, because it has not shipped.
- Analysis sandbox. The tool copies the repository into an isolated environment, then runs linters and scripts that search the code, e.g. for the callers of a changed function. The app does not start, so the tool observes only what those scripts report.
- Fix validation sandbox. The tool writes a patch for a problem it found and runs the project's build and tests on it before suggesting it. Some reviewers run the tests the same way to check a suspected bug before posting it. The changed code runs, but only for the behavior the tests check.
- App run. The tool builds and starts the software, then drives a workflow through requests or a browser. It observes what the running app does, e.g. a hydration error that appears only when a browser loads the page.
The table compares the four meanings:
| Meaning | What runs | Does the change run? |
|---|---|---|
| Production telemetry | Earlier releases, in production | No |
| Analysis sandbox | Linters and scripts that search the code | No |
| Fix validation sandbox | The project's build and tests, on the tool's patch | Partly, with the tool's patch, through the tests |
| App run | The built app, driven through a workflow | Yes |
What is an example of each kind of runtime?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change saves the address, then calls session.reload(), which rebuilds the session from saved data that does not include the cart. It meets each kind of runtime in turn:
- A reviewer that reads production telemetry reports no checkout errors in the last 30 days. Those records came from the code before the change.
- A reviewer with an analysis sandbox runs the linter and flags an unused import. Nothing in the checkout runs.
- A reviewer with a fix validation sandbox writes a patch that removes the import, and the project's 214 tests pass on it. No test checks the cart after an address change.
- A test runner starts the store in a sandbox, adds 2 items to the cart, and changes the address to "12 Elm Street" at checkout.
- The next page shows an empty cart. The address saved, but the cart emptied.
- The runner records the steps, the input, and the actual result, so the coding agent can repeat the failure.
Each reviewer reported correctly on what it examined, but only the app run used the cart. This example is simplified. A real project would also run the payment step and check the order total.
What should you ask a tool that says it runs your code?
A claim to "run your code" can mean an analysis sandbox, a fix validation sandbox, or an app run. The phrase "runtime bugs" names a kind of defect that some tools find by reading. These questions show what a tool does:
- What runs. Ask whether the tool runs its own scripts, the project's tests, or the app.
- Which code runs. Ask whether the run uses the author's change or a patch that the tool wrote. A passing patch does not show that the original change works.
- What the run records. Ask what the tool captures while the software runs, e.g. browser console errors in a web app. A run that keeps only a pass or fail result misses errors that no test checks.
- What comes back. Ask for the steps, the input, and the actual result of each failure, so a person or an agent can repeat it.
- What the run cannot reach. Ask how the tool handles test data, credentials, and outside services, e.g. a payment provider.
A team can add these questions to its code review checklist for AI-generated code. Published review benchmarks usually score comments against known issues with precision and recall, so a high score does not show that a tool ran anything.
What are the limits of runtime-aware review?
Each kind of runtime has limits:
- Telemetry covers only what already shipped. A code path that the change adds has no records, and a path that users rarely take has few.
- Existing tests may come from the author. When a coding agent wrote or edited the tests, a sandbox that runs them repeats the agent's own checks.
- A run covers only the paths it reaches. No findings is not the same as complete coverage. A run with gaps does not establish that the whole app works.
- Setup takes work. A sandbox needs build and start commands, test data, and replacements for outside services, or the run can stop before it reaches the change.
- Untrusted code needs isolation. A pull request can contain code that tries to read secrets or reach other hosts, so a tool that runs it needs a sandbox.
- Runs cost more and can vary. A run usually takes longer than a reading. A tool that explores the app with a language model can try different steps on each run.
How is runtime-aware review different from running the app?
Review that uses runtime information still ends in comments or suggested fixes on a change. Running the app tests the change by starting the software, driving a workflow, and recording what happened. The two overlap when a review tool runs the app itself, the fourth kind of runtime.
A comment about behavior is a suspicion until a run shows the failure. Suspicions that turn out wrong add to review noise. End-to-end testing and runtime verification run the app outside review. A full comparison of reading code and running code sets out which bugs each one finds.
How does RunStory help with running the change?
Two of the four kinds of runtime in review tools do not run the change, and a third runs it only through the existing tests. RunStory independently runs the software against your change and returns evidence to the coding agent. It runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
Is production telemetry a kind of testing?
Production telemetry is a record of how deployed code behaved for real users, not a test that someone designed and ran. A review tool that reads telemetry reads runtime context, and the change under review has not run.
Does a review tool's sandbox run the whole app?
A review tool's sandbox usually does not run the whole app. An analysis sandbox runs linters and search scripts, and a fix validation sandbox runs the project's tests on a patch. Only a tool that builds and starts the software and drives a workflow runs the whole app, and it still covers only the paths that workflow reaches.
What is an execution trace?
An execution trace is a record of the steps a program took during one run, e.g. the operations that one request passed through, in order. A review tool can read production traces as runtime context, but they describe the deployed code, not the change.
Why do some review tools run code in a sandbox?
Some review tools run code in a sandbox to check what reading the diff cannot, e.g. whether a suggested fix builds and passes the tests. The sandbox isolates that run, because code in a pull request can try to read secrets or reach other hosts.