# How to verify code you cannot read

Verifying AI code without reading it means checking what the running app does against results written in advance, and keeping a record of each check.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Explain why behavior is the test when code is unreadable
-   Apply behavior checks to software an AI built
-   Identify when to bring in a code reviewer

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [What counts as evidence that code works?](https://specstory.com/learning/verification/evidence-that-code-works)
-   [What is runtime verification?](https://specstory.com/learning/verification/runtime-verification)
-   [Common bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs)

## Key points

-   Write down the results a change must produce, in plain words, before the coding agent starts.
-   Try each result on a fresh copy of the app, as a customer and as a second user.
-   Have someone who reads code check logins, payments, and personal data, which behavior checks cannot fully show.

## How do you verify code you cannot read?

To verify AI-generated code you cannot read, check what the running software does instead of what the code says. Write the expected results before a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) starts. Then try each one on a fresh copy of the app, as 2 different customers, and keep a record. Have someone who reads code check logins and payments.

Checking a change does not require reading every line of it. Customers see only what the app does, and checking that needs no code knowledge. Developers also call such a check [runtime verification](https://specstory.com/learning/verification/runtime-verification). The check is black-box testing, one half of [black-box and white-box testing](https://specstory.com/learning/testing/black-box-vs-white-box-testing), because it judges software by inputs and outputs.

Code from an agent can look finished and still be wrong in one place. In [Stack Overflow's 2025 survey](https://survey.stackoverflow.co/2025/ai), the most common problem developers reported with AI tools was answers that are almost right but still wrong. Of the developers who answered that question, 66% had run into it.

People who build by [vibe coding](https://specstory.com/learning/ai-coding/vibe-coding) already judge an app by using it, and these steps make that repeatable. They are the running part of how teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing). Testing a whole app [before launch](https://specstory.com/learning/testing/pre-launch-testing) takes a longer checklist.

## What do you need before you start?

A behavior check needs 4 things:

-   **The request.** The exact prompt or task the agent received is the source of each check.
-   **A fresh copy of the app.** A preview built from the code the agent committed works best, with test data and no real customers.
-   **Test accounts.** Use 2 customer accounts, plus one for each other role, e.g. an admin.
-   **A place for records.** A shared document or ticket holds each check's steps and results, with a screenshot or screen recording.

## How do you verify code you cannot read step by step?

Here is an illustrative example. Acme Co. sells furniture online. Acme's founder, who does not write code, plans to ask a coding agent to "Let customers edit their delivery address during checkout."

### 1\. Write the expected results in plain words

Before the agent starts, write one line for each result that must be true when the change works. These lines are [acceptance criteria](https://specstory.com/learning/ai-coding/acceptance-criteria), which a person can mark true or false. Include what must stay the same. The founder writes 5 lines:

```text
1. At checkout, a customer can change the delivery address.
2. The changed address shows on the payment page and after a reload.
3. Changing the address keeps the items in the cart.
4. A second customer cannot see or change the first one's address.
5. Orders placed earlier keep the address they were placed with.
```

If the agent already finished, write the lines before opening the app, so its behavior cannot shape them.

### 2\. Run the change on a fresh copy of the app

Send the request. When the agent replies that the feature is "done," ask for a preview, a copy of the app built from the committed code, not the agent's working folder. That folder can hold files and settings that the commit lacks, so tests should run on a [clean checkout](https://specstory.com/learning/environments/clean-checkout-testing) of the code. The founder gets a preview at `preview.shop.example.com`, built from the agent's branch `address-edit`.

### 3\. Walk the changed journey as a customer

Start with the journeys that cost the most when they fail, e.g. paying. Follow each one as a customer would, which is a [user journey](https://specstory.com/learning/testing/user-journey-testing) test run by hand, and compare each screen with the written results, including after a reload.

The founder adds 2 items to the cart, starts checkout, and changes the address to 12 Elm Street. The payment page shows no items. The address saved, but the cart emptied.

### 4\. Repeat as a second user and as a stranger

Sign in as the second customer and paste the URL of the first customer's order page, e.g. `preview.shop.example.com/orders/A-1042`. The page should refuse or show nothing. Then paste the same URL in a private window, where nobody is signed in, and expect the login page. This [two-account test](https://specstory.com/learning/cli-and-web/two-account-access-test) checks behavior, not security as a whole.

### 5\. Try the paths a demo skips

Spend a fixed time, e.g. 30 minutes, on actions that a demo leaves out. Submit a blank address. Press Back halfway through checkout, then click Pay twice. This is [exploratory testing](https://specstory.com/learning/testing/exploratory-testing). Record each surprise with its steps.

### 6\. Rerun the journeys that worked before

A change to checkout can break the order history next to it, so rerun the checks from earlier changes. Order `A-1042`, placed before the change, still shows its old address.

### 7\. Send each failure back, then repeat the same steps

Give the agent the steps, the expected result, and the actual result, not a summary:

```text
Copy:     preview built from branch address-edit
Steps:    add 2 items to the cart, start checkout,
          change the delivery address to 12 Elm Street
Expected: the payment page shows 2 items
Actual:   the payment page shows 0 items
```

After the fix, repeat the exact steps on a fresh preview, then steps 3 to 6. Each record becomes [evidence](https://specstory.com/learning/verification/evidence-that-code-works) that another person can review or repeat. This example is simplified. A real store needs more checks, e.g. one for the order total.

## What are common mistakes?

These mistakes let a broken change pass a behavior check:

-   **Treating a demo as the check.** A demo repeats the path that the builder already tried, so it rarely shows what the builder missed.
-   **Trusting the agent's summary.** A message that says "done" is text the agent generated, not a record of a check.
-   **Checking only the screen.** A success message can appear when nothing saved, so reload and open the saved record.
-   **Letting the agent write the checks.** Tests from the same session tend to miss what the code missed, so keep the written results where the agent does not edit them.

## How do you check that it worked?

The check of a change is complete when these statements are true:

-   Each written result was tried on a fresh copy built from the committed code, after the last change.
-   Each check has a record of its steps, results, and the version it ran on.
-   The second account and the private window could not reach the first customer's data.
-   Each fixed failure passes on its original steps, and the journeys that passed before still pass.
-   Someone who reads code looked at the parts in the next section, or the record says nobody did.

No findings is not the same as complete coverage, and a path that nobody tried can still break.

## When do you need someone to read the code?

Some failures do not show in behavior until someone attacks the app or it holds real data. Bring in a person who reads code for these parts:

-   **Logins and access rules.** A login check that fails open lets requests through when the check itself breaks, and in normal use it behaves the same as a working one.
-   **Payments and personal data.** Code that charges cards or stores customer details can fail on paths that no manual check reached.
-   **Secret keys.** A key placed in the code that the browser downloads works in a demo, and any visitor can copy it.
-   **Fixes that keep breaking other features.** When each fix moves the bug, the code may carry [comprehension debt](https://specstory.com/learning/code-review/comprehension-debt), which only a person who reads it can pay down.

Get that review before the app takes real payments or stores real personal data, from a trusted developer or a paid code audit. The reviewer needs only the changes to these parts, not every line. Adding that reviewer moves a project from vibe coding toward [agentic coding](https://specstory.com/learning/ai-coding/agentic-coding-vs-vibe-coding). For someone who cannot read code, an AI code review of these parts is one more claim to check.

## How does RunStory help with code you cannot read?

When nobody reads the code, an agent saying "done" is a claim that needs to be verified. RunStory independently runs the software against your change and returns evidence to the coding agent. It runs your software in a separate environment, tries relevant workflows, and checks the results.

When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. RunStory is in private alpha for CLIs and web apps, and your team keeps the final release decision.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Do you need to learn to code to check an app an AI wrote?

Learning to code is not needed to check what an app does. Behavior checks need written expected results and a running copy of the app. Reading the parts that handle logins, payments, and personal data does need someone who reads code.

### What should you test first in an app you did not write?

The first tests in an app you did not write should cover the journeys that cost the most when they fail, e.g. paying. Then repeat them as a second customer, because a demo with one account cannot show one user reaching another's data.

### When is a paid code audit worth it?

A paid code audit is worth it before an app takes real payments or stores real personal data, and when fixes keep breaking other features. Ordinary behavior checks rarely show a login check that fails open or a secret key sent to the browser.

### Is a demo evidence that the app works?

A demo is evidence that one path worked once, for one account. It does not show what a second user, bad input, or a reload would do.

---

Source: [How to verify AI code without reading it | SpecStory](https://specstory.com/learning/verification/verifying-unread-code)
