Skip to content

Common bugs in AI-generated code

Common bugs in AI-generated code are recurring failures that a quick demo does not reveal, from open data access to features that never run.

Last updated , 9 min read

What are the most common bugs in AI-generated code?

The most common bugs in AI-generated code, also called vibe coding bugs, fall into 2 groups of 4 that a demo misses. Access bugs are open database rows, permission checks only in the page, login checks that fail open, and exposed secret keys. Behavior bugs are features that only look finished, errors shown as empty results, duplicate records, and regressions.

The name comes from vibe coding, the practice of building an app from prompts and accepting the code without reading it. Finding most of these bugs takes checks on the running app, which is part of how teams test AI-generated code. METR disclosed that a dashboard whose authentication silently failed open let an attacker extract an API key and run up about $600,000 in model usage over three weeks.

Why do vibe-coded apps break after the demo?

A demo follows one path through the app, with one account, working services, and a single try of the requested feature.

What a demo tries and what launch adds In the demo After launch Bug that shows One user, signed in Second user or stranger Access and data bugs Screen after a click Data after a reload Feature only looks finished Every service up A service fails Errors become empty data Each request sent once Retries and repeat clicks Duplicate records Only the requested feature Older workflows too Regressions
A demo tests one side of each condition. Launch adds the other side, and that is where these bugs appear.

A coding agent generates code from the prompt, and a prompt usually describes what the builder wants to see. Rules that a demo never shows are often missing from the prompt and the code, e.g. which customer may read which order.

Tests do not close that gap when the same agent writes them from the same prompt. OWASP guidance on coding with AI states that "a passing test suite generated by the same agent that produced the code provides no independent assurance."

Problems also build up as the app grows. On SlopCodeBench, a 2026 benchmark study where agents repeatedly extend their own code, no agent solved any problem end to end. The code usually grew more bloated and tangled.

Which access and data bugs are most common?

A demo misses access bugs because its one account may do everything it tries.

Can any user read every row in the database?

Row-level security is a database feature that limits which rows each user can read or change. Some hosted backends let the browser query the database with a public key, so these rules decide who reads what. A coding agent can leave the rules off, or write one that allows every row. The demo works because the app's pages filter rows by user, but a stranger with the public key can read the table.

Is the permission check only in the page?

The page hides other customers' orders, but the server returns whatever the request asks for. In an illustrative example at Acme Co., which sells furniture online, the API returns order A-1042 to any account that requests it. Security teams call this broken object level authorization. The demo misses it, because the builder has one account and clicks only the links the page shows.

Does the login check fail open?

A fail-open check lets a request through when the check itself errors. METR's report says the app behind its exposed dashboard was vibe coded and had this bug. One common form catches the error and continues:

async function requireUser(req, res, next) {
  try {
    const user = await verifySession(req.cookies.session);
    if (!user) return res.status(401).end();
    req.user = user;
  } catch (err) {
    console.warn("session check failed", err); // fails open
  }
  next();
}

The demo passes because the check does not throw. When the session secret goes missing after deploy, the check throws on each request and lets it through. The fix is to deny the request in the catch.

Are secret keys in browser code or the repository?

A secret key belongs on the server, but generated code can expose it. Build tools copy some values into browser code, e.g. Next.js copies each NEXT_PUBLIC_ environment variable that browser code reads. Keys also land in files committed to the repository. A leaked key breaks nothing the demo shows. A server key that bypasses row-level security opens the whole database when it ships to the browser.

Which behavior bugs are most common?

A demo misses behavior bugs because it shows the screen, not the data behind it.

Does the feature only look finished?

A ghost feature is code that nothing in the app calls, so tests that pass for it are a false pass. A coding agent can also fill a screen with sample data, or show a success message without saving anything. Either way, the screen shows what the builder asked for, so the demo passes. A reload shows that nothing was saved.

Do errors turn into empty results?

A common pattern in generated code catches an error and returns a default, e.g. an empty list, so the page keeps rendering. When the database fails, a customer gets an empty order history instead of an error, and nobody is alerted. The demo never triggers this, because every service is up.

Does a repeated request run twice?

A customer clicks Pay twice, or a payment provider resends a webhook event when the first delivery gets no success reply. Idempotency means a handler gives the same result when it runs twice. A handler without that property can create a second order or a second subscription. The demo sends each request once, so it never shows the duplicate.

Do changes break features that worked before?

A coding agent edits shared code to finish its task, and that code often serves other working features. In an illustrative example, a developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The address saved, but the cart emptied. That is a regression, and the demo misses it because it tries only the requested feature.

How do you check an AI-built app before launch?

Run these checks before launch in a staging environment with test data, not against production:

  • Use a second account. Sign in as one customer and request another customer's records by ID. This is a two-account test.
  • Call the API as a stranger. Send requests with only the public key, then remove the session secret and confirm that requests are refused.
  • Search for secrets. Search the built browser files for key formats, and run secret scanning on the repository history.
  • Break a dependency. Stop the database, and check that the page shows an error instead of empty data.
  • Send requests twice. Submit the signup form twice and replay one webhook event, and expect one record each time.
  • Reload and rerun. Use each feature from the page a customer reaches, reload it, and rerun the workflows that worked before.

At Acme, the two-account test is one request:

# Signed in as customer B, request customer A's order
curl -i https://shop.example.com/api/orders/A-1042 \
  -H "Cookie: session=CUSTOMER_B_SESSION"

A 403 or 404 is correct, and a 200 with customer A's order is the bug. End-to-end testing turns these checks into scripts, and each fixed bug's steps become a test. Exploratory testing, where a tester uses the app without a script, can find what nobody scripted.

What does this list leave out?

Other failures need checks beyond using the app:

  • Security flaws that normal use does not trigger. Injection and dependencies with known vulnerabilities need a security review and scanning tools. The OWASP Top 10 names the main categories.
  • Packages that do not exist. A model can generate an import for a package that does not exist, and an attacker can publish a package under that name.
  • Attacks on the coding agent. Prompt injection can target the agent that builds the app, not only the app itself.
  • Performance and accessibility. A demo by one builder shows nothing about heavy traffic or screen readers.

The order of this list is not a measured ranking. No findings is not the same as complete coverage.

How does RunStory help with vibe coding bugs?

Security flaws need a security review, but many behavior bugs show up when someone uses the running app beyond the demo path. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result.

Join the RunStory alpha →

FAQs

Is vibe coding insecure?

Vibe coding does not make an app insecure by definition, but it can produce access bugs that a quick demo misses. An app built that way can ship with database rows that any user can read, so it still needs a security review.

Can a coding agent find these bugs in its own app?

A coding agent that writes tests from the same prompt as its code often misses these bugs, because the rules a demo never shows are missing from both. Checks written from outside the prompt, e.g. a two-account test, can find them.

How do you stop regressions as a vibe-coded app grows?

You stop regressions in a growing vibe-coded app by rerunning the workflows that worked before, after each change. When a bug is fixed, its steps become a test, so the test fails if the bug returns.

Does a vibe-coded app need a security review as well as testing?

A vibe-coded app needs a security review as well as testing, because the two find different problems. Behavior checks show what the running app lets a user do. A security review and scanning tools look for flaws that normal use does not trigger, e.g. injection.