# Is AI-generated code secure?

AI-generated code carries the same security risks as code people write, plus a few of its own, so it needs review, scanning, and testing before release.

Last updated September 29, 2026, 9 min read

## Learning objectives

After reading this article you will be able to:

-   List the causes of security flaws in AI-written code
-   Explain why working code can still be insecure
-   Identify checks that reduce security risk

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [Common bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs)
-   [What are AI hallucinations in code?](https://specstory.com/learning/verification/ai-code-hallucinations)
-   [What is AI slop in code?](https://specstory.com/learning/verification/ai-slop)

## Is AI-generated code secure?

AI-generated code is not secure by default, and it is not safe to use unchecked. Code from a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) can contain the same flaws as code a person writes, plus a few of its own, e.g. imports of packages that do not exist. Security checks lower the risk, and each finds only the flaws it looks for.

A security flaw lets someone do what the app should refuse, e.g. one customer changing another customer's order. Finding these flaws is one part of how teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing).

An agent can write code faster than a team reviews it. In [vibe coding](https://specstory.com/learning/ai-coding/vibe-coding), where the builder accepts the code unread, nobody reviews it. [METR disclosed](https://metr.org/blog/2026-08-31-security-update/) that a dashboard whose authentication silently failed open let an attacker extract an API key and run up about $600,000 in model usage over three weeks. METR's report says the app was vibe coded, and a [fail-open](https://specstory.com/learning/cli-and-web/fail-open-authentication) login is one of the common [vibe coding bugs](https://specstory.com/learning/verification/vibe-coding-bugs).

## What causes security flaws in AI-written code?

Five causes come from how a coding agent and its model work.

### Does the model repeat patterns from older code?

A model generates code from patterns in its training data, including insecure ones, e.g. a SQL query built by joining strings. The Open Worldwide Application Security Project (OWASP) [cheat sheet](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Coding_with_AI_Cheat_Sheet.html) on secure coding with AI warns that models often suggest outdated library versions with known vulnerabilities.

### Does the prompt say who may do what?

A prompt describes a feature and often omits the rules around it, e.g. that only an order's owner may change its delivery address. A coding agent builds what the prompt asks, so the code can lack the rule and serve users it should refuse.

### Are the added packages real?

A model can generate an import for a package that does not exist, and an attacker can register that name with harmful code. This is called [package hallucination](https://specstory.com/learning/verification/ai-code-hallucinations), and the attack is slopsquatting. An agent can install the package before anyone checks it.

### Do the agent's tests check security?

An agent usually writes its tests from the same prompt as its code, so they check the feature, not the requests the app should refuse. OWASP says the agent's own tests give no independent assurance and asks for human review of each test an AI changes. A review by the same agent in the same session shares the context that missed the rule.

### Can someone steer the coding agent?

A coding agent reads files and web pages and runs commands with the access it has. Text an attacker places in that content can steer what the agent writes and runs, which is called [prompt injection](https://specstory.com/learning/environments/prompt-injection-in-coding-agents). OWASP ranks it first among [risks for apps](https://genai.owasp.org/llm-top-10/) built on language models.

## What does an insecure AI-written change look like?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds this route, shortened here:

```js
app.patch("/api/orders/:id/address", requireUser, async (req, res) => {
  const { address } = req.body;
  const { id } = req.params;
  await db.query(`UPDATE orders SET address = '${address}' WHERE id = '${id}'`);
  res.json({ ok: true });
});
```

The `requireUser` step checks that a customer is signed in, but nothing checks that the order is theirs. The change then goes through these checks:

1.  The agent's test signs in as a customer, changes that customer's address to 12 Elm Street, and passes.
2.  The developer tries the feature in the running store, and the address saves.
3.  A static application security testing ([SAST](https://specstory.com/learning/glossary#sast)) tool flags a possible SQL injection, because text from the request goes straight into the query.
4.  A [two-account test](https://specstory.com/learning/cli-and-web/two-account-access-test) signs in as customer B and sends the same request for customer A's order `A-1042`. The server answers `200`.
5.  The developer sends both findings to the agent, which changes the query as shown below.
6.  The scan and the two-account test run again, and customer B's request now gets a `404`.

```js
const result = await db.query(
  "UPDATE orders SET address = $1 WHERE id = $2 AND customer_id = $3",
  [address, id, req.user.id]
);
if (result.rowCount === 0) return res.status(404).end();
```

The scan reported nothing about the missing ownership check, because none of its rules encodes who owns which order. Only the second account found it. This example is simplified. A real review would also check the other routes that read or change orders.

## Why can working code still be insecure?

Working code does what its users ask. Secure code also refuses what nobody should be allowed to do, and a feature request rarely lists those refusals. A test written from the request usually checks only the request the feature was built for.

Diagram: Two requests to the same route

The agent's test sends the first request, the one the feature was built for. The flaw is in the second request, which none of the agent's tests sends.

Many security flaws sit in requests that a feature was not built for:

-   **Another user's data.** A request names a record that belongs to someone else, as customer B's request did.
-   **Input that becomes code.** Text from a form ends up inside a query or a page, e.g. a quote mark that ends a SQL string early.
-   **A check that errors.** A fail-open check lets a request through when the check itself throws an error.
-   **Repeated attempts.** A login form with no limit on attempts lets a script guess passwords.

## How can teams reduce security risks in AI-written code?

The [Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) from the National Institute of Standards and Technology (NIST) treats reviewing code (PW.7) and testing executable code (PW.8) as separate practices. For AI-written code, teams can combine these checks:

-   **Scan the source.** A SAST tool is a form of [static analysis](https://specstory.com/learning/code-review/static-analysis) that flags known flaw patterns, e.g. request text that flows into a query. Run it on each pull request, whoever wrote the code.
-   **Check each added package.** Before installing one, look it up in its registry and check when it was created and who maintains it. Run a dependency audit too, e.g. `npm audit`. OWASP recommends both steps. An audit finds only vulnerabilities that someone has already reported.
-   **Scan for secrets.** Run [secret scanning](https://specstory.com/learning/glossary#secret-scanning) on the repository and its history, because a key in an old commit stays readable after the file changes.
-   **State and test access rules.** Write each access rule into the prompt. Then run a two-account test on each route that reads or changes a customer's data, since a scanner rarely flags a missing ownership check.
-   **Review the security rules by hand.** A reviewer reads the access rules and the input handling. The reviewer also reads each change to build scripts or pipeline files, because those run with more access than the app. A [code review checklist](https://specstory.com/learning/code-review/code-review-checklist-for-ai-code) keeps these checks from being skipped.
-   **Limit what the agent can reach.** OWASP advises running agents in sandboxes with credentials scoped to the task. [Least privilege](https://specstory.com/learning/environments/least-privilege-for-ai-agents) and a way to [keep secrets away](https://specstory.com/learning/environments/agent-secrets) from the agent leave less to leak.
-   **Give each change a human owner.** OWASP asks for a person who approves each AI-generated change before merge and answers for its security.

A scan with no findings means only that its rules matched nothing, and [no bugs found](https://specstory.com/learning/verification/no-bugs-found) is not the same as no bugs.

## How is a security review different from functional testing?

Functional testing checks that the software does what its requirements ask, for users who are allowed to do it. A security review looks for ways to make the software do what nobody should be able to do. It can read the code for flaw patterns and try attacks on the running app.

The two overlap where a feature is itself a security control. A test that a login refuses a wrong password is both a functional test and a security test. Access rules are another overlap, because a functional test can check a rule once someone writes it down. Broken access control, the missing ownership check in the example, is first in the OWASP Top 10, a list of the main security risk categories for web apps.

## FAQs

### Is AI-generated code less secure than code people write?

Whether AI-generated code is less secure than code people write depends on how it is checked, because most of its flaws are the same kinds. It adds a few of its own, e.g. packages that do not exist, and an agent can write code faster than a team reviews it.

### Can a coding agent fix the security flaws it wrote?

A coding agent can fix a security flaw once a scanner, a test, or a reviewer reports it, and the same check should then rerun on the fix. A review by the same agent in the same session shares the context that missed the flaw, so the report should come from outside that session.

### What can a security scanner miss in AI-written code?

A security scanner can miss flaws that depend on rules only the app's owners know, e.g. that only an order's owner may change its address. A dependency audit also misses a harmful package that nobody has reported yet. Access tests on the running app and a human review cover part of that gap.

### What does OWASP recommend for AI-assisted coding?

OWASP's cheat sheet on secure coding with AI recommends checking that each suggested package exists, auditing dependencies for known vulnerabilities, and running agents in sandboxes with scoped credentials. It also asks for human review of test changes and a human owner for each AI-generated change.

---

Source: [Is AI-generated code safe? | Security risks | SpecStory](https://specstory.com/learning/verification/ai-generated-code-security)
