# What are AI hallucinations in code?

AI hallucinations in code are imports, calls, flags, or options that name a package or API that does not exist, generated in the same form as real code.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define code hallucinations and package hallucination
-   Explain which hallucinations fail loudly and which run quietly
-   Identify checks that surface each kind

## Related content

-   [How to test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing)
-   [Common bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs)
-   [What is AI slop in code?](https://specstory.com/learning/verification/ai-slop)
-   [Why do coding agents say "done" when the code doesn't work?](https://specstory.com/learning/verification/coding-agent-done-claims)

## What are AI hallucinations in code?

AI hallucinations in code are references to APIs or packages that do not exist, generated by a model in the same syntax as real ones. They are also called code hallucinations or hallucinated APIs. Some fail as soon as the code installs or runs. Others run without an error and ship.

A logic bug uses real names in the wrong way, and [AI slop](https://specstory.com/learning/verification/ai-slop) is real code that is bloated, duplicated, or poorly tested. A hallucination names something that does not exist in the version the project installs, e.g. a method that the library never defined. It can pass the agent's own tests, as other [bugs in AI-generated code](https://specstory.com/learning/verification/vibe-coding-bugs) do.

Finding hallucinations is part of how teams [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing). A [coding agent](https://specstory.com/learning/ai-coding/coding-agent) can install a package and run code before a person reads the name. In [vibe coding](https://specstory.com/learning/ai-coding/vibe-coding), nobody reads the code, so a hallucination surfaces only when something installs, checks, or runs it.

## How do hallucinations get into code?

A model generates code one token at a time, each a likely continuation of the prompt and code so far. An invented name reaches a project in five steps:

1.  The developer's request and the files in context go to the model.
2.  The model generates a call that fits patterns in its training data, e.g. an option that another library accepts.
3.  Nothing in generation looks the name up in the installed library or the registry.
4.  The coding agent writes the code to a file and may run a command, e.g. `npm test`.
5.  If a command fails on the name, the agent receives the error and can fix it. Otherwise the name stays in the change.

Where the name points decides how it fails:

-   **Invented package.** If the registry does not hold the name, the install stops with a "not found" error.
-   **Invented method.** A [type check](https://specstory.com/learning/glossary#type-checking) or a compiler reports it in typed code, and some [linting](https://specstory.com/learning/code-review/linting) rules report it, e.g. Pylint's `no-member` check. Otherwise the call fails only when its line runs, e.g. with `TypeError: order.setDeliveryAddress is not a function`.
-   **Invented flag.** Many argument parsers reject an unknown flag, so a [command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications) stops with an error. Python's argparse prints "unrecognized arguments" and ends with [exit code](https://specstory.com/learning/cli-and-web/exit-codes) 2.
-   **Invented option.** A key in an options object or a settings file is often ignored when nothing reads it, so the code runs and the option does nothing. In typed code, a type checker can reject the key.

An invented call can match an API that an older version had and the installed version removed. A deprecated API is different, because it still exists and runs in the installed version and may print a warning. Putting the installed version's type definitions in the agent's context gives the model the real names.

Diagram: Where each kind of hallucination fails

Each check stops different kinds of hallucination. An ignored option in untyped code and a registered package name produce no error at any stage.

## What is an example of a hallucinated API?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds a server call that sends the edited address to an address service:

```js
const res = await fetch("https://address.example.com/validate", {
  method: "POST",
  body: JSON.stringify({ address }),
  timeout: 3000,
});
```

The change then goes through these steps:

1.  The agent's test swaps `fetch` for a [stub](https://specstory.com/learning/testing/mocks-vs-stubs) that answers at once, and the test passes.
2.  Linting passes, because most lint rules check code patterns, and the object is valid JavaScript.
3.  When the developer tries the change, the address service answers in 90 milliseconds, and checkout continues.
4.  A reviewer approves `timeout: 3000` as a limit of 3 seconds.
5.  Later, the address service stops answering, and checkout hangs. The standard `fetch` has no `timeout` option, so it ignores the key and keeps waiting.
6.  A test that holds the address service open shows the request still waiting after 3 seconds.
7.  The fix passes `signal: AbortSignal.timeout(3000)`, and the same test now gets a `TimeoutError` after 3 seconds.

The key matches an option that version 2 of the `node-fetch` package accepted. A TypeScript version of the file fails the type check, a kind of [static analysis](https://specstory.com/learning/code-review/static-analysis), because `timeout` does not exist in type `RequestInit`. That check covers only an object written inline in the call, not an options object built earlier.

This example is simplified. A real checkout would also need to decide what the customer sees after a timeout.

## What is package hallucination?

Package hallucination is a code generation failure in which generated code imports or installs a package that does not exist. [A 2024 study](https://arxiv.org/abs/2406.10279) by Spracklen and colleagues examined 576,000 code samples from 16 models, in Python and JavaScript. The study found that on average at least 5.2% of package references from commercial models, and 21.7% from open-source models, named packages that do not exist.

The authors also found that many invented names came back when the same prompt ran again.

## What does a clean install catch and miss?

A clean install puts a project's declared dependencies into a fresh environment, e.g. a container built from scratch. It catches an invented name that nobody has registered, shortened here:

```text
$ pip install -r requirements.txt
ERROR: No matching distribution found for acme-address-check
```

Starting the app afterward, as a [smoke test](https://specstory.com/learning/testing/smoke-testing) does, can also catch an import that no declared package provides, e.g. an invented module name. A Python app stops with `ModuleNotFoundError`.

A clean install misses three kinds of hallucination:

-   **Registered invented names.** A name that an attacker registered installs normally, and nothing in the output marks it as wrong. Its code can do what the import expects while it also runs harmful code.
-   **Invented APIs in real packages.** The install checks package names and versions, not the calls the code makes into them.
-   **Invented flags and options.** The install never reads them. A flag often fails when its command runs, and an ignored option in untyped code may produce no error at all.

A registered package can also run code during the install itself, e.g. a `postinstall` script, which npm ran by default before version 12 unless given `--ignore-scripts`. Installing from source with pip can run a package's build code too. A clean install belongs in a [sandbox](https://specstory.com/learning/environments/ai-sandbox), where install code cannot reach real credentials or data.

## How is package hallucination different from slopsquatting?

Package hallucination is the failure, and slopsquatting is the attack that exploits it. In slopsquatting, an attacker registers package names that models tend to invent, so code that installs one of those names installs the attacker's code. In typosquatting, an attacker registers misspellings of real names instead.

The [OWASP cheat sheet](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Coding_with_AI_Cheat_Sheet.html) on secure coding with AI calls the attack "AI-assisted typosquatting." It advises checking that each package an AI suggests exists in the public registry before installing it. The lookups `npm view PACKAGE` and `pip index versions PACKAGE` fail for a name the registry does not hold.

OWASP also advises reading the package page, download count, maintainer history, and creation date. A reviewer can also compare the name with the library the code was meant to use.

## What does running the code tell you about invented APIs?

Running the code turns an invented method or flag into an error, but only on the lines that run. A reviewer or an [AI code review](https://specstory.com/learning/code-review/ai-code-review) tool can flag an unfamiliar name, but a plausible option reads the same as a real one. In untyped code, an ignored option surfaces only when a check exercises the behavior it promised. [Running code](https://specstory.com/learning/verification/reading-vs-running-code) also cannot show that a package that installs is safe.

An invented call can ship on a line that no test runs. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. When something breaks, your coding agent receives the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Does running the code catch slopsquatting?

Running the code does not reliably catch slopsquatting. A package that an attacker registered installs normally, and its code can do what the import expects while it runs harmful code. Checking each unfamiliar package in its registry before installing it is the safer step.

### Can a linter catch a hallucinated API?

A linter catches a hallucinated API only in some cases, e.g. an invented method in a library it can inspect. Most lint rules check code patterns, so an invented option passes them. In typed code, a type checker catches invented methods and can catch invented options.

### How do you check an unfamiliar dependency an agent added?

You check an unfamiliar dependency by looking it up in its public registry before anyone installs it, e.g. with npm view. Compare its name with the library the code was meant to use. Then read the package page, download count, maintainer history, and creation date, as OWASP advises.

### Is a hallucinated API the same as a deprecated one?

A hallucinated API is not the same as a deprecated one. A deprecated API still exists and runs in the installed version and may print a warning. A hallucinated API fails or does nothing in that version, even when an older version or another library had it.

---

Source: [AI hallucinations in code | Invented packages | SpecStory](https://specstory.com/learning/verification/ai-code-hallucinations)
