# What is a sandbox for AI agents?

A sandbox for AI agents is an isolated place to run an agent's commands, with limits on the files and hosts they can reach, so mistakes do less harm.

Last updated September 29, 2026, 10 min read

## Learning objectives

After reading this article you will be able to:

-   Define a sandbox for AI agents and what it isolates
-   Explain the main isolation technologies
-   Identify what a sandbox does and does not protect

## Related content

-   [What is a staging environment?](https://specstory.com/learning/environments/staging-environment)
-   [What is test isolation (hermetic tests)?](https://specstory.com/learning/environments/test-isolation)
-   [What is prompt injection in coding agents?](https://specstory.com/learning/environments/prompt-injection-in-coding-agents)

## Key points

-   A sandbox limits which files and network hosts a coding agent's commands can reach.
-   Isolation ranges from process sandboxes and containers to microVMs and full virtual machines.
-   A sandbox limits what code can reach, but it does not test whether the code works.

## What is a sandbox for AI agents?

A sandbox for AI agents is an isolated environment that limits which files and network hosts an agent's commands can reach. It is also called an AI sandbox or an agent sandbox. Inside it, a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) can install packages, edit files, and run tests, and a bad command can damage only what the sandbox lets it reach.

Outside coding agents, a sandbox environment is any isolated place to run code or try a change without affecting real systems. The same idea applies to any AI agent that runs code it generates. NIST's [Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) suggests running tests of executable code in a sandboxed environment (PW.8).

A coding agent runs commands with its user's permissions, unless something limits them. Those commands often run code that nobody has read, e.g. the install script of a package the agent added.

Agents can ask a person to approve each command first. Anthropic's [engineering post on sandboxing](https://www.anthropic.com/engineering/claude-code-sandboxing) warns that clicking "approve" constantly slows work and can lead to "approval fatigue," where users may not pay close attention. A sandbox limits what a command can reach while it runs, whoever approved it.

Approving commands automatically does not replace a sandbox either. In a 2026 test, a [security researcher's](https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/) prompt injection chain worked in 3 or 4 of 5 attempts against a coding agent's auto mode, which approves commands with a classifier. That mode had scored 0% attack success on a fixed benchmark that its vendor commissioned. The researcher writes that "a classifier is not a sandbox."

## How does a sandbox for AI agents work?

A sandbox puts two boundaries around the agent's commands. One covers files, and the other covers the network. In a container or virtual machine sandbox, a typical task moves through five steps:

1.  A developer or a tool creates the sandbox from a base image and a copy of the project.
2.  The developer keeps real secrets outside and gives the sandbox test credentials with narrow access, following least privilege.
3.  The agent's commands run inside the sandbox, which lets them change only its own files, e.g. the copy of the project.
4.  When a command opens a network connection, a proxy outside the sandbox checks the host against a list of approved hosts. This check is called [egress control](https://specstory.com/learning/glossary#egress-control).
5.  When the task ends, the developer copies out the changed files and logs, and deletes or resets the sandbox.

Diagram: What a sandbox lets an agent reach

A container or virtual machine sandbox limits what a command can reach. It does not check whether the code the agent wrote is correct.

A sandbox needs both boundaries. Anthropic's post explains that without network isolation, a compromised agent could send out sensitive files, e.g. SSH keys. Without filesystem isolation, the agent could escape the sandbox and gain network access. Both limits also apply to the scripts and subprocesses a command starts, not only to the command itself. The same post reports that sandboxing reduced permission prompts in Claude Code in Anthropic's internal use.

## What are the types of sandboxes for AI agents?

Sandboxes differ in where the boundary sits and how much they share with the host. Each type adds some setup time. A heavier type puts more between the agent and the host, and it starts more slowly.

### Process sandbox

A process sandbox uses the operating system's own controls to limit one program and the processes it starts, while sharing the host's kernel. On Linux, bubblewrap builds one from kernel namespaces, and macOS uses its Seatbelt framework. Some coding agents ship one, e.g. Claude Code, whose sandbox limits writes to the working folder and a temporary folder, and it sends network traffic through a proxy. Its default settings still let commands read most files on the machine, e.g. SSH keys, so the proxy is what keeps them from leaving.

### Container

A container runs a process with its own filesystem and resource limits, e.g. a Docker container. It shares the host's kernel, so it starts quickly and uses little memory. Teams often package an agent's tools and a copy of the project in one.

### Application kernel

An application kernel, e.g. gVisor, sits between a container and the host kernel. It handles the container's system calls itself, in user space, instead of passing them straight to the host kernel. Code in the container then has fewer ways to reach a bug in the host kernel. The extra layer adds overhead to programs that make many system calls.

### MicroVM

A [microVM](https://specstory.com/learning/glossary#microvm) is a small virtual machine with its own kernel and only a few emulated devices, e.g. one started by Firecracker. Hardware virtualization separates its kernel from the host's kernel. It starts faster and uses less memory than a full virtual machine.

### Full virtual machine

A full virtual machine gives the guest a whole virtual computer with its own kernel and operating system, e.g. one run by QEMU. It separates the guest from the host as a microVM does, and it has room for heavy tools, e.g. a desktop session. It takes longer to start and uses more memory than the other types.

### Hosted sandbox

A hosted sandbox is one of these types, run by a provider in the cloud instead of on the developer's machine. [Background coding agents](https://specstory.com/learning/glossary#background-coding-agent) usually work in one and return a branch when they finish. A field report on [dev sandbox options](https://specstory.com/blog/13-ways-to-run-a-dev-sandbox-for-agentic-coding) compares hosted sandboxes with a rented server.

## What is an example of a sandbox for an AI agent?

Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Let customers edit their delivery address during checkout." The agent works in a container sandbox:

1.  The developer starts the sandbox from the project's image. It holds a copy of the web store and a test database with fake customers.
2.  The agent runs `git clean -fdx` to clear old build files. The command deletes untracked and ignored files, e.g. `node_modules`, but only in the sandbox's copy.
3.  The agent runs `npm install`. The proxy allows the package registry, and it blocks one package's install script from reaching `metrics.example.com`, which is not on the list.
4.  The agent edits the checkout code, and the project's tests pass.
5.  The agent starts the web store inside the sandbox and runs an [end-to-end test](https://specstory.com/learning/testing/end-to-end-testing) with 2 items in the cart. The address saved, but the cart emptied.
6.  The agent fixes the cart code and runs the test again, and the test passes.
7.  The developer copies out the diff and the test output, and deletes the sandbox.

The sandbox kept the agent's commands away from the developer's files and real customer data. The end-to-end test, not the sandbox, found the bug. This example is simplified. A real project would also need test credentials for outside services, e.g. a payment provider's test mode.

## What are the limits of a sandbox?

A sandbox controls what a command can reach, not what the command does there. That leaves five limits:

-   **It does not test the code.** Code that runs without harm can still be wrong, so teams need to [test AI-generated code](https://specstory.com/learning/verification/ai-generated-code-testing).
-   **It is only as strict as its settings.** Each folder, secret, and host let into the sandbox is open to any command, including one planted by prompt injection.
-   **A shared kernel gives weaker isolation.** Process sandboxes and containers share the host's kernel, so a kernel bug can let a process escape.
-   **A mounted folder is not a copy.** Many container setups mount the project folder instead of copying it, so a command, e.g. `git clean -fdx`, deletes the real untracked files. Give the sandbox a fresh clone or a copy when the agent does not need to write to the developer's working folder.
-   **It does not check what leaves it.** The diff goes back to the host, where its code runs without the sandbox's limits, e.g. a changed `postinstall` script. Read changed install scripts, Git hooks, and continuous integration (CI) workflows before running the project on the host, and use [secret scanning](https://specstory.com/learning/glossary#secret-scanning) to catch credentials in the diff.

## How is a sandbox different from a staging environment?

A staging environment is a copy of production that a team shares and keeps running for final testing before release. A sandbox is an isolated place where unfinished work runs, and it is often deleted after one task. Staging checks whether a release works in a setup close to production, and a sandbox limits what code can reach while the work is in progress. The two overlap when a team starts a copy of the app that is close to production inside a disposable sandbox.

## What do test environments cover?

The other test environments articles each cover one part of where an agent's code runs and what it can reach:

-   [Staging environments](https://specstory.com/learning/environments/staging-environment) are copies of production that run a release for final testing before it ships.
-   [Ephemeral environments](https://specstory.com/learning/environments/ephemeral-environments) are temporary, complete copies of an app made for one change and deleted afterwards.
-   [Test isolation](https://specstory.com/learning/environments/test-isolation) lets each test set up its own state, so other tests cannot change its result.
-   [Containers, gVisor, and microVMs](https://specstory.com/learning/environments/containers-vs-gvisor-vs-microvms) differ in how much sits between an agent's code and the host kernel.
-   [Docker, dev containers](https://specstory.com/learning/environments/docker-vs-devcontainer-vs-vm), and virtual machines give a coding agent different boundaries around files, network, and secrets.
-   [Least privilege](https://specstory.com/learning/environments/least-privilege-for-ai-agents) gives an agent only the tools, files, credentials, and network access its task needs.
-   [Prompt injection](https://specstory.com/learning/environments/prompt-injection-in-coding-agents) hides instructions in content an agent reads, so the agent follows them instead of its user.
-   [Approval fatigue](https://specstory.com/learning/environments/approval-fatigue) is approving an agent's permission prompts out of habit, without reading what each one allows.

## Where should a coding agent's changes be tested?

Test the changes in a sandbox with test data and test credentials, away from real secrets. Start the app there and run it with real inputs, e.g. a [command-line application](https://specstory.com/learning/cli-and-web/testing-command-line-applications) with real arguments. Keep the output as evidence for the person who reviews the change.

RunStory, in private alpha for CLIs and web apps, runs your software in an isolated sandbox, tries relevant workflows, and checks the results.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Why test AI-written code in a sandbox?

AI-written code is safer to test in a sandbox because it often runs before anyone has read it, including the install scripts of packages the agent added. The sandbox limits the files and hosts that code can reach, and the whole sandbox can be deleted afterwards.

### What does a built-in agent sandbox do?

A built-in agent sandbox limits the commands a coding agent runs, using the operating system's own controls. The commands can usually change files only in the working folder and a temporary folder, and network access is blocked or limited to approved hosts. The sandbox can reduce approval prompts, but it does not test whether the code works.

### Does a sandbox slow a coding agent down?

A sandbox adds some setup time, e.g. building an image and installing packages inside it. Heavier types, e.g. full virtual machines, take longer to start than process sandboxes or containers. A built-in sandbox can also save time, because the agent stops for fewer approval prompts.

### Local, sandbox, or cloud, where should an agent run?

A coding agent that runs commands is safer in a local or hosted sandbox. Directly on a local machine, the agent can reach everything its user account can reach. Background coding agents usually run in a hosted sandbox and return a branch when they finish.

---

Source: [What is a sandbox for AI agents? | AI sandbox | SpecStory](https://specstory.com/learning/environments/ai-sandbox)
