What is a sandbox escape?
A sandbox escape is a failure in which code gets past its sandbox's limits and reaches the host or other systems the sandbox protects. In a container, the same failure is called a container escape or container breakout. Code gets out through a bug in the boundary itself or through a gap in how the sandbox was set up.
A sandbox is an isolated environment that limits which files and network hosts its code can reach. After an escape, the code can read the host's files, e.g. SSH keys, and run programs as a user on the host, sometimes as root. On a shared server, it can also reach other sandboxes on the same machine.
Teams run a coding agent in a sandbox so that its commands can run without a person approving each one. Those commands often run code that nobody has read, e.g. a package's install script. An escape lets that code act on the host with no approval prompt in the way.
What causes a sandbox escape?
A sandbox escape needs a path from the code to something outside the boundary. Some paths go through the boundary because of a bug, and others go around it.
Escapes usually come from one of four causes:
- A bug in the shared kernel. A container sends its system calls straight to the host kernel, so one kernel bug can let code out. The Open Worldwide Application Security Project (OWASP) puts host updates first in its Docker security cheat sheet for this reason.
- A bug in the runtime or monitor. In CVE-2019-5736, a flaw in runc, the program Docker uses to start containers, let code running as root in a container overwrite runc on the host. That gave the code root access on the host. In CVE-2024-21626, runc leaked a file descriptor that let a container process reach the host's files. A microVM can fail through a bug in the Kernel-based Virtual Machine (KVM) or in its virtual machine monitor.
- A setting that opens the boundary. Docker's security documentation warns that a container with the host's root folder mounted can change any file on the host. Mounting the Docker socket into a dev container lets code inside start a second container with host folders mounted.
- A shared folder that the host uses later. A mounted folder is the host's real folder. A file that the code writes there, e.g. a Git hook, runs later on the host.
Without any escape, data can still leave through a path that egress control does not check, e.g. Domain Name System (DNS) lookups.
What does a sandbox escape look like?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme runs a coding agent overnight in a dev container, with the project folder mounted from a laptop. The task is "Let customers edit their delivery address during checkout." The escape takes five steps:
- The container's firewall allows only the package registry and the model's API.
- The agent adds an address validation package, and
npm installruns the package's install script inside the container. - The script writes an executable
.git/hooks/pre-commitin the mounted folder, which the container allows. - The agent finishes the address form, and the tests pass. Git does not track
.git/hooks, so the hook appears in no diff. - In the morning, the developer commits on the laptop. Git runs the hook as the developer, outside the container, with access to
~/.sshand the open network.
When the registry removes the package as malware, the developer lists the hooks that are not Git's samples:
$ find .git/hooks -type f ! -name '*.sample'
.git/hooks/pre-commit
The code never broke the container's boundary. The host ran a file from a folder it shared. The developer deletes the hook, revokes the SSH key and tokens the laptop held, and starts the next task on a fresh clone. This example is simplified. A real response would also check .git/config and the laptop's shell startup files, which the hook could have changed.
What changes when a coding agent writes the code?
A coding agent receives each denied command as an error in its output, and its next tool call can try another way to finish the task. That loop needs no exploit when a setting leaves a way around the boundary. A file written during the task can also run later outside the sandbox, as the hook at Acme did.
Claude Code's sandboxing documentation describes an "escape hatch" for commands that fail under its sandbox. The agent may retry such a command with the dangerouslyDisableSandbox parameter. The retry runs outside the sandbox, under the normal permission rules instead of the sandbox's limits. The setting "allowUnsandboxedCommands": false turns the retry off. The same page says that sandboxing "is not a complete isolation boundary."
In a run that nobody watches, e.g. the task of a background coding agent, the permission settings can approve an unsandboxed retry with no person reading it. For unattended runs, turn the unsandboxed retry off. Then hold the whole agent in a boundary that its own settings cannot switch off, e.g. a virtual machine with a fresh clone.
How do you find or prevent a sandbox escape?
These practices close the common routes and limit the damage:
- Patch the host and the runtime. Kernel, runtime, and monitor updates close the known bugs that escapes use. Updates to runc fixed CVE-2019-5736 and CVE-2024-21626.
- Close the openings in the settings. Run containers as a user other than root, keep the Docker socket out, and leave off
--privileged. OWASP's cheat sheet also suggests--security-opt=no-new-privileges. - Add a second kernel for unread code. gVisor answers system calls with its own kernel, and a microVM boots a guest kernel, so fewer calls reach the host kernel.
- Copy the project in instead of mounting it. A fresh clone keeps the host's real folder out of reach. Where a mount stays, block writes to
.git/hooks,.git/config, and editor folders, e.g..vscode, as Claude Code's sandbox does. Before the next commit on the host, check.git/hooksand rungit config --local --listto look forcore.hooksPathorcore.fsmonitor, two settings that can make Git run other programs. - Limit what the host holds. Keep unneeded secrets off the machine that runs the sandbox, and give the task scoped tokens under least privilege. That caps the blast radius of an escape. Continuous integration (CI) runners that test an agent's branch run its code too, so CI for coding agents needs the same limits.
- Watch the boundary. Log the commands the sandbox denies and the hosts it refuses, and read those logs after unattended runs.
- Treat the host as reached. Stop the sandbox and keep its logs. Revoke the host's credentials, rebuild the host from a known image, and close the route before the next run.
None of these layers shows that a boundary holds, because a kernel, runtime, or monitor bug that nobody has reported can get past them. A sandbox that holds also says nothing about the code that leaves it, and AI-generated code still needs a security review.
How is a sandbox escape different from prompt injection?
A sandbox escape gets code past the boundary around it. Prompt injection hides instructions in content an agent reads, so the agent follows them instead of its user. It works inside the boundary. An injected agent can do harm without escaping, with the files, tokens, and hosts the sandbox already grants.
An escape can happen without any injection, e.g. through a kernel bug. The two combine when injected text tells the agent to try routes out. A sandbox limits what an injection can do only while the sandbox holds.
FAQs
What is a container escape?
A container escape is a sandbox escape from a container, in which code reaches the host's files or other containers on the same machine. Containers share the host's kernel, so one kernel bug can let code out. A runtime bug or a setting that opens the boundary can too.
Can a coding agent leave its sandbox without an exploit?
A coding agent can leave its sandbox without an exploit when a setting or a shared folder leaves a way out. Some built-in sandboxes let a blocked command retry outside the sandbox under the normal permission rules. A Git hook written into a mounted folder also runs later on the host.
Can code escape a microVM?
Code can escape a microVM through a bug in KVM or in the virtual machine monitor, or through a folder or setting that the host shares with it. A microVM gives the code its own guest kernel, so fewer of its calls reach the host kernel than in a container.
What should a team do after a sandbox escape?
After a sandbox escape, a team should treat the host and each credential it held as exposed. The team stops the sandbox, keeps its logs, revokes those credentials, and rebuilds the host from a known image. The team also removes what the code left in shared folders, e.g. a Git hook.