# What is the difference between Docker, dev containers, and VMs?

Docker, dev containers, and virtual machines put a coding agent behind different boundaries, and each exposes the files, network, and secrets it is given.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Compare Docker, dev containers, and VMs for agents
-   Explain what each protects and leaves exposed
-   Identify a setup for running an agent unattended

## Related content

-   [What is a sandbox for AI agents?](https://specstory.com/learning/environments/ai-sandbox)
-   [What is the difference between containers, gVisor, and microVMs?](https://specstory.com/learning/environments/containers-vs-gvisor-vs-microvms)
-   [What is approval fatigue in coding agents?](https://specstory.com/learning/environments/approval-fatigue)
-   [What is a staging environment?](https://specstory.com/learning/environments/staging-environment)

## What is the difference between Docker, dev containers, and VMs?

Docker and dev containers run a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) on the host's kernel, while a virtual machine (VM) gives the agent a kernel of its own. A Docker [container](https://specstory.com/learning/environments/containers-vs-gvisor-vs-microvms) isolates the agent's files and processes with limits set at each run. A dev container is a container that the repository defines, with the project's tools installed.

All three can hold a [sandbox](https://specstory.com/learning/environments/ai-sandbox) for an agent, and each is only as strict as its settings. To run Claude Code in Docker, start the container as a user other than root and mount only the project folder from the host. Claude Code's [guide to sandbox environments](https://code.claude.com/docs/en/sandbox-environments) suggests a container or a VM for unattended runs, and a dedicated VM for an untrusted repository.

Diagram: Where each boundary sits

A dev container is a Docker container that the repository defines, so both share the host's kernel. A VM brings its own kernel, and a hypervisor separates it from the host.

## What does Docker protect?

Docker runs software as a container, a process with its own filesystem and limits that shares the host's kernel. Commands in the container reach the image's files and any mounted folder, but not the rest of the host's files, e.g. SSH keys.

This command starts a shell for Claude Code:

```bash
docker run --rm -it \
  --user node \
  --mount type=bind,src="$PWD",dst=/work \
  --mount type=volume,src=agent-home,dst=/home/node \
  --workdir /work \
  node:22 bash
```

Inside the shell, install Claude Code with its native installer, which does not need root, and run `~/.local/bin/claude` to log in. The `--user node` flag avoids root, since Claude Code refuses its `--dangerously-skip-permissions` flag as root. The `--rm` flag deletes the container when it exits, but not the mounted folder or the `agent-home` volume, which keeps the install and the login. Three openings remain:

-   **The mounted folder.** A bind mount is writable by default, so deleting files in `/work` deletes them from the real project.
-   **The network.** Containers can make outgoing connections by default. [Egress control](https://specstory.com/learning/glossary#egress-control) limits them to approved hosts, e.g. the model's API.
-   **Secrets passed in.** Each command in the container can read each variable set with `--env`, each mounted credential file, and the login token in `agent-home`.

A fourth opening is the Docker socket, `/var/run/docker.sock`. Mounting it lets any command in the container start a second container that changes any file on the host.

## What is a dev container?

A dev container is a container defined in the repository that provides a development environment with the project's tools installed. The [Development Container Specification](https://containers.dev/) defines the file, usually `.devcontainer/devcontainer.json`, which names an image, a Dockerfile, or a Docker Compose file and adds features, e.g. a language runtime. A supporting editor or command-line tool builds it and opens the project inside.

Claude Code's documentation adds the agent as a feature:

```json
{
  "image": "mcr.microsoft.com/devcontainers/base:ubuntu",
  "features": {
    "ghcr.io/anthropics/devcontainer-features/claude-code:1.0": {}
  }
}
```

A dev container has the same openings, and the project folder is mounted by default. Some editors also forward the host's SSH agent and Git credentials into it.

Claude Code's [reference dev container](https://code.claude.com/docs/en/devcontainer) runs the agent as a user other than root. Its firewall rejects outbound traffic except DNS, SSH, the local Docker network, and a list of hosts. Its documentation says this setup supports unattended runs with prompts off. It also warns that the container does not stop a malicious project from sending out anything it can read, including the agent's credentials.

## When do you need a full VM?

A VM runs a whole guest operating system, with its own kernel, on virtual hardware that a hypervisor provides. A command that exploits a kernel bug reaches the VM's kernel, not the host's. It still leaves open any shared folder, any secret passed in, and its outbound network. It costs more setup, memory, and start time, so teams usually keep it for three cases:

-   **Untrusted code.** An unreviewed repository runs on a kernel of its own.
-   **Docker inside.** Tests that start containers need a Docker daemon, and a VM can run its own.
-   **A policy.** Some policies require a separate kernel between the agent and the host.

Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Let customers edit their delivery address during checkout." The developer wants the agent to run overnight with its prompts turned off:

1.  The agent starts in the repository's dev container, behind its firewall.
2.  The end-to-end tests run `docker compose up` for a test database and stop with `Cannot connect to the Docker daemon at unix:///var/run/docker.sock`.
3.  The developer moves the agent to a Linux VM with its own Docker daemon.
4.  The developer gives the VM a fresh clone, test credentials, and the dev container's outbound allowlist.
5.  In the VM, the end-to-end test starts the database and fails. The address saved, but the cart emptied.
6.  Overnight, the agent fixes the cart code and the test passes. The developer reviews the branch in the morning.

This example is simplified. A real project would also delete the VM after the review.

## What changes when a coding agent writes the code?

A coding agent with its prompts turned off runs the commands it generates without asking, so the boundary becomes the main check. People call this YOLO mode and often turn it on to escape [approval fatigue](https://specstory.com/learning/environments/approval-fatigue). It moves the [human-in-the-loop](https://specstory.com/learning/ai-coding/human-in-the-loop) check to the review at the end.

Claude Code's built-in sandbox limits only shell commands, so hooks and MCP servers run outside it. A container or VM holds the whole agent, [agent hooks](https://specstory.com/learning/ci-cd/agent-hooks) included.

Some files the agent writes into the mounted project take effect later on the host, outside any boundary:

-   **Dev container settings.** A changed `initializeCommand` in `devcontainer.json` runs on the host at the next start, and a changed `runArgs` entry can widen the container at the next rebuild.
-   **Git hooks.** A `pre-commit` [hook](https://specstory.com/learning/ci-cd/git-hooks) in `.git/hooks` runs on the host at the next commit, in each of the repository's [Git worktrees](https://specstory.com/learning/glossary#git-worktree).
-   **Install scripts.** A changed `postinstall` script runs on the host at the next `npm install`.

Review changes to these files before the next start, commit, or install on the host.

## Can Docker, dev containers, and VMs be used together?

Docker, dev containers, and VMs are often layered. On a Mac, Docker Desktop runs containers inside a [Linux VM](https://docs.docker.com/desktop/features/vmm/) that it manages. Claude Code's built-in sandbox, which uses Seatbelt on macOS, can also run inside a container or VM. Layering still has limits:

-   **A mounted folder crosses each layer.** A Mac folder mounted into a container passes through Docker Desktop's VM, so the VM does not protect it.
-   **Outbound access leaks data.** Any layer that allows outbound traffic can send out data the agent reads.
-   **Isolation does not test the code.** None of the three checks whether the agent's change works.

## How do Docker, dev containers, and VMs compare?

The three options differ on these attributes:

| Attribute | Docker container | Dev container | VM |
| --- | --- | --- | --- |
| Kernel | Shared with the host | Shared with the host | Its own guest kernel |
| Defined in | An image and `docker run` flags | `devcontainer.json` in the repository | The VM tool's settings |
| Project files | Mounted only if flags mount them | Mounted and writable by default | Cloned, copied, or shared |
| Outbound network | Open by default | Open unless a firewall limits it | Open unless its network limits it |
| Docker inside | Needs the host's socket or privileged mode | Same as a Docker container | Runs its own daemon |
| Fits | Quick runs set up by flags | Team setup, unattended runs behind a firewall | Untrusted code and tests that need Docker |

## Where should checks on agent code run?

Run the agent's tests inside its container or VM, not on the host, because they execute code the agent wrote. The same image can run continuous integration ([CI](https://specstory.com/learning/ci-cd/local-ci-checks)) checks locally, which narrows the gap when tests pass locally but [fail in CI](https://specstory.com/learning/ci-cd/tests-pass-locally-fail-in-ci).

Isolation does not check whether the change works. RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and the alpha tests your software in isolated sandboxes.

[Join the RunStory alpha →](https://specstory.com/runstory#alpha)

## FAQs

### Should an agent's container mount the Docker socket?

An agent's container should not mount the host's Docker socket. The socket controls the Docker daemon, which runs as root unless rootless mode is on. A command that reaches it can change any file on the host, so a VM with its own daemon is safer.

### What is the best sandboxing option on a Mac?

The best sandboxing option on a Mac depends on whether anyone watches the run. Claude Code's built-in sandbox uses macOS's own controls, but it limits only shell commands. For unattended runs, a dev container or a VM holds the whole agent, and a mounted Mac folder stays writable from inside.

### Is YOLO mode ever safe?

YOLO mode is reasonable only inside a boundary that holds while nobody watches, e.g. a VM with a fresh clone, no host secrets, and limited outbound traffic. Even there, the agent can send what it reads to an approved host.

### What does the agent leave behind after a run?

An agent's run leaves behind its changes in any mounted folder, including Git hooks, install scripts, and dev container settings that later take effect on the host. Files also stay in named volumes and in any container or VM that nobody deleted.

---

Source: [Claude Code in Docker vs. dev container vs. VM | SpecStory](https://specstory.com/learning/environments/docker-vs-devcontainer-vs-vm)
