Table of Contents
- What a sandbox is: aligning on the term
- Tier 1: container defaults — Namespace, cgroups, seccomp
- Tier 2: user-space kernel — gVisor
- Tier 3: lightweight virtualization — Firecracker, Kata Containers, Cloud Hypervisor
- Container-free sandboxes: Anthropic sandbox-runtime, nsjail, bubblewrap
- Cloud sandboxes for agents: E2B, Daytona, Modal, Vercel, Cloudflare
- Case study: how xAI Hades stacks the layers
- How to choose: decision tree and cost
- Overall architecture
- Overall
- References
🌏 中文版
Before letting a model run code for you, answer a more important question than model quality: where does the code run, what can it see, and whom can it touch. This post turns a long Q&A on sandboxes into a comparable spectrum, enriched with how major execution environments actually work today.
You will get three things: the four pillars and three build patterns of sandbox design, where each of Anthropic sandbox-runtime, Firecracker, gVisor, E2B, and AWS Lambda MicroVMs sits and what it trades off, and a decision tree you can use tonight.
What a sandbox is: aligning on the term
A sandbox is not a package name. It is a set of techniques that confine untrusted code to a bounded place. The metaphor is literal: sand stays in the pit even when you make a mess; it does not dirty the living room.
Three core actions define a sandbox:
- Isolation: the program cannot see what it should not see
- Restriction: only the minimal permissions and resources are granted
- Reset: throw it away after use and return to a clean state
And four pillars define its design:
- Isolation tier: determines hardness vs. cost (see the spectrum below)
- Access control and resource limits: read-only mounts or Copy-on-Write for files, no-egress by default for network, CPU/memory/PID caps
- Deception and mimicry: security sandboxes fake mouse trails, decoy documents, and time dilation to fool sandbox-aware malware
- Lifecycle: boot from a base image → execute and trace → destroy on timeout or completion; ephemeral is the default
Use the four pillars to audit any vendor: whichever is missing must be compensated elsewhere.
Tier 1: container defaults — Namespace, cgroups, seccomp
This is where most teams start, and where the most common misconception lives. Running docker run python:3.12 already has isolation, but it is not a sandbox designed for untrusted code.
Four primitives, each owning one question:
- Namespace → what can you see: PID, Network, Mount, User, IPC, and UTS are isolated so the process thinks it is the only tenant. It cannot see host processes or interfaces.
- cgroups → how much can you use: caps CPU, memory, disk I/O, and process count. An infinite
while True: x.append("hello")hits the limit and gets OOM-killed instead of taking down the host. - seccomp → what can you ask the kernel to do: filters system calls, blocking dangerous ones like
rebootormount. - capabilities / AppArmor / SELinux → how much privilege you can escalate: splits root into fine-grained capabilities and drops
SYS_ADMINby default.
# This has isolation
docker run python:3.12 python -c "print(1)"
# This is a sandbox for untrusted code
docker run --network none --memory 512m --cpus 1 \
--cap-drop ALL --read-only --pids-limit 64 \
--security-opt seccomp=default.json \
python:3.12 python /work/task.py
Philosophy: soft isolation inside one kernel, best performance and fastest start. vs. alternatives: cheaper than gVisor and Firecracker, but a shared kernel means a kernel vulnerability can lead to container escape. Good for: internal trusted workloads, CI. Not for: running arbitrary user-uploaded code directly, or hardware-grade multi-tenant isolation. Limitation: the default docker run has almost no limits; you must tighten them explicitly for it to count as a sandbox.
Tier 2: user-space kernel — gVisor
gVisor is Google's open-source sandbox runtime. Written in Go, it interposes a user-space kernel (Sentry) plus a filesystem proxy (Gofer) between the application and the host kernel.
Normal container: App → Linux Kernel
With gVisor: App → gVisor (Sentry/Gofer) → Linux Kernel
- Philosophy: keep the container ecosystem, add interception on the syscall path. Every
open,exec, andsocketgoes through Sentry first, never touching the host directly. - vs. alternatives: much safer than plain containers, lighter than microVMs; but compatibility is not 100%—rare syscalls, ioctls, and
/procbehaviors may differ. In practice 1–2% of long-tail tests hit differences. Workloads with many small file I/Os or frequent network calls see 10–30% extra latency. - Good for: teams that want stronger multi-tenant isolation on existing Docker / Kubernetes without changing images, running common languages (Python, Node, Go).
- Not for: workloads deeply dependent on kernel internals, full GPU passthrough, or extreme I/O sensitivity.
- How to use it: attach
runscvia a Kubernetes RuntimeClass:
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: gvisor
handler: runsc
---
apiVersion: v1
kind: Pod
metadata:
name: untrusted-job
spec:
runtimeClassName: gvisor
containers:
- name: worker
image: python:3.12-slim
resources:
limits: { memory: "512Mi", cpu: "1" }
GKE Sandbox is the managed version of this path, and Cloud Run also runs on gVisor. On EKS you can install containerd-shim-runsc-v1 on nodes and use the same RuntimeClass, but you own versioning and Karpenter image management.
Tier 3: lightweight virtualization — Firecracker, Kata Containers, Cloud Hypervisor
When the trust boundary says "even if they break the kernel, they should not get out," you need hardware virtualization. Each workload gets its own minimal kernel, no kernel shared.
- Firecracker: AWS's microVM virtual machine monitor (VMM) for Lambda and Fargate, backed by KVM, millisecond boot, memory overhead as low as a few MB, powering over 15 trillion Lambda invocations per month. Minimal device model and snapshot-based fast resume.
- Kata Containers: OCI-compliant lightweight VM containers—each Pod is a microVM. Docker UX with VM security guarantees, ideal for Kubernetes clusters that require hardware isolation.
- Cloud Hypervisor: a Rust-based modern VMM in the same space as Firecracker, focused on modularity and security.
Philosophy: trade hardware boundaries for the strongest isolation and near-100% Linux compatibility. vs. alternatives: slightly slower and more complex to operate than gVisor, but no longer limited by user-space kernel compatibility gaps. Good for: highly untrusted code, multi-tenant high-value data, compliance that explicitly requires VM isolation. Limitation: networking (tap/bridge), images (kernel + rootfs), lifecycle, and snapshots are yours to manage—unless you use a managed service.
Lambda MicroVMs, launched on 2026-06-22, makes this tier managed: build an image from a Dockerfile, upload to S3, and get a Firecracker snapshot. Each user or job gets an isolated microVM with its own HTTPS endpoint (HTTP/2, gRPC, WebSocket), pausable and resumable within 8 hours with state and memory preserved. Teams get Hades-like properties without building a Firecracker scheduler.
Container-free sandboxes: Anthropic sandbox-runtime, nsjail, bubblewrap
Not every sandbox needs a container. A recurring theme in the notes is "add a boundary directly around any process."
Anthropic sandbox-runtime (ASRT)
Anthropic sandbox-runtime (@anthropic-ai/sandbox-runtime, Apache-2.0, open-sourced 2025-10-20) is fully documented in Claude Code as "OS-level limits without requiring a container."
Its design is dual-track:
- Filesystem uses deny-then-allow: readable by default, then broad
denyRead: ["~/.ssh"]rules, withallowReadtaking precedence to re-open. Writes are allow-only, denied by default. - Network is forced through a proxy: on Linux, the network namespace is removed from the bubblewrap container so there is physically no network; all traffic must go via Unix domain sockets (
socatbridge) to external HTTP/SOCKS5 proxies; on macOS, a Seatbelt profile only allows connections to the local proxy port; on Windows, a Windows Filtering PlatformALE_AUTH_CONNECTfilter does the same.
The key insight: HTTP_PROXY environment variables are just guidance; the real boundary is OS-level filtering. Even if a tool clears the variable or ignores the proxy, Seatbelt / WFP / network namespace still blocks it. Inside Claude Code, this cut permission prompts by about 84% in auto-allow mode.
nsjail, bubblewrap, Isolate
- nsjail (Google): bundles Namespace, cgroups, seccomp-bpf, capabilities, and chroot into a single CLI. Best for understanding how a sandbox is assembled; common in CTFs and code hosting.
- bubblewrap (bwrap): the underlying tool for Linux desktop and CLI sandboxes; bind-mounts directories as read-only or writable, with seccomp filtering.
- Isolate (IOI): a high-performance code-execution sandbox in C++, also built on namespaces and seccomp.
This path is lightweight, fast, and composable—well suited as the local execution layer for coding agents that need to wrap bash calls in an auditable boundary without an image build pipeline.
Cloud sandboxes for agents: E2B, Daytona, Modal, Vercel, Cloudflare
The notes split "local sandbox" from "cloud sandbox" precisely: the former asks "how big is the blast radius on your machine," the latter asks "how do you safely move execution to someone else's machine, without persisting credentials, with lifecycle managed, and results returned safely."
| Priority | Look at first | Why |
|---|---|---|
| Agent-native SDK, full memory pause/resume, self-hostable | E2B | Only two objects (Template + Sandbox), pause/resume preserves files, memory, and processes; Firecracker microVMs, self-hostable on AWS / GCP |
| GPU inference or training on the same platform | Modal | Sandbox API can request GPUs (T4 to H100), sandboxes run on gVisor |
| Long-lived dev workspaces, multiple VM/container flavors, snapshots | Daytona | Filesystem persistence by default, hot snapshots, forks, multiple runtime classes |
| Evaluation, Blueprint, and Devbox workflows | Runloop | Devbox centers on repositories, snapshots, and dev-machine semantics |
| Already on Vercel, TypeScript-heavy | Vercel Sandbox | Shortest path with Vercel OIDC and preview workflows, GA 2026-01-30 |
| Control plane already on Workers, want Durable Objects / R2 | Cloudflare Sandboxes | Built directly on Workers + Containers, credentials injected via proxy so the sandbox never holds live keys |
| No ops and VM-grade isolation with 8-hour state retention | AWS Lambda MicroVMs | Managed Firecracker, no shared kernel, snapshot resume, launched 2026-06-22 |
Two engineering details to avoid pitfalls:
- Pause and billing: E2B's pause preserves memory and processes and does not bill for execution time while paused, but pausing 1 GiB takes about 4 seconds; do not treat it as zero-cost for large-memory workloads. Continuous execution caps are 1 hour (Hobby) and 24 hours (Pro); resume resets the clock.
- Credential boundary: baking a full
.envinto an image is the most common mistake. Give each sandbox only a short-lived, revocable token, or use a proxy-injection pattern like Cloudflare's credential proxy where the Worker injects secrets at request time.
Case study: how xAI Hades stacks the layers
The field probe of Hades in the notes is a strong reference because it is not theory but "what a platform for running arbitrary agent code looks like."
Layers:
Kubernetes cluster (hades-openbar)
└── Hades Runtime (custom)
├── Isolation backends (pluggable: gVisor / custom hypervisor / runc)
├── catatonit (PID 1, reaps zombies)
├── xai-hades-styx (statically linked Rust, exec/timeout/OOM/session)
├── grok-computer-server.mjs (Node.js control plane on 127.0.0.1:4242)
└── Ubuntu 24.04 userland + grok-files FUSE (remote workdir, JWT auth, virtually infinite capacity)
└── erofs + OverlayFS (/.hades-container-tools read-only compressed, root is OverlayFS)
The online instance currently uses a KVM lightweight VM managed by xai-hades-charon, evidenced by Hypervisor detected: KVM, hvc0, vsock, /dev/vda Virtio devices, and multi-backend error codes like HADES_CLHV_BOOT_TIMEOUT / HADES_RUNSC_START_TIMED_OUT. The workdir /home/workdir/artifacts is not a local disk but a grok-files remote filesystem mounted via FUSE with TERMINAL_JWT_VAL and gRPC.
Three takeaways: make isolation backends pluggable; separate control plane from execution plane (Node.js for tool calls, Rust for resources); remote the filesystem so destroying a sandbox does not lose artifacts. This is exactly the "prepare environment vs. use environment" split the previous section emphasized.
How to choose: decision tree and cost
Use trust boundary and workload traits for the first cut. Comparing feature lists first is less effective.
Three initial questions:
- How untrusted is the code? (your own tooling vs. user-uploaded or model-generated)
- Multi-tenant on shared hardware?
- Need long-lived state or GPU?
Highly untrusted multi-tenant arbitrary code ──→ Firecracker / Kata / Lambda MicroVMs / E2B
│
├── Already on Kubernetes, want to keep the ecosystem ──→ Kata + Firecracker or gVisor RuntimeClass
│
└── Speed and self-host flexibility first ──→ start with gVisor (runsc), upgrade to microVM when needed
Local dev machine agent ──→ sandbox-runtime / nsjail / bubblewrap / Docker SBX
Managed cloud, less ops ──→ E2B / Daytona / Modal / Cloudflare / Lambda MicroVMs (pick by GPU/state/ecosystem)
Costs are two ledgers:
| Cost | Docker + gVisor | Firecracker / microVM |
|---|---|---|
| Machine resources | ~50–100 MB extra per container, medium density | ~5 MB per microVM baseline, potentially higher density |
| Ops and engineering | Swap runtime, ecosystem ready | Own networking, images, snapshots, scheduling; harder to debug |
| Managed cloud | GKE Sandbox integrates easily | Lambda MicroVMs pay-per-use; short tasks can be cheaper |
In short: machine bills may not differ much; the difference is the human cost of running the system stably. That is why the notes conclude "start with Docker + gVisor at small scale, invest in microVM when you hit compatibility or isolation ceilings."
Residual risks even when separated from the main service
Separation is necessary but not sufficient. Still required:
- Sandbox escape: any software can have vulnerabilities; gVisor raises the bar but is not theoretically impossible to escape
- Resource exhaustion: CPU/memory/disk/network exhaustion that slows other workloads on the same node
- Lateral movement and credential leakage: if the sandbox can read Secrets, ConfigMap, or cloud IAM roles, leakage is still possible
- Egress abuse: a sandbox with free egress can be used for mining, external attacks, or data exfiltration; capabilities like webfetch amplify this—require proxy, allowlists, rate limits, and full logging
Mitigations: default-deny NetworkPolicy, per-sandbox ServiceAccount with least-privilege IAM, enforced resource limits, deny privileged and hostPath, mandatory egress proxy, runtime protection (e.g., Falco), and destroy-when-done.
Overall architecture
Trust boundary and cost increase left → right
┌──────────┐ ┌──────────┐ ┌──────────────┐ ┌──────────────┐
│ Regular │→ │ gVisor │→ │ microVM │→ │ Managed │
│ container│ │ runsc │ │ Firecracker │ │ E2B/Daytona │
│ Namespace│ │ Sentry │ │ Kata/CLH │ │ Modal/Vercel │
│ cgroups │ │ Gofer │ │ (KVM) │ │ Cloudflare/ │
│ seccomp │ │ │ │ │ │ Lambda µVMs │
└──────────┘ └──────────┘ └──────────────┘ └──────────────┘
↑ ↑ ↑ ↑
Local dev Hardened Self-built Low-ops
cluster strong isolation productized
└─────────────┴──────────────┴────────────────┘
Upper layer: sandbox-runtime / nsjail / bubblewrap
(boundary around any process, no image needed)
Local OS sandboxes answer "how big is the blast radius," cloud microVMs and managed sandboxes answer "how to move untrusted execution off your host." They are not either/or but defence in depth: even inside a Cloudflare Sandbox or Lambda MicroVM, the in-container process still benefits from a second wall of seccomp / Landlock.
Overall
A sandbox is not a "whether to do it" question but "at which layer, how tight, and who operates it." Container Namespace / cgroups / seccomp gets you started quickly, gVisor trades minimal changes for a much stronger boundary, Firecracker and Kata Containers trade hardware virtualization for the closest-to-VM guarantees, Anthropic sandbox-runtime and nsjail add "boundary around any process without a container," while E2B, Daytona, Modal, Vercel Sandbox, Cloudflare Sandboxes, and AWS Lambda MicroVMs productize those primitives so teams do not need to become virtualization experts first.
If you want to act tonight, do three things: funnel all untrusted execution through a single entry point with no network by default, replace per-task credentials with short-lived revocable tokens injected via proxy, and enforce explicit resource limits and a destroy policy per sandbox. Isolation tiers can be upgraded gradually, but without these three, no layer can make up for it.
References
- Anthropic sandbox-runtime (GitHub)
- Anthropic sandbox-runtime (npm)
- Claude Code — Sandbox environments
- Anthropic — Code execution tool (sandboxed containers)
- Claude Platform — Self-hosted sandboxes (AWS Lambda MicroVMs / E2B / Modal / Cloudflare)
- gVisor Documentation
- Firecracker
- Kata Containers
- Cloud Hypervisor
- nsjail
- bubblewrap
- Docker — seccomp security profiles
- Isolate — IOI sandbox
- Kubernetes — RuntimeClass
- GKE Sandbox (gVisor)
- E2B Documentation
- E2B — Sandbox persistence
- E2B — Internet access
- Daytona Documentation
- Modal — Sandbox guide
- Vercel Sandbox
- Cloudflare Sandbox SDK
- Cloudflare Sandbox — Security
- AWS — Introducing Lambda MicroVMs
- AWS — Run isolated sandboxes with full lifecycle control: Lambda introduces MicroVMs
- AWS Lambda MicroVMs — Product page
- Firecracker — firecracker-microvm.io
- Falco — Runtime security
- OWASP Top 10 for LLM Applications
- E2B Agent Sandbox: Confining Model-Generated Code in Restorable microVMs
- AI Agent Sandbox Escape and Permission Boundaries: Containers Are Not the Whole Security Boundary
- Learning from Mature Coding Agents (11): Sandboxes and Remote Execution — Cloudflare Sandbox in Practice
Loading...