AI Agent Sandboxes, Explained: Why Your AI Needs a Safe Room
Based on “5 things every AI engineer should know about agent sandboxes” by Ryan Ismert and Alan Blount of Google Cloud Tech (September 2026).
Table of Contents
- What is an agent sandbox?
- Why does it matter?
- How do you get it right? 5 lessons from the trenches
- The bottom line
Imagine hiring a brilliant intern who works at lightning speed, never sleeps, and occasionally does something completely unpredictable. Now imagine giving that intern the keys to your entire office on day one.
That’s roughly what happens when you let an AI agent run code directly on your machine or servers. Modern AI agents don’t just chat. They write code, install software, edit files, and reach out to the internet. That’s what makes them useful. It’s also what makes them risky.
The fix is a sandbox. Here’s what it is, why it matters, and how to pick one without getting fooled by marketing.
What is an agent sandbox?
A sandbox is an isolated, walled-off environment where an AI agent can run code without being able to touch anything outside it. Think of it as a padded room: the agent can do whatever it wants inside, but if something goes wrong (a buggy script, a malicious package, a hijacked prompt), the damage stays in the room.
Why does it matter?
Because AI-generated code is, by definition, untrusted code. A language model is a probabilistic engine. It doesn’t “intend” harm, but it can be tricked (through prompt injection) or simply make a dangerous mistake. Without isolation, one bad command or one poisoned package could compromise your whole infrastructure.
And there’s a twist: AI agents behave unlike anything cloud infrastructure was designed for. They’re not steady web services, and they’re not run-once batch jobs. The authors found agents spend about 95% of their time waiting (for the model to think or a tool to respond) and then burst into short, intense spurts of work. Standard cloud assumptions don’t fit that pattern well.
How do you get it right? 5 lessons from the trenches
1. Don’t trust the “100 ms startup” headline
Sandbox vendors love advertising lightning-fast cold starts. The catch? Those numbers usually measure booting an empty, bare-bones system that does nothing.
Real agents need a filesystem, a Python or Node runtime, networking, and often heavy libraries. In the authors’ tests, a modest sandbox took about 610 ms to be ready, roughly four times the advertised figure. A bigger desktop-style sandbox took 2.3 seconds, plus about 10 more seconds if you needed a headless browser.
The fix: Keep a pool of pre-warmed sandboxes ready to go, like a taxi rank instead of calling a cab from across town. With well-managed warm pools, allocation can drop to around 200 ms. Also keep your base images lean.
2. Isolation is about what you share with your neighbors
The key security question is simple: what layer of software does the agent share with the host machine? The less shared, the safer. There are four main tiers:
| Tier | Examples | Speed | Safety for untrusted code | Catch |
|---|---|---|---|---|
| V8 Isolates / WebAssembly | Cloudflare Workers, Deno | Ultra fast (under 5 ms) | Good | JavaScript/TypeScript/Wasm only |
| Standard containers | Docker, Kubernetes Pods | Native speed | ❌ Weak (shared kernel) | One kernel bug = host compromised |
| User-space kernels | gVisor | Near-native for most code | ✅ Strong | Some overhead on system-call-heavy work |
| MicroVMs | Firecracker, Kata Containers | Fast | ✅ Strongest | More memory per sandbox |
The big takeaway: regular Docker is not a security boundary for AI code. Containers share the host’s operating system kernel, so a single kernel exploit can hand an attacker the whole machine.
The fix: Plain hardened containers are fine for internal tools running trusted code. For anything running AI-generated code on behalf of users, use gVisor or microVMs, and layer on extra defenses.
3. The real danger is the internet connection
Security teams often obsess over whether attackers can “break out” of the sandbox. For AI agents, that’s the wrong priority.
An attacker who hijacks your agent via prompt injection doesn’t need a fancy breakout. If the sandbox has open internet access, the agent can simply send your files somewhere, query cloud metadata services (like the well-known 169.254.169.254 address), or grab credentials lying around in its environment. No exploit required, just ordinary web requests.
The fix:
- Block all outbound traffic by default. Only allow specific, approved domains.
- Block cloud metadata addresses (the
169.254.0.0/16range). - Strip secrets and credentials from the sandbox environment.
- Route any API calls through a gateway that handles authentication outside the sandbox.
4. Saving and restoring state beats fast booting
Real agent sessions span many steps: read a file, run a test, see it fail, fix the code, try again. If the sandbox wipes itself after every step, the agent loses its work.
But keeping a dedicated machine running the whole time is expensive (remember, it’s idle 95% of the time). And attaching storage disks on demand can add 20+ seconds per step.
The fix: Memory snapshots. When the agent goes quiet, freeze its state and save it to cloud storage. When it’s needed again, restore it in roughly 0.25 to 3 seconds. Think of it as hitting “pause” on a video game and resuming exactly where you left off. On optimized warm pools, per-call latency can drop to around 50 ms, versus 500 ms or more on setups without pooling.
5. Choose your sandbox with four questions
There’s no single best sandbox. Ask these in order:
- Is the work limited to JavaScript, TypeScript, or simple math? → Use V8 isolates or WebAssembly. Fastest and cheapest.
- Is it internal, single-user, with trusted code? → Hardened standard containers work fine.
- Many users, untrusted code, full Linux tools needed? → User-space kernels like gVisor (or managed options built on them).
- Need custom kernel modules or hardware-level isolation? → MicroVMs like Firecracker.
As the authors wryly hint: anyone claiming one technology wins every scenario might just be trying to sell you compute.
Practical tip: List the actual commands your agent runs. If most are standard Python data scripts, start with a managed user-space sandbox and benchmark from there with realistic traffic.
The bottom line
Letting AI agents run code is powerful, and risky. A good sandbox setup comes down to a few principles:
- Treat all AI-generated code as untrusted.
- Don’t rely on Docker alone for isolation.
- Lock down the network first, because that’s where real attacks happen.
- Use warm pools and snapshots to keep things fast and affordable.
- Benchmark real workloads, not vendor demos.
Get these right, and your agent gets the freedom to be useful without the keys to the kingdom.