An AI Agent sandbox is an enforceable execution boundary for code, shell commands, browser sessions, or tools proposed by an AI Agent. It limits what a compromised or mistaken workload can read, modify, contact, consume, and retain. A sandbox reduces blast radius; it does not decide whether a business action is authorized or whether model output is correct.
This distinction matters because an Agent can be influenced by repository files, web pages, retrieved documents, tool output, and user input. Treat generated actions as semi-trusted workload input, then choose isolation from the consequence of failure rather than from a product label.
Key Takeaways
- Start with assets, attackers, reachable interfaces, and unacceptable outcomes.
- A default container is a packaging boundary; its security depends on kernel sharing, privileges, mounts, devices, network, and daemon access.
- gVisor, microVMs, and WASM expose different compatibility, startup, density, and isolation trade-offs.
- Keep credentials outside the sandbox when possible and mediate approved requests through a policy-enforcing broker.
- Deny network egress by default, then allow exact destinations and request shapes required by the task.
- A snapshot speeds startup; it can also preserve secrets, malicious state, or tenant data.
- Test escape, exfiltration, exhaustion, lifecycle, and cross-tenant failures continuously.
Begin With the Threat Model
The sandbox boundary should contain the components an attacker can influence and exclude the assets that must remain trusted.
Ask:
- Can the Agent execute arbitrary native code, only approved commands, or a fixed tool API?
- Is the input private, public, attacker-controlled, or mixed?
- Are multiple customers colocated on one host?
- Which files, sockets, devices, cloud metadata, and network destinations exist nearby?
- Which credentials or delegated permissions are needed?
- What happens if the workload escapes, exhausts resources, or persists after cancellation?
- Can the task be retried without duplicating an external effect?
The model and sandbox are inside the untrusted side of the decision. The Agent Runtime control plane, proxy, credential broker, policy decision, and external service authorization must not accept identity or scope merely because the model supplied it.
Compare Isolation Technologies by Boundary
There is no universal "best sandbox." Select the smallest compatible environment whose residual risk is acceptable.
| Runtime | Isolation boundary | Strength | Typical constraint |
|---|---|---|---|
| Hardened process/container | Namespaces, cgroups, capabilities, seccomp | Low overhead and broad Linux compatibility | Shares the host kernel; weak configuration can expose host resources |
| gVisor | Userspace implementation of much of the Linux system interface | Reduces direct host-kernel syscall exposure | Some syscall and I/O workloads have compatibility or performance costs |
| MicroVM | Hardware virtualization with a separate guest kernel | Strong tenant and kernel boundary | Higher startup, memory, image, and orchestration cost |
| WASM/WASI | Capability-based runtime interface | Small capability surface for compatible components | Not a general Linux environment; ecosystem and interfaces are constrained |
| Dedicated VM or host | Machine boundary | Strong operational separation | Lowest density and highest provisioning cost |
gVisor explicitly states that its Sentry intercepts application system calls and exposes a reduced host surface, while still relying on host controls for resource exhaustion and network policy. Firecracker uses KVM microVMs and a minimal device model. WASI starts components without ambient authority and grants only declared capabilities. Those mechanisms solve different problems; benchmark compatibility and operating cost under the real workload.
Do not treat stronger isolation as perfect isolation. Hypervisors, userspace kernels, host kernels, firmware, device passthrough, orchestration APIs, and side channels all have attack surfaces.
Harden the Filesystem and Process Boundary
A useful filesystem policy starts empty:
- mount only the task workspace;
- mount inputs read-only unless mutation is required;
- place output in a separate bounded volume;
- do not mount home directories, SSH agents, cloud config, browser profiles, or package-manager credentials;
- reject host paths after canonicalization and symlink resolution;
- use a read-only root filesystem and ephemeral scratch space;
- run as a non-root identity with no added Linux capabilities;
- block privilege escalation, host namespaces, raw devices, and daemon sockets;
- enforce process, CPU, memory, disk, file-descriptor, and wall-clock limits.
Mounting /var/run/docker.sock gives a workload control over the Docker daemon and often amounts to host control. Nesting container execution must use a mediated service or a separately isolated worker rather than a privileged container.
Package installation is code execution. Pin registries, packages, versions, checksums, and build provenance. A sandbox limits impact if a dependency is malicious, but it does not make the resulting artifact trustworthy.
Make Network Egress Explicit
Network policy is part of the sandbox, not an optional perimeter feature. Prompt injection can turn one allowed shell command into data exfiltration.
Prefer:
- no network interface inside the execution environment;
- one controlled channel to a proxy outside the sandbox;
- exact destination, method, path, size, redirect, DNS, and content-type policy;
- credential injection only after policy succeeds;
- response size and type limits before data returns to the Agent;
- redacted audit events with request and policy identities.
A hostname allowlist alone is insufficient. Consider DNS rebinding, redirects, alternate IP forms, IPv6, local addresses, cloud metadata endpoints, proxy tunneling, and an allowed service that can fetch arbitrary URLs.
TLS tunneling also limits what a generic proxy can inspect. For sensitive integrations, expose a narrow tool adapter outside the sandbox rather than handing the Agent arbitrary HTTPS access.
Keep Secrets Outside the Boundary
Environment variables, files, command-line arguments, shell history, process memory, and crash dumps are all observable to sufficiently privileged code inside the sandbox. Prefer a credential broker that:
- authenticates the sandbox workload from trusted runtime evidence;
- receives the approved operation and normalized target;
- checks tenant, actor, resource, scope, and current policy;
- mints or injects a short-lived audience-bound credential;
- sends the request itself or returns a non-reusable handle;
- records revocation, expiry, and result status.
The AI Agent identity guide covers workload identity and delegated tokens. The tool security guide covers server-side authorization. Sandboxing complements both; it replaces neither.
Define a Machine-Checkable Sandbox Profile
The control plane should reject unsafe configurations before a workload starts. This dependency-free Go example checks a small admission contract. It is not a container runtime or a complete security policy.
package main
import (
"errors"
"fmt"
"strings"
)
type SandboxSpec struct {
RunAsRoot bool
ReadOnlyRoot bool
HostNetwork bool
Mounts []string
EgressHosts []string
MemoryMiB int
PIDs int
WallClockSecond int
}
func (s SandboxSpec) Validate() error {
if s.RunAsRoot || !s.ReadOnlyRoot || s.HostNetwork {
return errors.New("unsafe privilege or host boundary")
}
if s.MemoryMiB <= 0 || s.PIDs <= 0 || s.WallClockSecond <= 0 {
return errors.New("resource limits are required")
}
for _, mount := range s.Mounts {
clean := strings.ToLower(mount)
if strings.Contains(clean, "docker.sock") ||
strings.Contains(clean, "/.ssh") ||
strings.Contains(clean, "/.aws") {
return fmt.Errorf("forbidden mount: %s", mount)
}
}
for _, host := range s.EgressHosts {
if host == "*" || strings.HasPrefix(host, "127.") ||
host == "169.254.169.254" || host == "localhost" {
return fmt.Errorf("forbidden egress target: %s", host)
}
}
return nil
}
func main() {
spec := SandboxSpec{
ReadOnlyRoot: true,
Mounts: []string{"/workspace:ro", "/output:rw"},
EgressHosts: []string{"proxy.example"},
MemoryMiB: 2048,
PIDs: 100,
WallClockSecond: 300,
}
fmt.Println(spec.Validate())
}
Expected output:
<nil>
Production admission should also resolve paths, verify image digests and signatures, validate runtime classes, block privileged devices, constrain syscalls, enforce network policy, and bind the approved profile to the created instance.
Manage Lifecycle, Snapshots, and Recovery
An Agent sandbox is often stateful during a run and disposable afterward. Give every instance:
- immutable run, tenant, policy, image, and workspace identities;
- a creation deadline and hard maximum lifetime;
- lease renewal independent of model output;
- cancellation that reaches child processes and network operations;
- a bounded output collection phase;
- destruction and storage-deletion evidence.
Warm pools and snapshots reduce startup latency, but they add contamination risk. A reusable base snapshot must contain no tenant data, tokens, shell history, model conversation, mounted workspace, or runtime-generated secrets. Validate image and snapshot digests before assignment. Never return a previously assigned instance to a clean pool without a proven reset.
Cancellation does not undo an external API call. Keep business side effects behind an idempotent, policy-controlled service, record an operation before dispatch, and reconcile unknown outcomes.
Observe Without Capturing New Secrets
Useful events include:
- sandbox and policy identity;
- image and snapshot digest;
- process start, exit reason, resource limits, and peak use;
- normalized filesystem and egress policy decisions;
- broker operation IDs and result classes;
- cancellation, timeout, cleanup, and deletion status.
Do not default to collecting complete prompts, source files, shell output, tokens, or environment variables. Store hashes, counts, classifications, and bounded redacted excerpts when they are enough. Apply tenant access control and retention to the telemetry itself.
Test the Boundary
Release tests should attempt:
| Attack or failure | Expected result |
|---|---|
| Read parent paths and follow symlinks | Denied after canonical path resolution |
| Reach metadata, loopback, private IPs, or redirected hosts | Denied before connection |
| Read environment or mounted credentials | No reusable credential is present |
| Fork bomb, disk fill, memory growth | Hard resource limit and clean termination |
| Open host devices, namespaces, or daemon sockets | Not present or denied |
| Reuse another tenant's snapshot or output | Identity mismatch and allocation rejection |
| Cancel during a child process or network call | Bounded termination and reconciled effects |
| Escape attempt against the selected runtime | Alert, quarantine, patch, and regression case |
Measure boundary coverage, denied attempts, false blocks, startup p50/p95, active and idle resource use, cleanup latency, residual-data checks, and recovery success. A penetration test against one image does not certify all future runtimes and policies.
Frequently Asked Questions
Is sandboxing enough to stop prompt injection?
No. It limits reachable resources and consequences after an Agent follows hostile instructions. The system still needs provenance, tool authorization, output handling, user confirmation, and injection-focused evaluation.
Are containers unsafe for every Agent?
No. A hardened container can be appropriate for low-risk single-tenant work. Use stronger isolation when hostile native code, shared hosts, sensitive data, or high-impact escape consequences make shared-kernel risk unacceptable.
Does a microVM guarantee tenant isolation?
No. It adds a guest-kernel and virtualization boundary, but configuration, device exposure, host services, image provenance, networking, management APIs, and hypervisor vulnerabilities still matter.
Should a sandbox keep state between Agent turns?
Only deliberately. Keep task state in an authoritative store and treat workspace state as scoped, versioned, expiring data. Suspend or snapshot only the defined boundary and prove that another tenant cannot inherit it.
What is the safest default network policy?
No direct egress. Route the minimum required operations through a broker that validates destination and request shape, injects short-lived credentials outside the sandbox, limits responses, and records a redacted audit event.