Glossary · term

Agent sandboxes

An agent sandbox is an isolated execution environment in which an AI agent can run code, manipulate files or invoke tools with bounded access to the host system. The boundary may use operating-system controls, containers or microVMs, and can restrict files, processes, network destinations and credentials. A sandbox limits the consequences of an action; it does not decide whether that action is appropriate.

Agents2023-06-29Wave 2 · 2024Maturity: 4/5

Origin and context

Sandboxes are a longstanding security mechanism. Their agent-specific use expanded as coding agents and tool-using assistants began executing untrusted model-generated commands. E2B documented a cloud environment for a coding agent in June 2023; later documentation describes on-demand virtual machines. Anthropic describes filesystem and network isolation for Claude Code, while Cloudflare provides isolated containers through its Sandbox SDK. These implementations differ in lifetime, persistence and control surface, but independently establish the category.

Sources: s1, s2, s3, s4

Why it matters

Agents operate across a trust boundary: their commands are generated probabilistically and may also be influenced by untrusted retrieved content. Isolation can reduce blast radius by separating a task from developer laptops, production networks and unrelated secrets. It can also make runs reproducible by starting from a declared image or template. Effective containment still depends on configuration. Broad outbound network access, mounted credentials, persistent volumes or privileged host interfaces can undermine the boundary even when execution occurs inside a sandbox.

Sources: s1, s2, s3

Example

A coding agent receives an issue and checks out the repository into a fresh sandbox. The environment exposes only that checkout, a package mirror and a narrowly scoped token; production credentials and the user's home directory are absent. The agent runs tests and produces a patch, then a human reviews the result before merge. If a dependency contains malicious instructions, the sandbox can constrain access, but separate approval and credential policies are still needed.

Sources: s1, s2, s3

Maturity and evidence

Maturity is rated 4. Multiple independent providers document production implementations, and the underlying isolation mechanisms are well established. The practice remains below 5 because agent-specific threat models, portable policy formats and guarantees vary, while several offerings and SDK interfaces continue to change.

Sources: s1, s2, s3

Limits and open questions

A sandbox is not a complete defense against prompt injection, data leakage or harmful authorized actions. It may contain vulnerable software, allow approved network exfiltration, or expose secrets deliberately mounted for the task. Teams need least-privilege credentials, egress controls, resource and time limits, audit logs, patching, artifact review and tests showing that the boundary fails closed.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as agent sandboxing skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai data security skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as agent threat modeling maestro skill.