Skip to main content

Sandboxes in AI Agent Harnesses: Where Files, Commands, and Tools Run

Calculating read time…

A sandbox is the execution boundary that limits what an agent's generated actions can touch.

In an AI agent harness, that boundary matters because an agent may inspect files, create or modify artifacts, run commands, install packages, invoke tools, or reach external services. The model decides what action it wants to attempt; the harness decides whether that action is allowed and where it can actually run.

Imagine a release-triage agent receiving a software repository and a failing test report. The model may decide to inspect the repository, run a test, edit a file, and query an issue tracker. A production harness should not simply hand that agent the same authority as the developer's laptop. Instead, it can give the agent a bounded workspace, restrict outbound network access, keep sensitive credentials outside the execution environment, route selected tools through controlled adapters, require approval for consequential actions, and record what happened.

That is the central idea of this article: the sandbox is not the whole agent harness. It is one enforcement layer inside a larger runtime. Public architectures demonstrate different boundaries: some use operating-system controls, some use containers or virtual machines, and managed platforms may provide isolated per-session execution environments. The exact mechanism varies; the engineering question stays the same: what can this agent reach, from where, for how long, and under whose authority?

🟧 Child-friendly analogy

Think of the model as a person choosing what to do, the harness as the building's security desk, and the sandbox as a locked workshop. The person can ask to use a drill, but the workshop determines which tools and materials are physically available. The security desk can also stop requests before they reach the workshop.

🔀 Quick Comparison
Approach Where the boundary is enforced Useful when Important caution
OS-level sandbox Operating-system security controls restrict process capabilities, files, or network access. Local developer tooling where native isolation is available. Coverage depends on the operating system and the exact policy.
Container-based A container runtime and its host configuration form the execution boundary. Repeatable build environments and workloads that fit containerized infrastructure. Do not equate the word “container” with a universal security guarantee.
MicroVM A virtual-machine boundary separates the workload from the host or neighboring sessions. Higher-isolation multi-session execution and managed agent workloads. The VM boundary does not automatically solve credential, network, or authorization problems.
Managed sandbox A platform provider manages some or all of the execution isolation and lifecycle. Teams that want sandbox lifecycle and scaling without building the isolation layer themselves. You still own application-level authorization, data boundaries, tool policy, and workload configuration.

These are architecture patterns, not a universal ranking. A real system may combine several of them.

01
What a Sandbox Actually Means in a Harness

A sandbox is best understood as an enforced execution boundary. It answers questions such as: Which files can a process read? Which paths can it modify? Can it start child processes? Can those processes reach the internet? Can the process see host credentials? Can two sessions see each other's files?

The word “sandbox” therefore describes a control objective rather than one specific implementation. For example, OpenAI has publicly described local Codex sandboxes that use operating-system mechanisms, while managed agent environments can use isolated compute and separate environments for users or workloads. AWS AgentCore documents per-session microVM isolation for its code interpreter and runtime environments. These are different implementations of the same broad engineering idea: constrain execution and keep sensitive resources outside the agent's default reach.

A harness surrounds that boundary with other responsibilities. A useful conceptual decomposition is:

Layer Primary question Typical responsibility
Model What should I try next? Reasoning, action selection, interpretation of results.
Harness May this action happen, and under what policy? Tool routing, approvals, budgets, state handling, cancellation, tracing, completion criteria.
Sandbox What can executing code physically reach? Filesystem, process, device, network, and other runtime restrictions.
Tool adapter Which external capability is exposed? Authentication, input validation, authorization, downstream API invocation, result filtering.
Engineering principle

Do not make the sandbox carry responsibilities that belong to application authorization, and do not make prompts carry responsibilities that need technical enforcement.

Safety Check

Risk → The model generates an instruction that tries to read, modify, or connect to something outside the intended task boundary.

Control → Enforce filesystem, process, network, and authorization boundaries outside the model.

Remaining risk → A compromised or mistaken agent can still damage anything you intentionally exposed to it. The sandbox reduces blast radius; it does not prove the task is correct.

🎯 Use this when...

An agent can run code or commands, especially when the input contains files, repositories, generated code, browser output, documents, or other untrusted material.

02
What Runs Inside and What Stays Outside

A common beginner mistake is to picture a sandbox as a box containing the entire agent. Usually it is more useful to picture multiple execution planes.

🟧 Child-friendly analogy

Imagine a school workshop. The student can use the tools inside the room, but the teacher keeps the school's master keys outside. If the student needs something from another room, they ask through the teacher rather than receiving every key.

In our fictional Release Triage Agent, the model runs in a managed inference service. The harness receives the model's proposed action. A shell command is sent to a sandbox executor. The executor reads or writes only the permitted workspace. A separate issue-tracker tool is handled by an application-side adapter that holds the downstream credential. The final production change is not made directly by the model.

MODEL
  ↓ proposes action
HARNESS
  ├── sandbox executor ──→ files / shell / test process
  ├── tool adapter ──────→ issue tracker API
  ├── approval gate ─────→ human review for external change
  └── trace store ───────→ run history / policy decisions

Illustrative architecture, not a vendor SDK diagram.

This separation changes how you design permissions. A file-reading command inside the sandbox may need no external credential at all. A ticket lookup may need a narrowly scoped service identity. A ticket update may need a different authorization path. A deployment action may belong to a completely separate control plane.

OpenAI's public agent architecture is a useful documented example of this separation: the harness can run commands in a managed environment, while applications can also provide function tools and can connect their own environments through a separate executor. Its security guidance explicitly recommends keeping application credentials outside the sandbox when possible. OpenAI Agent Architecture and Sandbox Security.

🟩 Worked example — Release Triage Agent

Inside sandbox: repository files, test runner, temporary artifacts, compiler output.

Outside sandbox: issue-tracker credential, approval service, deployment credentials, central audit store.

Harness responsibility: decide whether the requested operation is read-only, workspace mutation, external communication, or consequential change.

Safety Check

Risk → A tool adapter accidentally exposes a broad credential to the sandbox.

Control → Keep long-lived credentials outside the sandbox and broker narrowly scoped requests through an application-side adapter or proxy where practical.

Remaining risk → The adapter itself can still be over-privileged. Review its identity, reachable endpoints, input validation, and result filtering.

03
Files, Workspaces, and Mounts

For many coding agents, the filesystem is the first important sandbox boundary. The critical design decision is not simply “give the agent a sandbox.” It is which paths should be visible, which are writable, which are read-only, and which should not exist from the agent's point of view at all?

A useful starting design is to give the agent one explicit working directory. Project files needed for the task appear there. Temporary files live there. Generated artifacts live there. Sensitive host directories do not.

/agent-workspace/
├── repository/
├── test-output/
├── generated/
└── task-metadata/

not exposed:
├── personal-home/
├── cloud-credentials/
├── production-secrets/
└── unrelated-customer-data/

Illustrative filesystem layout.

The distinction between read access and write access matters. A repository may be safe to inspect but not safe to modify automatically. A reference dataset may be readable but should not be mounted read-write. A generated-output directory may be disposable. A host directory containing credentials should normally not be exposed simply because a command-line tool might “need it.”

Some public sandbox implementations make this distinction explicit. OpenAI's published Codex safety material describes sandbox modes and writable workspace roots, while Docker's current sandbox documentation describes a direct workspace mount as an explicit connection across the host boundary and also documents a clone-based pattern where the host repository is exposed read-only while the agent works on its own copy. These are examples of architecture-specific controls, not universal requirements.

The harness should also decide what happens to files after the run. Temporary artifacts can often be destroyed. User-requested outputs may need to be copied out through a controlled boundary. Logs should not blindly capture file contents that may contain secrets or private customer data.

Engineering principle

Expose the smallest filesystem that makes the task possible, not the largest filesystem that makes the task convenient.

Safety Check

Risk → A prompt injection inside a repository convinces the agent to inspect neighboring directories or configuration files.

Control → Enforce filesystem permissions independently of model instructions. Prefer a dedicated workspace and deny unnecessary paths.

Remaining risk → Any sensitive data intentionally mounted into the workspace remains accessible to processes that can read that mount.

🎯 Use this when...

Your agent manipulates repositories, uploaded documents, generated files, data-analysis artifacts, or any other local content.

04
Commands, Processes, and the Execution Tree

When an agent runs a shell command, the important question is not only “which command?” It is also which process launches it, which child processes it can create, which environment variables they inherit, which files they can see, and whether they can communicate with the network or host?

🟧 Child-friendly analogy

Giving an agent permission to start a program is like giving a worker permission to open one toolbox. The tricky part is that the program may open more tools of its own. A good sandbox controls the whole process family, not just the first command typed by the agent.

This is why effective sandboxes typically enforce restrictions at the operating-system, container, VM, or equivalent runtime layer. OpenAI has publicly described Codex commands as sandboxed from the start in its Windows sandbox engineering discussion, with restrictions intended to propagate through the process tree. Its public safety material also documents platform-specific approaches rather than pretending one mechanism works identically everywhere.

A harness therefore benefits from explicit process limits. Typical controls include a maximum execution duration, cancellation support, output-size limits, process-count limits, temporary-disk limits, and rules about whether the command may access the network. Not every agent needs all of these controls, but each one should have a reason tied to the workload.

Illustrative harness policy

workspace = "/agent-workspace/repository"
network = "deny-by-default"
max_runtime = ""
max_output = ""
max_processes = ""
allow_shell = true
allow_host_files = false

Illustrative configuration only. These names are not a vendor configuration format.

The goal is not to make the environment incapable of doing useful work. The goal is to make the environment predictably capable of the work you intended.

Safety Check

Risk → A generated program starts a second process that has more access than the original command.

Control → Apply restrictions at the execution boundary so child processes inherit the intended limitations.

Remaining risk → Runtime isolation can still be weakened by host integrations, privileged mounts, exposed sockets, or incorrectly configured helper services.

05
Tools, Network Access, and Credentials

A sandbox can dramatically reduce risk while still failing to protect the system if the network and tool layer are too powerful. An agent that cannot read the host filesystem directly may still leak a file through a network-enabled tool. An agent without a production credential may still invoke a tool whose server holds a production credential.

This is why filesystem isolation and tool authorization must be designed together.

🟩 Worked example — Read versus write

The Release Triage Agent may be allowed to call get_issue without approval. The corresponding update_issue capability is a different operation and can require a separate authorization rule. The sandbox is still relevant for commands, but the issue tracker itself is protected by the tool adapter.

The network policy is equally important. A default-deny posture can reduce accidental access to arbitrary services and can limit exfiltration paths. OpenAI's public description of its managed Codex environment emphasizes constrained outbound access and approval or policy around unfamiliar destinations. AWS AgentCore documents multiple network configurations for code-execution environments, illustrating that the required network boundary depends on the workload.

A good harness does not assume that “network enabled” is a single permission. It can distinguish between:

  • no outbound network access;
  • access only to approved package or artifact endpoints;
  • access to an internal API through a controlled proxy;
  • selected public endpoints for a specific task.

Credentials deserve another boundary. OpenAI's current security guidance explicitly recommends keeping application API keys outside the sandbox and describes patterns where a proxy or application layer supplies scoped credentials to approved destinations. This is a strong general design pattern even when the implementation differs across platforms.

Agent → harness → policy check → tool adapter
                ↓
           scoped credential
                ↓
           approved endpoint

Illustrative sequence. The credential should not need to become model-visible just because the tool needs it.

MCP is a useful example of why tool description and runtime enforcement are different concerns. The current MCP specification and its tool-annotation guidance describe ways for tools to communicate behavioral information, but metadata should not be mistaken for a security boundary. A harness still needs its own authorization and execution policy.

MCP 2026-07-28 specification overview and MCP tool-annotation security discussion.

Safety Check

Risk → Untrusted content persuades the agent to use a powerful network-enabled tool.

Control → Separate read and write capabilities, scope tool identities, restrict network destinations, and require approval for consequential operations.

Remaining risk → A legitimately authorized tool can still perform harmful work if its own validation or authorization is weak. Prompt injection is therefore not eliminated by sandboxing alone.

🎯 Use this when...

The agent can browse, call APIs, use MCP tools, access private services, download dependencies, or communicate with systems outside its workspace.

06
State, Persistence, Cleanup, and Isolation

A sandbox is not only about where code runs. It is also about how long its state exists.

Some tasks need temporary state. Others need a persistent workspace across multiple turns. A long-running agent may need to resume after a restart. A multi-tenant platform may require one user's state to remain inaccessible to another user's session.

AWS AgentCore's public documentation gives a concrete example: its code-interpreter sessions maintain files and data during the session, then terminate after the configured session lifetime; its runtime documentation describes per-user-session microVM isolation. The key lesson is architectural rather than product-specific: session lifetime, storage lifetime, and identity scope should be explicit.

State type Example Typical lifecycle question
Ephemeral Temporary files and command output. Can it be destroyed when the task ends?
Session state A working directory and intermediate artifacts. Who can resume it and for how long?
Durable artifact A user-requested report or patch. Where is it stored, and who owns it?
Audit record Tool calls, policy decisions, approvals, outcomes. How long must evidence remain available?

Cleanup is part of security. A sandbox that disappears while a separate durable store retains confidential outputs is not actually ephemeral from a data-governance perspective. Likewise, deleting a VM does not automatically delete external artifacts copied to object storage, issue trackers, logs, or caches.

For multi-user systems, isolation must be tested as a property of the whole lifecycle. A good test does not merely ask whether user A can access user B's files while both sessions are running. It also asks whether a reused workspace, cache, temporary artifact, tool session, network connection, or resume token can accidentally cross identities.

Safety Check

Risk → Session state from one user survives long enough to become visible to another user.

Control → Bind workspace, credentials, tool sessions, and durable artifacts to an explicit workload or user identity; destroy or reset state according to policy.

Remaining risk → Shared infrastructure may still contain caches or logs that require separate tenant-isolation controls.

07
Failure, Recovery, Cancellation, and Completion Checks

A sandbox protects the execution environment, but it does not make the workflow reliable. Commands can fail. Processes can hang. Network calls can time out. The model can misunderstand an error. The environment can disappear midway through a task.

This is where harness engineering becomes more than “put the agent in a VM.” A production harness needs a disciplined lifecycle.

1. create isolated execution context
2. attach only required workspace/data
3. initialize approved tools and policy
4. run task step
5. capture result + policy decision + trace
6. verify expected state change
7. retry only within bounded policy
8. cancel on timeout or unsafe condition
9. preserve approved artifacts
10. destroy or recycle the execution context

Illustrative lifecycle, not a claim about one platform's internal implementation.

Notice the word verify. “The command returned exit code zero” is not always the same as “the task succeeded.” A test command may pass while the requested file was never changed. A report may be generated but contain no records. An external API call may succeed while the intended business state remains unchanged.

Completion criteria should therefore be explicit. For the Release Triage Agent, successful completion might mean: the selected test is green, the intended source file changed, no unrelated files changed, the generated patch is available, and no external ticket was updated without authorization.

🟩 Worked example — Recovery

Suppose a build process exceeds the runtime limit. The harness terminates the sandbox process, records the timeout, preserves the last approved artifact, and starts a clean retry only if the retry policy permits it. The model receives a structured failure result instead of being allowed to continue indefinitely.

Engineering principle

A sandbox should fail closed where practical: when a boundary condition, approval, or resource limit is uncertain, the safer default is to stop the consequential action rather than silently expand authority.

Safety Check

Risk → A timeout or partial failure leaves a process, lock, artifact, or external operation in an uncertain state.

Control → Support cancellation, idempotent operations where possible, explicit completion checks, bounded retries, and post-action verification.

Remaining risk → Some external systems cannot be rolled back automatically. Those actions need stronger authorization and reconciliation procedures.

🎯 Use this when...

Tasks can run for more than a few seconds, invoke external systems, create durable artifacts, or fail in the middle of multi-step workflows.

08
A Practical Sandboxed Harness Design

Now combine the ideas into one concrete architecture. The following is a fictional teaching scenario, not a description of a private production system.

🟧 Scenario

A Release Triage Agent receives a repository and a failing test report. Its job is to investigate the failure and prepare a patch. It may read files, run tests, create temporary files, and query a read-only issue tracker. It must not deploy code or modify production records.

Step 1 — Create the boundary. Start a fresh execution environment for the task. Attach only the repository needed for the investigation.

Step 2 — Establish permissions. Allow workspace writes if patch creation is required. Keep unrelated host paths unavailable. Start with network disabled unless the task actually needs a specific destination.

Step 3 — Separate external tools. Expose a read-only issue lookup tool through the harness. The sandbox does not receive the issue-tracker secret directly.

Step 4 — Enforce runtime limits. Bound command duration, output size, and total run time. Provide a cancellation path.

Step 5 — Trace consequential actions. Record the task ID, selected tool, sandbox identity, policy decision, approval decision, command result, and final verification. Do not store sensitive payloads indiscriminately just because they are convenient for debugging.

Step 6 — Verify before completion. Confirm that the patch changed the intended files, the relevant test passes, and no forbidden external action occurred.

Illustrative pseudocode — NOT a working SDK example

run(task):
  sandbox = create_isolated_environment(task.workspace)
  policy = load_policy(task.type)

  while not complete:
    action = model.next_action(context)

    if action.kind == "shell":
      result = sandbox.execute(action, policy.shell_rules)

    elif action.kind == "tool":
      result = tool_gateway.call(action, policy.tool_rules)

    elif action.kind == "external_change":
      approval = approval_service.request(action)
      result = execute_only_if_approved(approval, action)

    record_trace(action, result, policy)
    check_budget_timeout_and_cancellation()

  verify_completion(task)
  export_approved_artifacts()
  destroy_or_recycle(sandbox)

The function names above are deliberately illustrative. They are not vendor APIs.

This design also shows why context engineering is not the same thing as sandbox engineering. The context may tell the model which repository files are relevant. The sandbox determines what the executing process can physically open. The harness uses both, but they solve different problems.

Question Context Sandbox / Harness
What does the model see? Selected information in model input. May limit what execution can actually access.
What can code touch? Not guaranteed by context alone. Enforced by runtime boundaries and permissions.
Who decides authority? May describe intended behavior. Policy, identity, approvals, and technical enforcement.
Safety Check

Risk → Engineers assume the model will obey the intended task and therefore expose more files, tools, or network access than necessary.

Control → Treat model instructions as one layer of guidance; enforce authority with runtime and application controls.

Remaining risk → No single control eliminates instruction injection, tool misuse, or human approval mistakes. Defense in depth remains necessary.

09
Enterprise Rollout

At enterprise scale, the question changes from “Can the agent run safely?” to “Can the organization operate and audit this capability consistently?”

The first step is to classify workloads. A documentation agent that only reads a disposable workspace is different from a coding agent that modifies a shared repository. A data-analysis agent with access to customer data requires a different data boundary from a synthetic-data experimentation agent.

Control area Enterprise question Evidence to retain
Identity Which user, workload, or service identity started the run? Run identity and authorization context.
Sandbox policy Which filesystem, network, and runtime restrictions applied? Resolved policy or policy version.
Tool use Which tool calls occurred and under which identity? Tool name, authorization decision, result metadata.
Outcome Did the task satisfy its completion criteria? Verification result and approved artifacts.

Change control matters too. Sandbox policies are security-sensitive configuration, not ordinary application preferences. A change from “network disabled” to “internet enabled” can materially change the threat surface. A change from read-only repository access to read-write access can change the blast radius of a model mistake.

Logging should be useful without becoming a second data-leak surface. Public documentation from OpenAI describes agent-aware telemetry that can include prompts, tool approval decisions, tool execution results, MCP usage, and network policy events. OWASP's 2026 Agent Control Standard similarly emphasizes agents being inspectable, traceable, instrumentable, and controllable at runtime. These are useful principles, but an enterprise still needs to decide what to retain, for how long, and under what privacy policy.

OWASP Agent Control Standard provides current guidance on runtime visibility and control for agent systems.

Enterprise rollout sequence

Start with read-only workloads → add bounded file writes → add selected tools → add narrowly scoped network access → introduce approval gates for consequential actions → measure failures and policy violations → expand only where evidence supports it.

Safety Check

Risk → A broadly permissive sandbox becomes the default template because it is easier for developers.

Control → Use workload-specific policy, change control, centralized review of high-impact permissions, and measurable audit evidence.

Remaining risk → Governance does not compensate for a fundamentally unsafe execution boundary. Technical controls still need to be enforced.

10
Common Mistakes

1. “The agent is sandboxed, so everything is safe.”

Cause: Treating isolation as the only control. Consequence: Over-privileged tools, credentials, or network paths remain dangerous. Correction: Pair sandboxing with authorization, network policy, tool controls, approvals, and completion checks.

2. Mounting the whole developer home directory.

Cause: Convenience. Consequence: A task that only needs one repository gains visibility into unrelated data. Correction: Mount the smallest useful workspace and keep unrelated paths outside the boundary.

3. Putting production credentials into the sandbox.

Cause: Making API calls from inside generated code. Consequence: Agent-generated code can potentially read and misuse the credential. Correction: Prefer a credential-brokering adapter or narrowly scoped tool identity outside the sandbox.

4. Enabling unrestricted network access because package installation is convenient.

Cause: The first blocked dependency download creates friction. Consequence: The execution environment gains many new outbound paths. Correction: Allow only the destinations required for the task, or prebuild dependencies into the environment.

5. Reusing a stateful workspace without identity checks.

Cause: Reuse is cheaper than recreation. Consequence: Residual files or credentials can cross task boundaries. Correction: Explicitly bind reusable environments to a tenant, workload, or session identity and test the reset path.

6. Treating approval as the only security boundary.

Cause: Humans are visible and intuitive controls. Consequence: Approval fatigue or a mistaken approval can still authorize the wrong action. Correction: Keep low-risk work inside technical boundaries and reserve human approval for actions where human judgment is actually useful.

7. Logging everything without a data policy.

Cause: “More logs are always better.” Consequence: Secrets or customer data may become durable telemetry. Correction: Define fields, retention, redaction, access, and legal requirements before production rollout.

8. Skipping post-action verification.

Cause: Assuming a successful command means a successful task. Consequence: The harness reports completion when the intended outcome did not happen. Correction: Verify business or task-level completion independently of the model's narrative.

11
❓ FAQ

Q1. Is a sandbox the same thing as an AI agent harness?

No. A sandbox is an execution boundary. A harness can also manage the model loop, tool routing, approvals, context selection, state, cancellation, tracing, budgets, recovery, and completion checks. Some architectures place more responsibilities inside one runtime, but the concepts remain distinct.

Q2. Does putting an agent in a sandbox stop prompt injection?

No. A sandbox can reduce the damage that a compromised agent can cause by limiting files, commands, network access, and credentials. Prompt injection can still manipulate the model into using the capabilities that remain available, so defense in depth is required.

Q3. Should every tool execute inside the sandbox?

No. A useful architecture may execute shell commands inside the sandbox while routing external API tools through a separate application service. What matters is that each tool has an explicit trust boundary, identity, authorization policy, and input/output contract.

Q4. Is a container automatically a strong enough sandbox for untrusted agent code?

Not as a universal rule. The security properties depend on the container runtime, host configuration, privileges, mounts, exposed sockets, network, and surrounding infrastructure. Some workloads use stronger isolation such as microVMs or dedicated managed environments. Choose the boundary based on the threat model rather than the label alone.

Q5. What is the simplest sensible sandbox design for a first agent?

Start with a disposable execution environment, one explicit workspace, no unrelated host mounts, minimal network access, no production credentials, bounded runtime and output, and a clear completion check. Add more permissions only when the task demonstrates a real need for them.

12
🔗 References & Further Reading

The following primary sources were used to verify architecture and security claims. The article's explanations, examples, structure, and pseudocode are original teaching material.

Vendor and project names belong to their respective owners. This article does not imply that any one architecture is universal or that a documented control guarantees complete safety.

13
📝 Summary

  • A sandbox is an execution boundary, not the entire agent harness.
  • The model can choose an action, but the harness should enforce authority.
  • Files, commands, tools, network paths, credentials, and session state need explicit boundaries.
  • Keep powerful credentials outside the execution environment whenever the architecture allows it.
  • Isolation, authorization, approvals, verification, tracing, cancellation, and cleanup solve different parts of the problem.
  • No single sandbox eliminates prompt injection or guarantees correct agent behavior.
  • The most useful sandbox is the smallest environment that can reliably complete the intended task.

A strong agent harness does not merely ask, “Can the agent do this?” It asks, “Where will it run, what can it reach, which identity will it use, how long will that authority exist, and how will we know the result was actually correct?”

Keep learning. Keep testing. Keep the boundary explicit.

Comments