Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test an AI sandbox by checking whether its deployed workload can cross specific, documented boundaries—not by relying on the word “sandbox” or a generic checklist. Use an authorized, disposable environment, synthetic data, and a written threat model; then probe host, tenant, control-plane, network, credential, workspace, persistence, and resource boundaries with harmless canaries and clearly defined stop conditions.

What counts as an escape vulnerability?

A sandbox’s security depends on the files, credentials, network destinations, tools, and system interfaces available to the workload in its actual deployment. Code generated by an agent can access what its environment exposes, so a product label alone does not establish which boundaries are enforced. OpenAI’s sandbox security guidance makes this point while recommending isolated compute, outbound allowlists, and keeping application keys outside the sandbox where possible.

For an assessment, define “escape” in terms of a boundary and an asset. A workload reaching a host-only file, another tenant’s data, or a control-plane API it should not access is a boundary failure. A model being influenced by untrusted content to misuse an allowed tool or transmit information is a separate agent-action risk; it can matter just as much operationally, but it is not by itself proof of an operating-system or runtime escape.

Start with the threat boundaries identified in the Kubernetes SIGs Agent Sandbox threat model and adapt them to your system: workload-to-host, workload-to-control-plane, cross-tenant, network, credential, workspace, shared-service, and resource boundaries. For each, write down what the workload may and may not reach.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you prepare a safe test?

  1. Authorize and scope it. Record the deployment and environment, workload image and runtime, tenant boundaries, test window, and every system or network in scope. Get authorization from the owners of those systems.
  2. Use a disposable environment. Keep production credentials, production data, and unrelated systems out of reach. Use synthetic test data and harmless canaries—unique markers that reveal access without exposing a real secret. Choose bounded tests that cannot exhaust shared resources.
  3. Map intended connections. Diagram the host or node, sandbox workload, orchestrator or control plane, other tenants, mounted workspaces, shared services, external network, credential broker, and MCP or other tool integrations. For each connection, record the intended allow or deny behavior.
  4. Set stop conditions before testing. Decide who can stop the run and what triggers a stop, such as a canary appearing outside its expected boundary, an attempt to reach an out-of-scope system, or resource consumption exceeding the test limit.

A nested evaluation can help contain a test: SANDBOXESCAPEBENCH describes an outer sandbox holding a flag while inner container tasks run. You can adapt the containment idea by placing synthetic canaries in an outer test boundary and treating any unexpected access as a failure signal. This is an assessment pattern, not a drop-in product or certification. The paper covers misconfiguration, privilege-allocation mistakes, kernel flaws, and runtime or orchestration weaknesses; its reported benchmark results do not establish the security of any particular deployment.

What should you test?

Before running probes, make a test matrix that states the allowed and forbidden outcome, the harmless probe, the evidence to retain, and the stop condition for every boundary. The examples below are categories to assess, not exploit instructions.

Boundary Expected denied outcome Safe probe and useful evidence
Host files and processes The workload cannot read host-only canary data or inspect host processes beyond what policy permits. Use a synthetic host canary and a benign, policy-approved visibility check. Retain the canary access result and relevant runtime or audit logs.
Tenant separation One workload cannot access another tenant’s workspace, data, or resources. Place unique synthetic markers in isolated test tenants. Check for marker visibility and retain tenant-scoped access logs.
Control plane and service identity The workload cannot call unauthorized orchestrator or internal APIs, or use an identity with excess permissions. Review service-account settings and test only against explicitly approved, synthetic endpoints. Keep policy and access-denial logs.
Network and metadata The workload cannot reach internal destinations, metadata endpoints, or unapproved external hosts. Compare attempted connections with the documented egress allowlist, using controlled destinations. Record proxy or network-policy decisions.
Credentials and brokers Application secrets are not readable by generated code, and brokers expose only the approved operations. Use dummy credentials or a test broker. Verify which values and operations are visible to the workload and retain broker audit records.
Workspace and persistence Writes stay within the intended workspace, and state does not persist beyond the documented lifecycle. Use a synthetic marker in a disposable workspace. Check the allowed write area and verify cleanup after the run using approved inspection methods.
Resource limits A workload cannot exceed its assigned CPU, memory, storage, or execution limits in a way that affects other services. Run bounded resource tests in isolation. Record configured limits, observed enforcement, and whether cleanup completed.

How do you check the deployed controls?

Inspect effective runtime settings rather than assuming a template or default is active. Compare what is running with the boundary map and the written policy.

  • Runtime and privileges: Record the runtime and version, workload privileges, capabilities, and any configured secure runtime.
  • Identity and control plane: Review service-account permissions and whether tokens are mounted. The Kubernetes Agent Sandbox documentation lists disabling service-account token mounting by default in the template path it describes; verify the configuration actually used in your deployment.
  • Mounts and workspace: Inventory host and workspace mounts, their access modes, and what survives workload termination.
  • Network: Inspect network policy, egress proxy rules, DNS and metadata reachability, and any permitted internal destinations.
  • Secrets and tools: Find where credentials are stored and which brokers, MCP servers, or other tools the agent can invoke. Check the permissions and audit trail for each path.
  • Limits and cleanup: Review resource requests and limits, lifecycle behavior, and evidence that temporary state is removed.

In its threat model, the Kubernetes SIGs project distinguishes a trusted controller/router from untrusted workload pods and lists secure runtime configuration such as gVisor or Kata Containers, managed network policy, token-mount settings, and resource requests and limits as mitigations. The project also explicitly says it does not itself implement isolation. Those details describe that project and its configuration options; they do not guarantee that a separate deployment has the same controls or outcomes. Read the project threat model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you test agent-level paths?

Run a separate test track for whether untrusted content can steer the agent toward an action that violates policy. Provide benign test content that represents an untrusted source, then check whether the agent attempts a forbidden tool call, follows an unsafe instruction, or sends a synthetic marker to an unapproved destination. Do not use real secrets or uncontrolled external recipients.

OpenAI’s prompt-injection guidance frames this risk as a source that can influence an agent combined with a sink: an action such as transmitting information, following a link, or using a tool. Record the agent’s behavior, tool requests, approvals, and resulting network or broker logs. Report a tool-mediated disclosure separately from a runtime escape so the finding identifies the control that failed.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do architecture choices affect the assessment?

Architecture changes which boundaries deserve the closest scrutiny, but an architecture label is not a security verdict. Compare the deployed configuration and integration paths, not product names alone.

Approach described by its documentation Boundary facts to account for Assessment focus
Container-based workloads Containers share the host kernel. The Kubernetes Agent Sandbox project describes secure runtimes such as gVisor and Kata Containers as configurable mitigations, while noting that the project itself does not implement isolation. Kubernetes threat model Verify the runtime actually selected, privileges, host mounts, network policy, tenant separation, and control-plane access.
Docker AI Sandboxes Docker describes a microVM with a separate Linux kernel, policy-controlled outbound TCP, and a separate Docker Engine per sandbox. Its documented layers include the hypervisor, network, Docker Engine, workspace, and credential proxy. Docker isolation layers Check the effective network and credential policies, writable workspace mounts, cleanup, and integrations. Docker notes that directly mounted workspaces are shared read-write and that local stdio MCP servers run on the host outside the VM. Docker security overview

These are documentation claims about specific architectures and configurations, not a universal ranking. Your test still needs to establish which controls are enabled and what the workload can reach.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you document findings?

For each result, preserve the deployment and runtime versions, configuration snapshots, applicable policies, test inputs, observed outputs, logs, and cleanup evidence. Classify the finding by the asset and trust boundary crossed, then fix the configuration or design and repeat the same bounded test. State exactly which deployment and configuration you assessed and what the test did not cover; passing these probes is evidence about that test, not proof that a sandbox is secure against every escape path.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.