Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI agent containment is the practice of limiting what an agent can access and what damage it can cause if it behaves unexpectedly or follows malicious instructions. It is not a single setting: use a narrowly permissioned identity, isolate agent-directed execution, restrict files and network access, protect credentials, review consequential actions, and have a tested way to stop work and revoke access. These controls limit an agent’s capabilities; they do not guarantee that the model will resist every prompt injection or make every safe decision.

What containment does—and what it cannot do

Instructions and model safeguards can influence what an agent tends to do. Permissions and environment boundaries determine what it can actually reach. Anthropic’s security guidance distinguishes those model-layer defenses from environmental controls and cautions against relying on model safeguards alone.

That distinction matters when an agent reads webpages, documents, or tool output. Such content can contain prompt injection: malicious instructions embedded in otherwise ordinary material. If the agent follows them, it may misuse tools it is legitimately authorized to call. Containment is therefore about limiting the consequences of a mistake or manipulation, not proving that the model will never be manipulated.

A useful design question is not simply, “Will the agent follow its instructions?” Ask instead, “If it does something unintended, which files, services, people, and systems can it affect?”

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Build containment in layers

Each layer addresses a different route to impact. A sandbox does not compensate for a powerful credential inside it, and a confirmation prompt does not compensate for unrestricted access if the prompt is bypassed.

Layer What it limits What to check
Agent identity and tool permissions Which resources and operations the agent is authorized to use Roles, API scopes, connected tools, delegated agents, and credential lifetime
Execution isolation Which processes, files, and host resources model-directed work can reach Mounts, privileges, writable paths, persistence, and separation from orchestration
Network controls Which external destinations the agent can contact Whether egress is disabled, allowlisted, or open, and whether indirect routes are possible
Credential handling Whether agent-directed code can read or expose secrets Secrets in environment variables, files, mounts, logs, or tool responses
Human review and monitoring Whether sensitive actions are blocked pending scrutiny and whether activity can be reconstructed Approval gates, audit visibility, responsible approvers, and incident procedures

1. Give each agent a bounded identity

Create a distinct identity for each agent or workload rather than sharing a broadly privileged service account. Grant only the roles, files, endpoints, and operations needed for that task. Apply the same least-privilege rule to connected tools and delegated sub-agents: a restricted top-level identity does not help if a tool or child agent has wider authority.

Google Cloud recommends an agent identity with only necessary roles. Google’s Gemini documentation also recommends least-privilege credentials and short-lived tokens where available. Limit each token’s resource and API scope, rotate it, and revoke it if exposure is suspected.

2. Keep orchestration separate from execution

The control plane, or harness, commonly handles model calls, tool routing, approvals, tracing, run state, and recovery. The execution plane is where model-directed work reads or writes files, runs commands, installs packages, or uses mounted data. OpenAI’s Agents SDK guidance describes this separation and warns that placing the harness and execution in the same compute boundary puts orchestration alongside model-directed activity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep sensitive application authentication, billing controls, audit records, and recovery functions outside the execution environment where possible. If the agent can alter the system that records or stops its activity, the containment boundary is weaker.

3. Restrict files, processes, and egress

A container, virtual machine, or hosted sandbox can limit process and filesystem access, but the label “sandbox” does not tell you how strong the boundary is. Inspect what the agent can see and change: host mounts, repositories, artifacts, prior-session data, user privileges, persistence, and exposed ports.

Configure network egress separately from filesystem isolation. Google’s documentation for its managed-agent environment describes OS isolation but says outbound networking is unrestricted by default; allowlists can restrict or disable that access. That is a product-specific default, not a general rule for all Google services or agent environments. OpenAI’s sandbox security guidance likewise recommends restricting network access and isolating workloads.

Consider whether DNS or other indirect network paths could defeat the intended destination restrictions. The point is to verify the effective boundary, not infer it from a product name or a single configuration toggle.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Keep secrets outside agent-readable environments

If agent-generated code can read a credential, unexpected behavior or prompt injection may cause the code to use or expose it. OpenAI cautions that injecting a stored secret into an environment still makes it accessible to agent-generated code in that environment.

Where practical, keep application-wide keys out of the sandbox. A trusted proxy or credential broker can make a narrowly scoped request to an approved destination without disclosing the underlying secret to the agent. Review tool responses and logs too: credentials can leak through outputs even when they were not mounted as files.

5. Treat outside content as data, not authority

Webpages, user documents, database records, and tool results should be treated as untrusted input. Google Cloud advises treating user-provided and database-derived content as data rather than instructions. OpenAI describes prompt injection as an evolving challenge and recommends layered defenses.

Keep the task narrow, limit the data and tools the agent needs, and constrain reachable destinations. For consequential actions, require confirmation and show the reviewer the target, requested operation, and relevant information that will be shared. Detection can help surface suspicious content, but it does not replace limits on what the agent can do.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Use approval gates selectively

Require approval for actions with meaningful consequences, such as sending external communications, changing production data, making purchases, or moving money. The action should remain technically blocked until approval arrives; an instruction asking the model to “wait for approval” is not itself an enforcement mechanism.

Anthropic reported that users approved roughly 93% of Claude Code permission prompts in its 2026 telemetry. That is a vendor-reported figure for that product, not an industry-wide estimate. Anthropic warns that frequent prompts can reduce attention. Google Cloud also notes that human-in-the-middle approval can fail when people approve malicious or destructive suggestions without proper verification.

Make the approval view specific enough to support a real decision: identify the destination or record, the change or message, and the data being sent. Avoid prompting for every low-risk tool call, which can train reviewers to approve without checking.

Choose an architecture by its actual boundary

An in-process tool runner, container, VM, or hosted sandbox cannot be ranked safely by name alone. Product labels do not establish how much the agent can reach, and the reviewed vendor materials do not provide an independent head-to-head benchmark ranking these options. Compare the configured boundary and its operator instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Boundary enforcement: Is isolation enforced by an operating-system or virtualization boundary, or does safety mainly depend on agent instructions?
  • Filesystem and data exposure: Which host paths, repositories, mounts, artifacts, and prior-run data are visible or writable?
  • Credential boundary: Can the agent read the secret itself, or does a trusted service broker a scoped request?
  • Network egress: Is outbound access disabled, allowlisted, or open? Can DNS or another indirect route bypass the restriction?
  • Control-plane separation: Are model calls, approvals, audit logs, credentials, and recovery functions outside agent-directed compute?
  • Persistence and cleanup: What survives a run, who can resume it, and how are credentials and queued calls invalidated?
  • Operational visibility: Can responders inspect a human-readable timeline of tool calls, permission changes, and external effects?
  • Human intervention: Which actions require confirmation, who can authorize them, and can the agent bypass the gate?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make the kill switch an incident procedure

A kill switch is not useful merely because a stop button exists. The deployment needs an owner, a known control path, and a tested understanding of what stopping does. The Cloud Security Alliance’s May 2026 rapid research note recommends incident-response procedures that include kill-switch activation protocols and clear accountability. The note is AI-assisted rapid research, not a primary regulator standard, and the reviewed sources do not establish a universal technical design or response-time standard.

Define and test a deployment-specific stop sequence. A practical sequence to evaluate is:

  1. Stop execution: identify the control that disables the active run or worker, and verify whether in-flight processes actually terminate.
  2. Block further actions: disable tool routing and network egress so a stopped or resumed process cannot continue making external calls.
  3. Revoke access: revoke or expire credentials that could remain usable beyond the run, including tokens held by connected services.
  4. Check queued work: determine whether scheduled, retried, or already queued tool calls can still execute, and invalidate them if needed.
  5. Preserve evidence: retain the relevant tool-use sequence and privilege changes so responders can reconstruct what happened.
  6. Assign authority: document who can invoke each step, where the controls are, and how the team communicates that the agent has been contained.

These are design and testing questions inferred from credential, network, and incident-response controls—not a universal specification. Test the path under realistic conditions rather than assuming that ending a user-facing session also revokes access or cancels queued work.

How to judge claims about agent safety

Vendor-reported model or product results can inform a deployment decision, but they are not substitutes for testing the configured system. Anthropic reported roughly 0.1% attack success on single attempts and around 5–6% after 100 adaptive attempts for Claude Opus 4.7 on Gray Swan’s Agent Red Teaming benchmark, as well as roughly 83% detection of “overeager behaviors” by Claude Code auto mode. These are vendor-reported, product- and benchmark-specific results; they should not be compared across vendors without matched independent testing, and they do not establish that a particular deployment is contained.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a real deployment, assess the identity, data, tools, network, credentials, approval gates, and recovery path the agent will actually have. Then test whether the controls block the unwanted action, not merely whether the model recognizes it as risky.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.