Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control an AI agent’s tools in the application and infrastructure that execute them—not just in its prompt. Give the agent a narrowly scoped identity, check each tool call before it runs, require human approval for consequential actions, and isolate code execution, credentials, and network access. Then test and audit the controls as tools and workflows change.

1. Inventory what the agent can do

Start by listing each tool and the authority behind it. A tool name is not a permission boundary: a broadly defined “database” or “shell” tool may expose far more than the task requires. Record the actual operations, resources, and identities available to the agent.

  • What data can the tool read, and what state can it change?
  • Which account, service identity, tenant, project, or role does it use?
  • What network destinations can it reach?
  • Can its effects be undone, and what would reversal require?
  • Does it send external messages, run code, move money, or change production systems?

Separate read-only access from writes and other side effects. Narrow each tool’s schema and server-side authorization to the smallest supported action and resource scope; do not rely on a descriptive tool name to limit what it can do.

Classify risk by consequence

Use read-versus-write access, reversibility, required account permissions, and financial impact as a starting risk rubric, as recommended in OpenAI’s practical agent guide. Map each category to an actual control: automatic execution for acceptable low-risk actions, a deterministic policy check, human approval, or denial.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Give the agent a least-privilege identity

Create a distinct service account, workload identity, or equivalent for the agent instead of handing it broad user or administrator credentials. Scope that identity to the project, tenant, records, and actions needed for its task. Google Cloud’s AI security guidance recommends creating an agent identity and granting only the roles and permissions necessary to complete the task.

Remove tools the workflow does not need. This is separate from controlling how enabled tools run: Anthropic’s managed-agent documentation notes that a permission policy only governs an enabled tool, so a tool that should never be available must be disabled. Review the identity and enabled-tool list whenever the workflow or tool inventory changes.

3. Enforce rules at every tool invocation

Place authorization and validation in the trusted application, tool server, IAM policy, or equivalent execution boundary. A prompt such as “do not delete records” may guide the model, but it cannot reliably prevent a proposed deletion. The enforcement point must be able to reject the call regardless of what the model says.

Check the actual proposed action

Before dispatch, evaluate the tool, arguments, target resource, calling identity, and task scope. Apply deterministic checks for rules such as allowed hosts, file paths, record ownership, transaction limits, or production environments. Validate tool results before returning sensitive data to the model or user.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI distinguishes input and output guardrails from tool guardrails. Agent-level checks do not automatically cover every call in manager-style workflows, so put checks beside every side-effecting tool, including calls reached through handoffs or nested agents. A check on the initial request or final response does not secure intermediate invocations.

Fail safely when enforcement is unavailable

If a required policy service cannot evaluate a call, do not silently run it. Deny or pause the action and surface the failure for recovery. Apply the same principle to malformed arguments or missing authorization context rather than falling back to broader access.

4. Choose automatic execution, policy checks, approval, or denial

Match the decision mode to the consequence of the action. An automatic decision is not the same as a human checkpoint: Anthropic’s managed-agent documentation says that under auto, calls judged safe may run before anyone sees them. Its documentation recommends always_ask when a person must review every call to a tool.

Control mode What happens Use it when
Automatic execution The action runs without per-call human review. The tool is narrowly scoped and its risk is acceptable without a person reviewing each call.
Deterministic policy check A server-side rule allows, denies, or escalates a call based on its context. The decision can be made reliably from enforceable conditions such as resource scope or transaction limits.
Mandatory human approval The run waits for a person to approve or reject the proposed call before execution. Every call to that tool needs human judgment before it runs.
Denial The call is rejected and does not execute. The action is out of scope, violates policy, or cannot be safely evaluated.

Make approvals specific and fail closed

Show the reviewer the exact proposed tool, arguments, target, and relevant context. Bind the decision to that proposal; if a delayed review could make the underlying state stale, revalidate important preconditions when executing. Deny policy violations instead of asking a person to approve them. OpenAI’s guardrails guidance recommends human approval for ambiguous or high-risk actions and failing closed when review is unavailable or times out.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Isolate execution and restrict network access

Run model-directed shell or code work in isolated compute, rather than in the same environment as orchestration, production services, or unrelated users’ data. Use separate environments where users or workloads must not share data, and limit outbound connections to explicitly approved destinations.

Keep orchestration, approval decisions, audit records, billing authority, and recovery controls in a trusted harness or service where possible. The sandbox should execute the task without gaining unnecessary control over those surrounding systems. OpenAI’s sandbox security guidance warns that “Agent-generated code can access the files, credentials, and network available to its environment.”

6. Keep powerful credentials out of agent-directed code

Do not place long-lived application credentials in prompts, source code, images, or logs. Keep application API keys in the trusted application that handles tool calls. If sandboxed code needs a third-party API, broker the request through a trusted proxy or scoped secret mechanism restricted to approved destinations.

Secret injection is not a complete safeguard: code able to read an injected environment variable may expose it. Limit what credentials can do and where they can be used, and rotate or revoke them if exposure is suspected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

7. Treat content and changing tools as active risks

User submissions, webpages, database records, and MCP results are untrusted data, not instructions that should override policy. Separate those contents from system instructions and isolate memory or state across users and tenants. These controls address risks including prompt injection and unsafe tool chaining described in Google Cloud’s AI security guidance.

MCP tool inventories can change: Google Cloud notes that even trusted MCP servers may add tools dynamically. Verify server provenance, periodically review available tools, and allow only the specific tools the workflow needs. Block production reads or writes unless they are required, and reassess permissions when a tool is added or its behavior changes.

8. Log decisions, test failure cases, and revise controls

Record enough information to investigate both harmful and blocked calls: the proposed action, policy decision, approval or denial, identity, execution result, and relevant configuration or version. Protect logs themselves from containing unnecessary secrets or sensitive data.

Test permitted and rejected actions, not only the happy path. Include prompt-injection attempts, unexpected tool additions, malformed arguments, approval timeouts, and unavailable policy services. OpenAI’s practical guide recommends adding guardrails as real-world edge cases and failures emerge, while balancing protection with usability.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to evaluate an agent framework or managed platform

Documentation and feature names alone do not establish that a platform meets every security requirement. Use these questions to evaluate the controls you will actually deploy:

  • Permission granularity: Can you disable tools and scope access by tool, action, resource, user, tenant, and environment?
  • Decision modes: Is it clear whether a call runs automatically, receives a policy decision, waits for explicit approval, or is denied?
  • Coverage: Do checks run before and after each custom tool call, including nested agents, handoffs, and MCP tools?
  • Approval quality: Does the reviewer see the exact action and arguments? Can the run pause and resume, and can execution revalidate state?
  • Isolation: Can you constrain filesystem, compute, and network access to approved resources?
  • Credential handling: Can a trusted harness broker access without exposing application-wide secrets to generated code?
  • Audit and availability: Are decisions and outcomes recorded, and does the system fail closed if review or policy evaluation is unavailable?

For Anthropic managed agents, treat behavior as version-sensitive: the documentation labels the feature beta and identifies the permission-policy interface as managed-agents-2026-04-01. Check the current documentation and behavior for the version you plan to implement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.