Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build the agent so the model can propose an external action, but a trusted execution layer decides whether that action is allowed. Before every call, enforce the current actor’s authorization, the permitted service and resource, the operation, and its parameters. Then apply least privilege, action-specific approval where needed, safe logging, and adversarial testing as parts of the same workflow—not as instructions the model is expected to obey.

1. Define what the agent is allowed to do

Start with the task, not with a list of tools you happen to have. Inventory the external services the agent needs, the data it may read, and the effects it may create. For each service, distinguish retrieval from actions such as drafting, sending, updating, deleting, spending, or changing access.

Write down both the allowed boundary and explicit prohibitions. Specify relevant targets—for example, which mailbox, records, repositories, or audiences are in scope—and identify actions that are externally visible, hard to reverse, or consequential. OWASP recommends giving an agent only the tools and permissions needed for its task; separating mailbox reading from message sending is one practical example.

Classify actions by impact and recoverability

Risk depends on what an action can affect and how well the result can be recovered, not just on the tool’s name. OWASP offers the following illustrative classification; it is not a universal standard, so adapt it to your data, impact, and recovery model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Example action Illustrative risk level
Document search and reading Low
Writing a file Medium
Sending email or executing code High
Deleting database data or transferring funds Critical

2. Give the workflow a constrained identity

A tool’s availability is not authorization. Avoid giving an agent a person’s broad, general-purpose credentials when the service can support a narrower agent or delegated identity. Use service-supported identity mechanisms with the smallest useful scopes, and limit access to the resources and operations the task requires.

Where supported, prefer short-lived credentials restricted to the intended audience and task over static, broadly reusable secrets. NIST warns that static API keys and bearer tokens can be presented by whoever obtains them, and that API keys may grant broad, unscoped access without granular authorization. OAuth 2.0, SPIFFE, JWT, and X.509 can provide starting points for identity and credentials; adopting a protocol by itself does not define which actions an agent should be authorized to take.

Keep credentials out of model-visible prompts, tool results, and routine logs. Arrange for the trusted execution component to obtain or use secrets without exposing their values to the model.

3. Enforce authorization at the execution boundary

Treat a model-generated tool call as a request, not as permission. A deterministic execution service—or the downstream service itself—should validate every call before it runs. Check the current actor, the agent or session identity, the tool, the target resource, the requested operation, and normalized parameters against current policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not let a model instruction, a model-generated risk score, or a bare user_confirmed flag substitute for that check. The enforcement component should make its decision independently of the model’s claim that an action is safe or approved. Where the downstream service supports authorization checks, use them as well; a check made only earlier in a workflow can become stale if the actor’s permissions or the requested action changes.

Make capabilities narrow by construction

  • Expose only the tools needed for the task, with read and write capabilities separated where practical.
  • Limit each tool to the required resources and operations instead of granting access across an entire account or service.
  • Prefer a specific, validated function to open-ended shell or URL access when the task can be done without those broader capabilities.
  • Validate and normalize parameters before use, including the destination and target resource. Recheck authorization against the normalized action that will actually execute.

4. Treat external content as untrusted data

Email, web pages, documents, tool descriptions, and API responses can contain instructions intended to redirect an agent. That risk remains even when the assigned task is only to summarize or process the content. NIST describes this form of indirect prompt injection as agent hijacking.

Keep trusted policy separate from retrieved content where the architecture allows. One design option is to parse untrusted material in a component that has no action tools, then pass constrained data to a privileged planner or execution path. OWASP also describes quarantined parsing and capability tracking as possible approaches, while noting their threat-model assumptions and implementation limitations. Neither architectural separation nor a model-based guardrail replaces narrow permissions, input validation, and enforcement at the action boundary.

5. Require approval for consequential actions

Use human approval for actions with significant external effects, such as sending, deleting, spending, deploying, or changing permissions. Keep routine, low-risk work within a clearly defined policy so that users are not asked to approve every ordinary step; a high volume of low-value prompts can encourage reflexive approval.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bind approval to the action that will run

Show the reviewer the operation, destination, target resource, and relevant parameters—not a vague request to approve what the agent plans to do. Bind approval to the current actor and exact action, give it an expiry, prevent replay, and check and consume it atomically immediately before execution. If the target or parameters change, require fresh approval. Unknown or unclassified high-risk actions should fail closed rather than proceeding on an assumption.

An approval interface is useful only if the execution boundary verifies that the approval still applies to the call being made. OWASP’s agent security guidance emphasizes checking approval for the current actor and exact tool call at that boundary.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

6. Log activity without creating another secret store

Record enough security-relevant information to reconstruct a workflow: who initiated it, which agent or session acted, the tool and target invoked, whether policy or approval allowed the call, and its outcome. Send security logs to a system outside the agent’s control.

Do not log credentials or secrets, and avoid retaining unnecessary sensitive prompt or response content. Add rate limits and alerts for patterns such as unusual destinations, unexpected tools, bulk operations, or repeated failures. If audit logging is a required condition for a consequential action, do not execute that action when the logging system is unavailable.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Test the workflow against realistic attacks

Test ordinary tasks as well as hostile instructions embedded in realistic emails, documents, web pages, and service responses. Include repeat attempts: an agent that resists one malicious instruction may still be induced to act after retries or a changed context.

Measure whether the attack caused a prohibited external action, not merely whether the model produced suspicious text. Include task-specific cases, then refresh the tests when tools, permissions, or workflows change. NIST CAISI’s evaluation overview recommends adaptive evaluation that accounts for task-specific attack performance and may measure success across multiple attempts.

Review the controls as one system

For each workflow, verify that the tool set and identity are limited to the task, the execution boundary checks every call, untrusted content cannot expand capabilities, consequential actions receive action-specific approval, and operators can detect and investigate failures without exposing secrets. A weakness in any one of these stages can undermine the rest.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.