Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A system prompt can tell an AI agent what it should do; it cannot reliably prevent the agent from doing something else. If an agent reads malicious instructions in a webpage, email, or document, it may misuse tools that are available to it. Enforce permissions in the code and environment that execute those tools, then test whether those controls hold when the model is manipulated.

Why an agent’s written rules are not a security boundary

An agent may receive developer instructions alongside task data, then use tools to act on what it reads. That data can include attacker-controlled directions embedded in an otherwise ordinary webpage, email, or file. NIST calls this kind of manipulation agent hijacking: malicious content attempts to steer an agent away from its intended task by exploiting the difficulty of separating trusted instructions from untrusted information.

The practical risk depends on what happens after the model is steered. If it has only read access to a narrow set of documents, a bad decision may be limited. If it can send messages, modify records, run commands, or reach sensitive credentials, the same kind of mistake can have much greater consequences. As Anthropic puts it in its response to NIST on agentic security, “Agent security is a property of the whole system, not just the model.”

Prompt-injection filters, instruction hierarchy, and labels that identify untrusted content can help, but they are not authorization controls. OWASP explicitly cautions that labeling untrusted input alone does not enforce a security boundary. OpenAI’s March 11, 2026 guidance likewise emphasizes that manipulation can rely on context and social engineering, not just recognizable hostile strings. Design as though a model-level defense can fail.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Enforce permissions where tools execute

The model can propose a tool call, but ordinary application code should decide whether that call is allowed. At the execution boundary, check who or what is making the request, which resource it targets, what operation it requests, and whether the arguments are valid. Do not let a model-generated statement such as “the user approved this” expand its own authority.

Give each task the smallest useful tool set

Expose only the operations and resources the task needs. Separate read-only interfaces from write-capable ones; avoid broad or wildcard access. A research agent that needs to read selected files should not automatically inherit permission to edit them, execute arbitrary processes, or access unrelated accounts. OWASP’s AI Agent Security Cheat Sheet recommends least privilege and authorization checks outside the model.

Bind approval to the actual action

Require review for actions that are sensitive, irreversible, financial, administrative, or externally visible. Show the reviewer the precise operation and its parameters—for example, the recipient and message of an email or the record and values being changed—not a vague request to approve “the agent’s next step.” An approval prompt is useful only if it authorizes a specific action rather than becoming a general permission slip.

Constrain the runtime and its connections

Limit the files, processes, credentials, and network destinations reachable from the agent’s execution environment. Use isolation appropriate to the task, such as a sandbox or container, and restrict outbound network access to destinations the task requires. Keep secrets outside the runtime when they are not needed: a prompt injection cannot retrieve credentials that the agent’s environment cannot reach. Anthropic describes this containment principle in its account of how it contains Claude across products.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate data at every handoff

A trusted connector can return untrusted content because the source it reads may be attacker-controlled. Validate tool inputs and the resulting action rather than assuming that an approved integration makes its content safe. Treat model output as untrusted when passing it to another system, too: use parameterized database queries, safe rendering, and other controls suited to the destination.

Keep authorization independent across agents

In a multi-agent system, the receiving service should check its own permissions before acting on another agent’s request. Validate inter-agent messages, but do not mistake message authenticity for authorization. OWASP states: “A valid message signature does not grant permission to perform the requested action.”

Describe the authority an agent actually has

NIST’s 2025 tool-use taxonomy offers teams a vocabulary for comparing deployments. It distinguishes both the agent’s level of tool capability and whether its environment is trusted or untrusted. NIST presents the taxonomy as something to adapt, not as a definitive standard or a ranking of products.

Dimension Category What to establish for your deployment
Tool capability Read-only Which resources can the agent inspect, and can any available path change state?
Tool capability Constrained-write Which specific changes are allowed, on which resources, and what validation or review applies?
Tool capability Write Which systems can the agent modify, and what limits prevent an erroneous action from spreading?
Environment Trusted or untrusted Which files, processes, credentials, inputs, and network destinations can the runtime reach, and which are outside its boundary?

Use those questions to document the real deployment, not just the model’s intended role. Include connectors, tool descriptions, credentials, and any downstream service that can turn an output into an action. Also establish how tool calls and authorization decisions are logged, how access can be revoked, and how the agent can be stopped; these determine whether a problem can be investigated and contained.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test the deployed boundaries, not just the prompt

Build abuse cases around every external content channel the agent reads and every tool that can change state or send information. For each case, define the legitimate task, the prohibited outcome, and the observable evidence that would show whether the attempt succeeded. Use dummy data and instrumented or sandboxed tool substitutes rather than exposing production systems during adversarial tests.

  • Try direct prompt injection and indirect instructions embedded in webpages, emails, documents, or tool results.
  • Test harmful or malformed tool arguments, attempts to reach data outside the task, and paths that could exfiltrate information.
  • Attempt privilege escalation and bypasses of approval or other review gates.
  • Check each side-effect path, including connectors and downstream systems, rather than testing only the model’s text response.
  • Repeat attempts and vary the attack. A control that blocks a familiar string may still fail when an attack changes its wording or context.

NIST’s Center for AI Standards and Innovation recommends adaptive evaluations because resistance to known attacks does not establish resistance to new ones. Its January 2025 evaluation discussion describes task-specific testing and multiple attempts as informative, while its experiments used then-current models and AgentDojo-derived scenarios. Those results should not be read as a current, universal failure rate for agents. OWASP also warns that its sample prompt-injection smoke tests are illustrative, not a representative security benchmark.

Read vendor benchmark results as limited evidence

Vendor figures can describe a particular system and evaluation, but they do not prove that another deployment is safe. In material accessed October 7, 2026, Anthropic reports that Claude Opus 4.7 had roughly 0.1% attack success on single attempts and roughly 5–6% after 100 adaptive attempts on Gray Swan’s Agent Red Teaming benchmark. Anthropic also reports that Claude Code auto mode catches roughly 83% of “overeager behaviors” before execution. These are vendor-reported results for the named systems and evaluations, not independent comparisons or guarantees for agents generally.

Anthropic’s broader point about containment is useful when interpreting any such result: “The failure is identical. The consequences are not.” A benchmark can inform a risk assessment, but only tests of your actual tools, permissions, data, and runtime can show how a failure could affect your deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.