A system prompt can tell an AI agent what it should do; it cannot reliably prevent the agent from doing something else. If an agent reads malicious instructions in a webpage, email, or document, it may misuse tools that are available to it. Enforce permissions in the code and environment that execute those tools, then test whether those controls hold when the model is manipulated.
Why an agent’s written rules are not a security boundary
An agent may receive developer instructions alongside task data, then use tools to act on what it reads. That data can include attacker-controlled directions embedded in an otherwise ordinary webpage, email, or file. NIST calls this kind of manipulation agent hijacking: malicious content attempts to steer an agent away from its intended task by exploiting the difficulty of separating trusted instructions from untrusted information.
The practical risk depends on what happens after the model is steered. If it has only read access to a narrow set of documents, a bad decision may be limited. If it can send messages, modify records, run commands, or reach sensitive credentials, the same kind of mistake can have much greater consequences. As Anthropic puts it in its response to NIST on agentic security, “Agent security is a property of the whole system, not just the model.”
Prompt-injection filters, instruction hierarchy, and labels that identify untrusted content can help, but they are not authorization controls. OWASP explicitly cautions that labeling untrusted input alone does not enforce a security boundary. OpenAI’s March 11, 2026 guidance likewise emphasizes that manipulation can rely on context and social engineering, not just recognizable hostile strings. Design as though a model-level defense can fail.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Enforce permissions where tools execute
The model can propose a tool call, but ordinary application code should decide whether that call is allowed. At the execution boundary, check who or what is making the request, which resource it targets, what operation it requests, and whether the arguments are valid. Do not let a model-generated statement such as “the user approved this” expand its own authority.
Give each task the smallest useful tool set
Expose only the operations and resources the task needs. Separate read-only interfaces from write-capable ones; avoid broad or wildcard access. A research agent that needs to read selected files should not automatically inherit permission to edit them, execute arbitrary processes, or access unrelated accounts. OWASP’s AI Agent Security Cheat Sheet recommends least privilege and authorization checks outside the model.
Rank #2
Bind approval to the actual action
Require review for actions that are sensitive, irreversible, financial, administrative, or externally visible. Show the reviewer the precise operation and its parameters—for example, the recipient and message of an email or the record and values being changed—not a vague request to approve “the agent’s next step.” An approval prompt is useful only if it authorizes a specific action rather than becoming a general permission slip.
Constrain the runtime and its connections
Limit the files, processes, credentials, and network destinations reachable from the agent’s execution environment. Use isolation appropriate to the task, such as a sandbox or container, and restrict outbound network access to destinations the task requires. Keep secrets outside the runtime when they are not needed: a prompt injection cannot retrieve credentials that the agent’s environment cannot reach. Anthropic describes this containment principle in its account of how it contains Claude across products.
Recommended Free Tools
Validate data at every handoff
A trusted connector can return untrusted content because the source it reads may be attacker-controlled. Validate tool inputs and the resulting action rather than assuming that an approved integration makes its content safe. Treat model output as untrusted when passing it to another system, too: use parameterized database queries, safe rendering, and other controls suited to the destination.
Keep authorization independent across agents
In a multi-agent system, the receiving service should check its own permissions before acting on another agent’s request. Validate inter-agent messages, but do not mistake message authenticity for authorization. OWASP states: “A valid message signature does not grant permission to perform the requested action.”
Rank #4
Describe the authority an agent actually has
NIST’s 2025 tool-use taxonomy offers teams a vocabulary for comparing deployments. It distinguishes both the agent’s level of tool capability and whether its environment is trusted or untrusted. NIST presents the taxonomy as something to adapt, not as a definitive standard or a ranking of products.
| Dimension | Category | What to establish for your deployment |
|---|---|---|
| Tool capability | Read-only | Which resources can the agent inspect, and can any available path change state? |
| Tool capability | Constrained-write | Which specific changes are allowed, on which resources, and what validation or review applies? |
| Tool capability | Write | Which systems can the agent modify, and what limits prevent an erroneous action from spreading? |
| Environment | Trusted or untrusted | Which files, processes, credentials, inputs, and network destinations can the runtime reach, and which are outside its boundary? |
Use those questions to document the real deployment, not just the model’s intended role. Include connectors, tool descriptions, credentials, and any downstream service that can turn an output into an action. Also establish how tool calls and authorization decisions are logged, how access can be revoked, and how the agent can be stopped; these determine whether a problem can be investigated and contained.
Free tools Windows power users keep installed
One-click scans. No signup required.
Test the deployed boundaries, not just the prompt
Build abuse cases around every external content channel the agent reads and every tool that can change state or send information. For each case, define the legitimate task, the prohibited outcome, and the observable evidence that would show whether the attempt succeeded. Use dummy data and instrumented or sandboxed tool substitutes rather than exposing production systems during adversarial tests.
- Try direct prompt injection and indirect instructions embedded in webpages, emails, documents, or tool results.
- Test harmful or malformed tool arguments, attempts to reach data outside the task, and paths that could exfiltrate information.
- Attempt privilege escalation and bypasses of approval or other review gates.
- Check each side-effect path, including connectors and downstream systems, rather than testing only the model’s text response.
- Repeat attempts and vary the attack. A control that blocks a familiar string may still fail when an attack changes its wording or context.
NIST’s Center for AI Standards and Innovation recommends adaptive evaluations because resistance to known attacks does not establish resistance to new ones. Its January 2025 evaluation discussion describes task-specific testing and multiple attempts as informative, while its experiments used then-current models and AgentDojo-derived scenarios. Those results should not be read as a current, universal failure rate for agents. OWASP also warns that its sample prompt-injection smoke tests are illustrative, not a representative security benchmark.
Read vendor benchmark results as limited evidence
Vendor figures can describe a particular system and evaluation, but they do not prove that another deployment is safe. In material accessed October 7, 2026, Anthropic reports that Claude Opus 4.7 had roughly 0.1% attack success on single attempts and roughly 5–6% after 100 adaptive attempts on Gray Swan’s Agent Red Teaming benchmark. Anthropic also reports that Claude Code auto mode catches roughly 83% of “overeager behaviors” before execution. These are vendor-reported results for the named systems and evaluations, not independent comparisons or guarantees for agents generally.
Anthropic’s broader point about containment is useful when interpreting any such result: “The failure is identical. The consequences are not.” A benchmark can inform a risk assessment, but only tests of your actual tools, permissions, data, and runtime can show how a failure could affect your deployment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

