Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An AI agent can become a confused deputy when it uses its legitimate access—such as a mailbox, repository, payment system, or cloud API—to carry out an attacker’s request. The key risk is not simply that a model reads malicious text: it is that the surrounding application executes a proposed action with credentials the person or content that influenced it does not have. Preventing that escalation requires authorization checks at the point of action, not just a list of allowed tools.

What is a confused deputy?

A confused deputy is a program with legitimate authority that is tricked into using that authority for someone who lacks it. The deputy is “confused” about whose request it is serving or which resource its permission was meant to protect.

A classic example involves a compiler allowed to write usage data in a protected system directory. If a user could choose the compiler’s debug-output filename, the user might name a protected billing file. The compiler would then overwrite a file the user could not write directly, using its own permission on the user’s behalf. The underlying failure is the program’s use of its authority without preserving the boundary between its own privileges and the requester’s authority. Cosmonic’s capability-security explainer gives a concise description of the problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How can an AI agent become a confused deputy?

An agent’s application may connect it to tools that carry substantial authority: email, a code repository, a CRM, payment operations, a browser, or infrastructure APIs. Meanwhile, the agent may read content that is untrusted or partly controlled by an attacker, including web pages, emails, support tickets, documents, retrieved passages, tool results, or handoffs.

#1 Best Overall

If instructions embedded in that content influence the agent to propose an action, and the runtime executes the action using the application’s credentials without checking whether it is authorized for the relevant user, session, resource, and values, the agent can serve as the deputy. The attacker may never have direct access to the tool. The escalation happens because the system lets a less-authorized source steer a more-authorized intermediary.

This is a system-design risk, not a claim that every model or agent is inherently vulnerable. Exposure depends on what the agent can access, which inputs can influence it, and whether the application independently enforces authorization before side effects occur. The 2024 ConfusedPilot preprint describes studied risks in retrieval-augmented generation (RAG), including malicious text in modified prompts, retrieval-cache-related secret leakage, and effects on enterprise response integrity or confidentiality. Those mechanisms do not establish that every RAG deployment has the same weaknesses.

Why tool lists and schemas are not authorization

An allowed operation can still be the wrong operation

A tool allowlist determines which operations an agent can see or request. By itself, it does not decide whether a particular call is permitted. A payment tool may be approved in general while a specific transfer, recipient, or amount is not. A repository tool may be allowed while access to a particular private repository is not.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Valid arguments can still exceed authority

A schema can confirm that an amount is numeric or a destination is a string. It cannot establish that the current principal may send that amount to that destination. Syntax and type checks are useful for correctness; they are not substitutes for authorization over the concrete values.

A 2026 arXiv preprint by David Mellafe Zuvic audited pinned public-source commits for LangChain/LangGraph, LlamaIndex, and Stripe Agent Toolkit. Under the audit’s specified conditions, it reports capability gating in the audited defaults but no deterministic, fail-closed authorization of the model’s concrete argument values by default. This finding is limited to the public code and conditions examined; it does not establish that every version or private production integration behaves the same way.

RAG broadens the set of possible instruction sources

In a retrieval system, the person chatting with the agent may not be the person who influenced the document or passage it retrieves. Treating retrieved content as trusted policy can let an attacker-controlled source steer an agent that has broader credentials. RAG is not automatically unsafe, but retrieved text should not be allowed to define or override the rules that govern tool access.

What the reported attack-attempt figures do—and do not—show

The 2026 preprint also describes a companion sweep across 27 models. It reports mean task-aligned attempted unauthorized-call rates of 0.603 for cost-optimized deployment-tier models and 0.189 for flagship models. These are attempted calls in that benchmark, not successful breaches or real-world compromise probabilities. The deployment-tier aggregate has no paired confidence intervals, and the figures should not be generalized to a provider’s full model fleet or to deployed agents as a whole.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The sources reviewed do not establish how prevalent vulnerable deployments are, what share of agents can be exploited, or a market-wide loss rate. The two preprints are research preprints, not evidence of peer review or independent reproduction.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to prevent an agent from using your permissions on an attacker’s behalf

Put authorization between the model’s proposal and any side effect. The application—not the model and not instructions found in retrieved content—should decide whether the specific action is permitted.

  1. Reduce the authority available to the agent. Use narrow, task-specific credentials or capabilities. Restrict access to the smallest useful set of resources and operations rather than supplying broad ambient credentials.
  2. Keep policy separate from untrusted content. Treat emails, web pages, documents, tickets, retrieved passages, and tool results as data, not as instructions that can change access policy. Store trusted policy outside model-controlled context where possible.
  3. Authorize every proposed side effect. Before executing a call, check its operation and concrete arguments against the relevant principal and session. A decision should account for the resource, destination, amount, or other values that determine what the call actually does.
  4. Fail closed. If the policy check fails, cannot be completed, or encounters an error, deny the action rather than executing it by default.
  5. Constrain sensitive actions further. Depending on the operation, policy can impose scopes, resource allowlists, amount ceilings, and replay protection. These controls should be enforced by the application or tool boundary, not by asking the model to follow them.
  6. Use human approval for high-impact actions where appropriate. Approval can add a useful checkpoint, but it should supplement least authority and deterministic enforcement, not replace them.

How to assess an agent implementation

When comparing designs, assess the authorization boundary rather than relying on the presence of a tool registry or a permission setting. Useful questions include:

  • How broad is the authority? Does each tool have access only to the resources and operations needed for its task?
  • Is each call checked? Does enforcement evaluate every proposed action and its values, or does it only approve the tool during setup?
  • Where does policy come from? Is the decision based on trusted application policy independent of model-controlled text?
  • What happens on errors? Does a missed policy, timeout, or enforcement failure deny the action?
  • What operational friction is justified? Consider approval requirements and audit needs alongside usability. The cited sources do not provide a comprehensive independent comparison of products’ latency, cost, or usability.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.