Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Prompt injection becomes a serious security problem when an AI model can do more than write text. A malicious instruction hidden in a web page, document or email can only cause real harm through the tools the agent is allowed to use. The practical question is therefore not only whether a model can be misled, but what a misled model is able to do, and which control decides whether that action goes ahead.

What prompt injection is

OpenAI defines prompt injection as a third party misleading a model by inserting malicious instructions into the conversation context. Its safety guidance, Understanding prompt injections, describes it as “a type of social engineering attack specific to conversational AI.” The attacker does not need access to the system. They only need to place text where the model will read it.

Direct and indirect injection

In a direct injection, the hostile text comes from the person typing to the assistant. In an indirect prompt injection, the text sits inside material the agent reads on its own: a page it browses, a file it summarizes, an email in a mailbox it triages, or the output of another tool. NIST uses the term agent hijacking for this pattern, describing it as indirect prompt injection in which malicious instructions inserted into data an agent ingests lead it toward unintended, harmful actions. That is the form most relevant to agents, because the user never typed the attack and may never see it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Not every injection leads anywhere harmful. A chatbot with no tools can be pushed into a misleading or inappropriate reply, but it cannot itself send a message, change a record or complete a purchase. The risk grows with capability.

The three layers that decide the outcome

A safe design keeps three layers distinct. Mixing them is where most agent failures start.

Layer What it contains Who controls it Failure if treated as trusted
Untrusted content Web pages, emails, documents, search results, tool outputs Third parties and anyone who can publish to the agent’s inputs Carries injected instructions into the agent’s context
Model decision The model’s proposed next step, including which tool to call and with what arguments The model, shaped by everything in its context Can be steered toward a harmful tool call by injected text
Authorized action Whether the tool call actually runs, against which resource, for which actor The execution layer and its policy, outside the model Must be the only thing that grants permission

The middle layer is where injection operates. The third layer is where defenses have to operate. A model that follows an injected instruction is not, by itself, an authorization decision.

Why tools turn a bad answer into an action

OWASP lists prompt injection, tool abuse or privilege escalation, and data exfiltration among the core risks for AI agents. Its Excessive Agency entry (LLM06:2025) makes the link explicit: when an agent holds more permissions than its task needs, including excessive third-party tool permissions, a manipulated agent has more room to cause harm. The same entry includes an indirect injection example, which is the pattern most teams should model first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read access and data leaving the system

A read-capable agent is not harmless. If it can open confidential files and also send outbound messages or call an external URL, an injected instruction can try to move that data somewhere else. This is the data exfiltration risk OWASP names. The leak does not need a write permission; the ability to transmit is enough.

Write and execution

Write-capable tools can change records, send messages, delete content or trigger workflows. Execution tools are more sensitive still. OWASP’s MCP Top 10 entry MCP05:2025, Command Injection and Execution, describes how untrusted input that reaches a command or code execution path can compromise the host. Its suggested mitigation is an allowlist of permitted commands rather than trying to filter every malicious string.

Comparing two designs

The table below is illustrative, not a measured test. It shows how the same injected email could produce very different outcomes depending on the permissions around the model.

Axis Summarizer with read-only folder access Assistant that can read and send mail
Data and tool permissions Read one labeled folder; no outbound tools Read the full inbox and send email to any address
Read, write, execute or transmit Read only Read and transmit
Authorization enforced at execution Recommended, limits the folder scope Required for each send, checked against the recipient and the actor
Confirmation for sensitive actions Not needed for summaries Required before any external send
Sandboxed tool operation Optional Recommended for the sending path, using a test mailbox during evaluation

The first design can still produce a misleading summary. The second can send data outside the organization, which is the consequential failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Authorization belongs in the execution layer

An instruction generated by the model should never be the thing that grants permission. Classification, system prompts and confirmation dialogs all help, but none of them is an access control. The check that matters runs in the component that actually performs the action, and it should evaluate:

  • The actor: which user, service or session is asking for the action.
  • The action: the specific tool and operation, such as “send email” rather than “use the mail tool.”
  • The resource: the recipient, file path, account or record affected.
  • The policy: an allowlist of permitted actions and destinations where one can be defined.
  • The confirmation state: whether a human has approved this specific call when the action is sensitive.

OWASP’s AI Agent Security Cheat Sheet describes this pattern, with authorization checks applied at execution for each action. OpenAI’s guidance likewise points to user confirmation and sandboxing as controls around consequential actions.

Practical controls that reduce risk

Grant only the data and tools the task needs

Scope each agent to one clearly defined task. A calendar-scheduling agent does not need access to the full mailbox archive, and a document summarizer does not need network egress. Excessive agency is the condition that turns an injection from a bad answer into a data problem, so removing unused permissions is usually the highest-value change.

Separate read access from consequential actions

Where the architecture allows it, keep read tools and write or transmit tools in different components with different permissions. An agent that reads untrusted content should not automatically hold the credentials needed to send, delete or pay.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Require review before sensitive actions

Route sending information, completing a purchase, changing access rights and similar actions through explicit human review. Show the reviewer the exact action and its parameters, not a general summary, so that an injected recipient or amount is visible before approval.

Sandbox code and tools that can cause change

Run code execution and tools that could make harmful changes in a sandbox with restricted network access and limited file scope. A sandbox limits what a successful injection can reach; it does not stop the injection from being read.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to test indirect injection safely

Testing should follow the same path that untrusted content takes in production. The OWASP LLM Prompt Injection Prevention guidance recommends dummy data and sandboxed tool substitutes so that test payloads cannot cause real-world effects.

  1. Create test documents, web pages and emails that contain planted instructions, using dummy data only.
  2. Deliver them through the real input channel the production agent uses, such as the same browser tool, parser, or mailbox connector.
  3. Replace consequential tools with substitutes. For example, point a send-email tool at a logging stub or a test mailbox instead of a live address.
  4. Run tasks the application actually performs, with the same permissions the production agent would hold.
  5. Record whether the agent proposes a harmful tool call, and whether the execution layer blocks it.
  6. Repeat the tests whenever a tool, permission or input source is added.

What these controls do not guarantee

Scoping permissions, enforcing authorization outside the model, separating read and write paths, requiring review and sandboxing all reduce the chance or the impact of a successful injection. None of the cited guidance from OpenAI, OWASP or NIST supports a claim that these steps prevent every injection. The sources also do not give a reliable production prevalence figure for indirect injection, so the risk should be treated as a design constraint rather than a measured rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The durable rule is simple: the model may read anything, but it should only be able to do what an independent authorization check permits for that specific actor and action.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.