Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An AI agent can be manipulated through an email, webpage, document, or tool result it reads. The danger is not just that it might produce a bad answer: if it has permission to send messages, access files, run code, or change records, a successful manipulation may turn into an action. Whether that happens depends on the agent’s design, available tools, and permissions—not on a malicious string alone.

What “AI hacks AI” means—and what it does not

The phrase describes two different developments. In one, a human security tester uses an AI system to help automate bounded penetration-testing tasks. In the other, an attacker puts misleading instructions in material an AI agent processes, trying to make the agent misuse capabilities it already has. The first is about using AI as a security tool; the second is about attacking an AI-enabled application.

Neither, by itself, shows that a general-purpose agent can independently break into real organizations at scale. Research on RedTeamLLM, for example, proposes a summarize/reason/act framework for automating penetration-testing tasks and evaluates it on entry-level but non-trivial capture-the-flag challenges. That is evidence about bounded task automation, not proof of autonomous criminal intrusions or large-scale zero-day discovery.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The practical change is that an agent may do more than generate text. OWASP describes agents as systems that can reason, plan, use tools, maintain memory, and take actions. Their security boundary therefore includes integrations and permissions, not just the model’s responses.

How an email or webpage can hijack an agent

Indirect prompt injection occurs when attacker-controlled instructions are embedded in material an agent may read, such as an email, webpage, file, or tool result. NIST’s Center for AI Standards and Innovation (CAISI) describes the weakness as a failure to reliably separate trusted instructions from untrusted task data. In its January 17, 2025 technical blog, CAISI characterized agent hijacking as malicious instructions inserted into data the agent may ingest, causing unintended and harmful actions.

The attack path is easier to understand as a sequence:

  1. Attacker-controlled content enters the task. A user asks an assistant to summarize a page or review an email, and that content contains misleading instructions.
  2. The agent processes the content in context. The application may present external text alongside legitimate instructions, and the model may treat some of that text as directions rather than merely data.
  3. The agent has an action available. A connected tool might send a message, share a file, execute code, modify a record, or make another change.
  4. A side effect occurs only if the pieces line up. The model must be influenced, the application must permit the action, and the agent’s tools and permissions must allow it. A malicious string does not automatically produce execution.

OWASP illustrates the risk with a personal assistant that can read email and also has a sending-capable plugin. A malicious email tries to persuade the assistant to search messages for sensitive information and forward it. The relevant security question is not just whether the assistant recognizes the email as suspicious; it is whether the application has given it a path to send anything without the user’s review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What an agent can do depends on its permissions

Reading a message and sending one are different authorities. A read-only mail integration can expose private content if misused, but it cannot itself forward that content. Giving the same assistant a sending tool creates a separate route from reading to disclosure. Similar distinctions apply to viewing versus editing files, drafting versus publishing, and inspecting code versus executing it.

OWASP’s AI Agent Security Cheat Sheet identifies a broad set of risks, including prompt injection, tool abuse and privilege escalation, data exfiltration, memory poisoning, goal hijacking, excessive autonomy, high-impact action abuse, cascading failures, supply-chain attacks, sensitive-data exposure, and denial-of-wallet risks. This is a risk taxonomy and guidance for defenders, not evidence that every category has occurred as a deployed incident.

For a specific assistant, ask what it can actually reach and change:

  • Which files, messages, accounts, or services can it read?
  • Can it write, send, share, purchase, execute, or delete—or only prepare a proposed action?
  • Are permissions limited to the task, or do they cover broader data and actions?
  • Does it retain memory or state that could affect later tasks?
  • Can the user inspect and approve consequential actions before they take effect?

What security evaluations have found

Published evaluations show that agent-hijacking attacks can succeed under test conditions, but their results should not be read as real-world compromise rates. The following figures come from different test settings and measure different things.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Evaluation Reported result What the result describes
NIST CAISI, 2025 11% for the strongest baseline attack; 81% for the strongest novel attack Attack success on held-out Workspace tasks in an evaluation of an upgraded Claude 3.5 Sonnet agent, using attacks developed for that model.
UK AI Security Institute (AISI), 2025 22 agents across 44 scenarios; more than 60,000 successful policy violations among 1.8 million submitted attacks A competition in which participants submitted prompt-injection attacks against agents in realistic deployment scenarios. Reported violations included unauthorized data access, illicit financial actions, and regulatory noncompliance.
UK AISI, 2025 Policy violations appeared for most tested behaviors within 10–100 queries Behavior observed in the competition benchmark; it is not a prediction of how many attempts would compromise a deployed organization.

The CAISI comparison illustrates why the attack method matters: a baseline attack and a novel attack adapted to the tested model produced very different results in that evaluation. It does not establish that agents generally have an 81% compromise rate. Likewise, the AISI counts describe attacks and violations in a competition, not field incidents.

AISI’s study summary also reports limited correlation between robustness and model size, capability, or inference-time compute in its evaluation. A larger or more capable model, on its own, is therefore not a sufficient safety argument; robustness has to be assessed in the context of the agent and its tasks.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to reduce the risk of an agent acting on manipulation

The strongest practical defense is to limit the consequences of a hijack, including when detection fails. No single filter can guarantee that an agent will correctly distinguish every instruction from every piece of untrusted content.

Give each tool the narrowest useful authority

  • Grant only the tools and data access required for the task.
  • Use per-tool scopes rather than broad access to an entire account or service.
  • Prefer read-only capabilities when the job is to search, summarize, or review.
  • Remove unnecessary actions, such as sending or deleting, instead of relying on the model not to use them.

In the email example, OWASP recommends a read-only mail capability and scope, removing unnecessary sending functionality, and having the assistant draft a message for the user to review and send.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Put approval and reversibility around consequential actions

Require a person to confirm high-impact actions, especially when they transmit sensitive information, spend money, change access, execute code, or alter important records. Where possible, make the action inspectable before it takes effect and provide a way to undo it. OpenAI describes using source-to-sink analysis to identify sensitive third-party transmissions, then showing users what would be sent for confirmation or blocking the transmission. It also describes sandboxing certain agent features to detect unexpected communications. These are OpenAI’s reported design approaches, not a universal standard or an independent guarantee of safety.

Keep external content separate from authority

Treat text from emails, websites, documents, and tool results as untrusted data, even when it is relevant to the task. Application design should avoid granting that content authority to override system or developer instructions. Because manipulation can depend on context rather than one recognizable phrase, input filtering alone is not a complete defense.

Monitor use and limit the damage from repeated attempts

Log tool calls and consequential actions, watch for unexpected behavior, and apply rate limits where appropriate. These measures can make abuse easier to detect and constrain its impact. OWASP cautions that monitoring and rate limiting do not, by themselves, prevent excessive agency.

Test the actual agent, tools, and tasks

Security testing should cover the deployed configuration—not only a model in isolation. Include task-specific adversarial cases, try attacks adapted to the target, and test whether an injection can move from content the agent reads to an action it can take. NIST CAISI recommends adaptive evaluations, multiple attack attempts, task-specific analysis, and shared evaluation frameworks. OWASP calls for adversarial testing and regression checks as systems change.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

After changing a prompt, tool, permission, model, or workflow, repeat the relevant tests. A fixed list of known attack strings can miss attacks that exploit a particular task or integration; evaluation should examine both whether the agent is manipulated and what its available tools allow it to do afterward.

What the evidence says about autonomous cyberattacks

There is evidence that AI can help automate bounded offensive-security work and that manipulated agents can violate policies in controlled evaluations. That is different from demonstrating that a general AI agent can independently compromise real organizations at scale. RedTeamLLM’s CTF evaluation, the CAISI Workspace test, and the AISI competition each have defined tasks and settings; their results should stay attached to those settings.

For defenders, the most useful question is concrete: if this particular agent is fooled, what can it do next? The answer lies in its tools, data access, and permissions as much as in the model. Restrict those capabilities, require review where consequences matter, and test the full workflow against adaptive attacks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.