iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
There is no single expert consensus that one threat makes every AI agent unsafe. OWASP and NIST map concrete attack paths, including indirect prompt injection, hijacking and excessive agency. OpenAI emphasizes limiting the damage an attacker can cause even if manipulation succeeds. The AI Now Institute argues that agents handling untrusted data should not be used in certain sensitive settings when they also have powerful capabilities. Their disagreement is best understood as a difference in focus, confidence in mitigations and tolerance for residual risk—not as two opposing camps.
Why experts can look at the same agent and rank its risks differently
An agent is more than a model that generates text: it can receive information from outside sources and use tools to take actions. Security analysis can therefore begin at different points in the same chain.
- Application design: OWASP’s Excessive Agency guidance asks what tools and permissions a developer has granted, and whether those capabilities exceed the task’s needs.
- Trust boundaries: NIST’s agent-hijacking work focuses on how untrusted content—such as text in a document or webpage—can influence an agent that treats it as an instruction.
- Attack consequences: OpenAI describes the risk in terms of a source that can influence an agent and a “sink,” or action capability, that makes that influence consequential.
- Deployment suitability: AI Now Institute asks whether the remaining risks make whole categories of sensitive use inappropriate, even with safeguards.
These are related questions, not mutually exclusive explanations. Untrusted content can manipulate an agent; excessive permissions can make that manipulation more damaging; and an organization can still disagree about whether the resulting risk is acceptable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What is the biggest security risk of AI agents?
The most useful answer is often the combination of untrusted input and excessive capability. Prompt injection matters because it can turn content the agent was meant to read into an instruction it follows. Tool access, permissions and autonomy determine what the agent can do as a result.
#1 Best Overall
OWASP’s LLM06:2025 Excessive Agency guidance illustrates the combination with a mail assistant. An attacker can place malicious instructions in an email; if a plugin lets the agent both read and send mail, the agent might be influenced to send a message. The issue is not only whether a model detects the injected text. It is also whether reading mail required permission to send it.
OWASP’s Agentic Applications Top 10 announcement, published in December 2025, identifies behavior hijacking, tool misuse, and identity or privilege abuse among the concerns addressed by the project. It says more than 100 security researchers, industry practitioners, user organizations and technology providers contributed input. That figure describes the project’s input process, not how common attacks are.
Can an AI agent be hacked through an email or webpage?
It can be exposed to an attack through content it reads, but exposure is not the same as a guaranteed compromise. In an indirect prompt-injection attack, hostile text is embedded in material the agent is asked to process—a message, file or webpage, for example. If the agent follows that text and has a relevant tool, the attempted manipulation may lead to an unintended action.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchNIST’s Center for AI Standards and Innovation (CAISI) describes agent hijacking as indirect prompt injection and reports evaluations in simulated Workspace, Travel, Slack and Banking environments. Those evaluations help examine attack paths in controlled settings; they do not establish how every commercial agent behaves or prove that an agent is safe in a live deployment.
OpenAI’s technical perspective adds that advanced attacks are not usually caught by input-classifying firewall systems. That makes a defense based only on detecting suspicious text fragile: the system should also limit the impact of instructions it mistakenly follows.
Threats, likely consequences and the question each framing can miss
| Threat or failure mode | What can happen | Useful counter-question |
|---|---|---|
| Indirect prompt injection or agent hijacking | Untrusted content in an email, file or website influences the agent to take an unintended action. | Beyond detecting hostile text, what permissions, tools and destination controls limit the consequences? |
| Excessive agency | An integration provides more functions, permissions or autonomy than a task requires. A read-mail task may also expose send-mail capability. | Could ordinary external content become the attacker’s delivery channel for those capabilities? |
| Tool misuse or identity and privilege abuse | An agent misuses a legitimate tool or acts through an identity with more access than it needs. | Are credentials scoped and access limited, or is the failure being treated only as a model problem? |
| Oversight failure | A person approves an action without meaningful review or cannot intervene effectively under time pressure. | Does approval reveal the actual action and its scope, and can the reviewer stop it? |
| Use in a sensitive setting | A misdirected or compromised agent can affect security-critical work or other high-consequence decisions. | Do the task’s consequences and reversibility justify this level of access and autonomy? |
The counter-questions are analytical prompts, not claims that a named organization has overlooked a specific issue. They show why an emphasis on one layer should not be mistaken for a complete threat model.
Can human approval make AI agents safe?
Approval can reduce risk, but the label “human in the loop” does not establish that a person is meaningfully controlling an action. A reviewer needs to see what the agent will do, which account or data it will affect, and the relevant scope—not merely a reassuring natural-language summary. They also need a practical opportunity to reject or change the action.
Recommended Free Tools
OWASP recommends manual review for sending mail in its example of an agent with mail-reading and sending capabilities. AI Now Institute warns that oversight can be weakened by automation bias and prompt fatigue. Taken together, these perspectives support treating review as one safeguard within a broader design, not as a substitute for limiting capability.
Why the evidence does not settle the disagreement
The cited work includes different kinds of evidence, which support different conclusions:
- Threat taxonomies and guidance: OWASP’s Top 10 is a community-developed security taxonomy. It helps teams identify and discuss failure modes; it is not a population-level study of incident frequency.
- Simulated evaluations: NIST CAISI’s AgentDojo work examines agent hijacking in simulated environments. Results can inform evaluation design, but they do not guarantee performance across products, tasks or deployments.
- Vendor research: OpenAI presents a technical position on designing agents to resist prompt injection. Anthropic reports observations from its own products and API. These are informative perspectives, but their scope is not every provider’s systems.
- Policy analysis: AI Now Institute takes a precautionary position on sensitive uses, drawing on its research interpretation. A policy recommendation about acceptable deployment is not the same thing as an estimate of attack probability.
Anthropic’s February 2026 report, “Measuring AI agent autonomy in practice,” describes millions of human-agent interactions across Claude Code and its public API. It says nearly 50% of the agentic activity analyzed was software engineering, and that the longest-running Claude Code sessions rose from under 25 minutes to over 45 minutes over three months. These findings describe Anthropic’s observed activity, not the whole agent market. The company also notes that there is no agreed definition of an agent, API requests cannot reliably be grouped into sessions, and providers have limited visibility into customer architectures.
In that same report, Anthropic says roughly 20% of new-user Claude Code sessions used full auto-approval, compared with more than 40% among experienced users. Those figures describe sessions in that product context, not the prevalence of auto-approval across all agent deployments. Anthropic also says most public API agent actions it observed were low-risk and reversible; it does not give a universal estimate for deployed agents.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteWhere the strongest recommendations overlap
The perspectives differ on how much residual risk is acceptable, but several practical safeguards follow from their shared concern with the path from input to action:
Best Value
- Start with capability: Inventory tools, permissions, identities and autonomous actions. Remove access the task does not need; consider read-only or narrower alternatives.
- Handle outside content as untrusted: Familiar sources such as a coworker’s email or a commonly used website can still contain hostile instructions.
- Make consequential actions reviewable: Show the specific action and its scope, and make approval a genuine opportunity to intervene.
- Limit downstream damage: Monitor actions and consider rate limits. OWASP describes monitoring and rate limiting as ways to limit damage, not as measures that prevent excessive agency.
- Evaluate the intended use: Test in conditions resembling the actual tasks, tools and data. Simulated evaluation is useful, but it does not by itself establish safety in deployment.
- Scale safeguards to consequences: A reversible, low-impact task is different from an agent with broad access or the ability to make irreversible changes. The cited sources do not support a universal rule that every agent is safe—or unsafe—in every setting.
What the disagreement means for deployment decisions
For a deployer, the practical question is not simply whether an agent can resist prompt injection. Ask what could happen if it fails to resist, which permissions make that outcome possible, how readily the action can be reversed, and whether a reviewer can meaningfully stop it. A system with narrow access and reversible actions presents a different decision from one that processes untrusted data while holding powerful privileges.
AI Now Institute’s July 2026 policy brief, “Friendly Fire,” recommends against using AI agents that ingest untrusted data when they can execute arbitrary code, access security-critical environments, feed unsanitized output into automated pipelines, or inform security- and safety-critical decisions. That is the institute’s stated policy position—not a consensus rule adopted by OWASP, NIST or all security experts. Its value in the debate is to make the acceptable-risk question explicit: some organizations may judge residual risk unacceptable even when safeguards reduce it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

