What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Blocking known bad domains can stop some malicious links, but it cannot by itself keep an AI agent safe. An agent may be manipulated by instructions hidden in a webpage, document, or email; a permitted site may redirect it elsewhere; and a URL can leak information simply because the agent requests it. Protect the whole path: restrict where the agent can go, treat retrieved content as untrusted, limit what tools and data it can access, and require approval before sensitive actions.

What “phishing an AI agent” means

Phishing is a useful analogy, but one common mechanism is indirect prompt injection. Someone places instructions in content the agent later reads—a webpage, email, document, or database record. If the agent mistakes those instructions for trusted commands, it might change its response, follow a link, use a tool, or disclose information.

The risk comes from a chain: an attacker-controlled source, an agent that can access useful data or tools, and an action that carries information somewhere or changes something. A domain block addresses only part of that chain. It does not tell the agent whether text on an allowed page is trustworthy, nor does it decide whether the agent should be allowed to submit a form or send a message.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why a bad-domain list is not enough

A permitted site can lead somewhere else

A link hosted on an allowed domain may redirect to another host. A control that checks only the first hostname can therefore miss the destination that ultimately receives the request. Check the actual URL and re-check after redirects; if the destination is unknown or cannot be verified, stop, use a safer source, or ask the user.

A familiar domain can contain hostile content

A reputable service may host content supplied by other people. Even when the domain is legitimate, instructions on a particular page, in an uploaded file, or in an embedded item may be adversarial. Domain reputation cannot establish that the page’s contents are safe to follow.

A URL can carry information

Information placed in a URL’s path or query string is sent as part of the request. The receiving service may record it in server logs or analytics, and a background request can transmit it without producing an obvious reply in the chat. Do not put passwords, access tokens, private records, or other sensitive values in a URL an agent might open.

OpenAI’s January 28, 2026 description says its approach checks whether an exact URL has previously been independently observed as public. An unverified URL may require user action or a different source. OpenAI explains the rationale this way: “If a URL is already known to exist publicly on the web, independently of any user’s conversation, then it’s much less likely to contain that user’s private data.” That is a safeguard aimed at quiet disclosure through the URL itself—not a finding that the destination or its page content is harmless.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build protection in layers

1. Limit destinations to what the task needs

Use the narrowest destination controls available in the agent, browser, or organization’s network setup. Allow only the origins required for the task rather than granting unrestricted browsing by default. A destination decision should consider the final URL after redirects, not merely the initial link or its apparent brand.

For unknown or unverified destinations, define a fallback: block the request, obtain the information from an approved source, or pause for a person to decide. A blocklist can help identify known threats; an allowlist can reduce exposure to unapproved destinations. Neither should be treated as a complete verdict on content or behavior. Overly broad restrictions can also create friction that leads people to ignore warnings, so make the allowed set fit the actual workflow.

2. Treat everything retrieved as untrusted data

Keep webpage text, emails, files, and other external content separate from system and user instructions. The agent may summarize or analyze that material, but instructions found inside it should not gain authority merely because the agent can read them. Where the workflow permits, structure or sanitize inputs and use tools that distinguish retrieved data from trusted instructions.

This distinction still matters on an allowed domain. A safe destination is not proof of safe content, and a clean-looking page does not authorize the agent to follow its requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Restrict data, tools, and identities

Give the agent only the information and capabilities its task requires. Avoid exposing unrelated private data, broad account access, or powerful tools to an agent that only needs to read or summarize. Use appropriately scoped service identities and isolate the agent’s working environment where possible. If malicious content does influence the agent, these limits reduce what it can reach or do.

4. Put approval in front of consequential actions

Require a person to confirm sensitive communications, disclosure of private data, payments, and other consequential actions. A confirmation should make clear what will be sent or changed and to whom; a vague “continue?” prompt is less useful. OpenAI’s March 11, 2026 article states that “potentially dangerous actions, or transmissions of potentially sensitive information, should not happen silently or without appropriate safeguards.”

5. Test the actual workflow

Test the task flows, browser versions, tools, permissions, and destination rules you actually deploy. Include hostile instructions in content the agent is expected to read, links that redirect, and attempts to induce an unauthorized disclosure or action. Record the versions and configuration tested, and repeat the checks when they change. Google describes automated red-teaming and tracking attack success rates for its own engineering; that is Google’s account of its approach, not independent proof that any other deployment is protected.

Do not rely on a prompt that tells the agent to “be careful” as the only defense. Clear instructions can express intent, but they do not remove the agent’s access to destinations, sensitive data, or powerful actions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose controls by the risk they address

Control What it checks or limits What it does not establish
Domain blocklist or allowlist Whether a hostname is known to be blocked or permitted Whether a permitted page contains hostile instructions, or whether a link redirects elsewhere
Exact-URL verification Whether the exact URL has been independently observed as public, in the OpenAI approach described January 28, 2026 Whether the page’s content is trustworthy or safe to follow
Redirect-aware destination checks Whether the final destination remains within the permitted set Whether the destination’s content is benign
Content separation or inspection How retrieved material is treated as data rather than authoritative instruction Whether the agent is prevented from taking an action through some other route
Scoped permissions and action approvals What data and tools are available, and whether high-impact actions require confirmation Whether every attack attempt will be detected before it reaches the agent

Use more than one control because they address different failure points. A destination check can reduce unsafe navigation; content handling can reduce the influence of hostile instructions; restricted permissions and human approval can limit the consequences if earlier layers fail.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the browser research does—and does not—show

A University of Washington study page, updated April 15, 2026, reports experiments on seven agentic browsers using their latest stable versions in late January and early February 2026 on macOS Sequoia. The researchers describe a proof-of-concept cross-origin data-theft attack on ChatGPT Atlas in Agent Mode. They also say preconditions for an attack existed in Chrome with Gemini, Claude for Chrome, and Perplexity Comet if prompt injection succeeded.

Those findings describe tested versions, a specific operating system and testing window, and conditions relevant to the demonstrated attack, including framing and cookie-policy details. They do not establish that every browser or current version is vulnerable in the same way, or give a general rate of real-world agent phishing. Treat the study as evidence that browser-agent security boundaries deserve testing—not as a verdict on every deployment.

Other vendor accounts have narrower scopes too. Google’s December 8, 2025 article describes its own Chrome agent defenses. Microsoft Learn discusses an earlier Bing Chat URL-exfiltration example and says Microsoft fixed that specific issue; that historical fix should not be read as eliminating other agent risks. OpenAI’s descriptions of URL verification and safeguards around sensitive actions explain its stated design and intended scope, rather than independently validating all products or workflows.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical pre-click checklist

  1. Check the destination: Is this exact URL expected for the task? Will the control check the final destination after redirects?
  2. Check the data in the request: Could the URL itself expose private information to the receiving service or its logs?
  3. Check the content boundary: Will the agent treat instructions in the page, email, or file as untrusted material rather than commands?
  4. Check the agent’s reach: Does it have more account access, data, or tool capability than this task requires?
  5. Check the action gate: Will a person review a sensitive disclosure or consequential action before it happens?
  6. Check the tested setup: Have these protections been tested against the deployed workflow and its current browser and agent configuration?

If any answer is unclear, pause the task rather than asking the agent to proceed on trust alone. Use a verified alternative source, narrow the agent’s access, or require a human decision before continuing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.