What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If an AI crawler keeps requesting pages after you disallow it in robots.txt, the key point is that the file is a request, not a technical barrier. To stop public access, require authentication or remove the material; to deny requests before they reach your site, use a CDN, WAF, firewall, or bot-management rule. Choose the control according to whether you want to discourage crawling, change search visibility, protect private content, or block traffic.

Why robots.txt may not stop a crawler

The Robots Exclusion Protocol asks crawlers to follow a site’s rules. It does not authenticate visitors or prevent a client from sending a request. RFC 9309 states, “These rules are not a form of access authorization.” A crawler that ignores the rule can still try to fetch the URL; the file alone cannot deny it. RFC 9309

That distinction matters: a robots rule can communicate a preference to compliant crawlers, but it is not a privacy control or a guarantee that content stays out of search or model training.

Choose the control that matches your goal

Goal Control What it does and does not do
Ask a compliant crawler to stop fetching selected pages A crawler-specific robots.txt rule Communicates a crawl preference; it does not force a noncompliant client to stop.
Keep a page out of Google Search An indexing control such as noindex Addresses indexing, not access. Google must be able to crawl the page to see the directive, so a robots disallow can prevent it from seeing noindex.
Make content genuinely private Remove it from public service or require authentication Restricts access rather than relying on crawler cooperation.
Deny matching requests before they reach the origin CDN, WAF, firewall, or bot-management rule Can block requests at the network edge, but rules need monitoring for false positives and unintended effects.

Google cautions that blocking Googlebot in robots.txt prevents crawling, but a page URL may still appear in search results. For privacy, Google recommends access controls such as password protection or removing the content; noindex is not a substitute for authentication. Google Search Central: Introduction to robots.txt Google Search Technical Requirements

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
SonicWall Content Filtering Service for TZ370-1 Year License (02-SSC-6565) - URL Filtering & Web Access Control for Safe, Compliant, and Productive Internet Use
  • SonicWall Content Filtering Service for TZ370 - 1 Year License (02-SSC-6565)
  • Website Access Management: Blocks access to inappropriate, unproductive, or harmful websites across more than 50 predefined categories.
  • Real-Time URL Classification: SonicWall’s cloud-based Dynamic Rating Engine keeps URL ratings accurate and up to date with no manual intervention.
  • User & Group-Based Policies: Enforce browsing rules by identity, department, or role with integration into directory services like Active Directory.
  • Easy Setup & Built-In Integration: Works natively on SonicWall firewalls—no additional hardware or endpoint software required.

Check that the rule applies to the requests you are seeing

  1. Fetch the live file for the exact site address. Check https://example.com/robots.txt or the corresponding HTTP address for the host receiving requests. The file belongs at the site’s top-level /robots.txt. Rules do not automatically carry across a different host, subdomain, protocol, or port. Google describes robots rules as applying to the host, protocol, and port where the file is hosted. RFC 9309 Google Search Central: Introduction to robots.txt
  2. Inspect the actual contents. Confirm that the intended crawler product token and disallowed path are present, correctly spelled, and served in production—not merely in a local file or an unpublished CMS setting.
  3. Check the serving and proxy layers. A CMS or managed robots feature may generate the response, while a CDN can provide its own robots behavior and separate request-blocking controls. Verify the file visible to a client and the active policy at the edge, not only the origin configuration. Cloudflare: Verified bots Cloudflare WAF custom rules
  4. Match the rule to the crawler’s purpose. Vendors can publish separate identities for different uses. OpenAI documents OAI-SearchBot for search and GPTBot for training; Anthropic documents ClaudeBot and says to apply its robots restriction on each subdomain to be covered. Check current vendor documentation before maintaining a long-lived list. OpenAI crawler documentation Anthropic: About ClaudeBot

Identify the requester before blocking it

A user-agent string is a declaration, not proof of identity; a requester can use a misleading string. Compare access logs with the vendor’s current crawler information and verification guidance where available. For Googlebot, Google recommends reverse DNS checks or matching the source IP against its published IP ranges because its user-agent can be spoofed. Google Search Central: Verify Googlebot and other Google crawlers

Identification also helps avoid blocking a function you want to keep. For example, a site owner may want to disallow a training crawler while allowing a search crawler that helps users discover pages. Review what each documented identity does before writing rules or edge policies; purposes and vendor guidance can change. OpenAI crawler documentation Anthropic: About ClaudeBot

Rank #2
Sale
TP-Link ER605, Wired Gigabit VPN Router
  • 【Five Gigabit Ports】1 Gigabit WAN Port plus 2 Gigabit WAN/LAN Ports plus 2 Gigabit LAN Port. Up to 3 WAN ports optimize bandwidth usage through one device.
  • 【One USB WAN Port】Mobile broadband via 4G/3G modem is supported for WAN backup by connecting to the USB port. For complete list of compatible 4G/3G modems, please visit TP-Link website.
  • 【Abundant Security Features】Advanced firewall policies, DoS defense, IP/MAC/URL filtering, speed test and more security functions protect your network and data.
  • 【Highly Secure VPN】Supports up to 20× LAN-to-LAN IPsec, 16× OpenVPN, 16× L2TP, and 16× PPTP VPN connections.
  • Security - SPI Firewall, VPN Pass through, FTP/H.323/PPTP/SIP/IPsec ALG, DoS Defence, Ping of Death and Local Management. Standards and Protocols IEEE 802.3, 802.3u, 802.3ab, IEEE 802.3x, IEEE 802.1q

Block unwanted requests at the edge

If the requirement is to deny traffic rather than ask a crawler to cooperate, configure a CDN, WAF, firewall, or bot-management system to block matching requests. Cloudflare documents AI bot policies and WAF custom rules; its available controls and classifications depend on current product configuration. Review the active policy and test its effect on legitimate visitors and services before relying on it. Cloudflare: Verified bots Cloudflare WAF custom rules

Do not assume a vendor default is permanent. Cloudflare documents a default change for new domains effective September 15, 2026. Check the provider’s current documentation and your own account settings, since defaults and bot classifications can change. Cloudflare: Verified bots

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
SonicWall Content Filtering Service for TZ350-1 Year License (02-SSC-1791) - URL Filtering & Web Access Control for Safe, Compliant, and Productive Internet Use
  • SonicWall Content Filtering Service for TZ350 - 1 Year License (02-SSC-1791)
  • Website Access Management: Blocks access to inappropriate, unproductive, or harmful websites across more than 50 predefined categories.
  • Real-Time URL Classification: SonicWall’s cloud-based Dynamic Rating Engine keeps URL ratings accurate and up to date with no manual intervention.
  • User & Group-Based Policies: Enforce browsing rules by identity, department, or role with integration into directory services like Active Directory.
  • Easy Setup & Built-In Integration: Works natively on SonicWall firewalls—no additional hardware or endpoint software required.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep an incident record if requests continue

Preserve enough detail to diagnose what happened and to explain the issue to your provider or qualified counsel. Keep unmodified logs where possible and record:

  • Timestamp and timezone, requested URL or path, and response status.
  • Source IP, user-agent, and request headers available in your logs.
  • The version of robots.txt served at the time and the relevant CDN, WAF, or firewall rules and events.
  • Any rate limits, challenges, or blocks applied, including when they took effect.

These details can help distinguish a declared bot identity from the actual requester and show whether a request reached the origin or was handled at the edge. Whether ignoring a robots rule creates a legal claim or remedy depends on jurisdiction and the specific facts; the protocol itself does not establish one. For a dispute, document the facts and seek jurisdiction-specific legal advice. RFC 9309

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.