Cloudflare said in August 2025 that it detected an undeclared crawler it attributed to Perplexity continuing to request pages after the company’s declared crawlers were blocked. Perplexity denied that account, saying Cloudflare may have mistaken traffic from BrowserBase, a third-party cloud-browser service, for its own. The dispute is unresolved in the material available here; it is not an independently adjudicated finding.
What Cloudflare said it observed
In an August 4, 2025 post, Cloudflare said customers had blocked Perplexity’s declared crawlers, PerplexityBot and Perplexity-User, using robots.txt instructions and Cloudflare rules. Cloudflare then reported seeing an undeclared crawler with a Chrome-like macOS user agent, using multiple IP addresses outside Perplexity’s published range and continuing to try to reach content after blocks were applied. Cloudflare alleged that behavior was Perplexity’s; Perplexity denied the characterization and said Cloudflare may have confused it with BrowserBase traffic.
Cloudflare also said its test domains were not indexed or publicly discoverable, yet information from them allegedly appeared in Perplexity answers after robots.txt and network blocks were in place. It attributed a pattern of roughly 3–6 million requests per day to the undeclared crawler. Those figures and observations are Cloudflare’s account, not independently verified measurements or a finding that Perplexity was responsible.
Cloudflare said it used machine-learning and network signals to identify the behavior and added matching signatures to a managed rule. Perplexity’s explanation was that Cloudflare may have attributed BrowserBase’s 3–6 million daily requests to Perplexity; Perplexity says it uses that cloud-browser service only occasionally. The company also distinguishes retrieving pages to answer a particular user’s question from crawling content for model training.
#1 Best Overall
- 【Five Gigabit Ports】1 Gigabit WAN Port plus 2 Gigabit WAN/LAN Ports plus 2 Gigabit LAN Port. Up to 3 WAN ports optimize bandwidth usage through one device.
- 【One USB WAN Port】Mobile broadband via 4G/3G modem is supported for WAN backup by connecting to the USB port. For complete list of compatible 4G/3G modems, please visit TP-Link website.
- 【Abundant Security Features】Advanced firewall policies, DoS defense, IP/MAC/URL filtering, speed test and more security functions protect your network and data.
- 【Highly Secure VPN】Supports up to 20× LAN-to-LAN IPsec, 16× OpenVPN, 16× L2TP, and 16× PPTP VPN connections.
- Security - SPI Firewall, VPN Pass through, FTP/H.323/PPTP/SIP/IPsec ALG, DoS Defence, Ping of Death and Local Management. Standards and Protocols IEEE 802.3, 802.3u, 802.3ab, IEEE 802.3x, IEEE 802.1q
Why the dispute matters to site owners
The key distinction is between a crawler’s declared identity and the traffic a site or security provider observes. Perplexity’s crawler documentation provides user-agent strings and IP ranges for its declared crawlers, along with robots.txt guidance and AWS WAF allowlisting advice. Cloudflare’s report, by contrast, described traffic with a generic browser-like user agent and IPs outside that published range. A request that does not match a declared identity is harder to attribute from its user-agent string alone.
The disagreement also concerns purpose. Perplexity says its retrieval is prompted by specific user questions and is not the same as storing pages for model training. For a publisher, however, a request’s stated purpose does not itself determine whether to permit it: the relevant policy may be whether to allow search visibility, user-directed retrieval, training, or none of these.
Rank #2
- Easier-Than-Ever Setup — Convenient and easy router management via web browser or the ASUS ExpertWiFi mobile app through Bluetooth setup.
- VLAN for Added Security —Each of the Ethernet ports can be assigned to one or more VLAN IDs that provides additional security for your business.
- Up to 3 WAN Ethernet Ports – 1 gigabit WAN port and 2 gigabit WAN/LAN ports with load balancing optimize multi-line broadband usage.
- Backup WAN for Stable Connectivity –The USB port can be used as a backup WAN by connecting it to a mobile phone with hotspot to maintain a reliable internet connection.
- Commercial-Grade Network Security and VPN — Secure public WiFi connections with Safe Browsing and VPN features. Enjoy a free-subscription ASUS AiProtection Pro, including robust intrusion prevention system (IPS) features like deep packet inspection (DPI) and virtual patching to block malicious traffic.
Robots.txt, WAF rules, and bot controls do different jobs
| Control or signal | What it does | What it cannot establish by itself |
|---|---|---|
| robots.txt | Publishes instructions for crawlers that choose to honor them. Cloudflare’s managed robots.txt feature can prepend disallow rules for known AI crawlers when a site has no robots.txt of its own. | It is not an access-control barrier; some operators may ignore its instructions. A robots.txt rule alone does not prevent a request from reaching a page. |
| WAF or edge rule | Applies a network-layer policy to requests that reach the site through the provider, such as blocking or otherwise handling matching traffic. | A declared user-agent string is not proof of identity. A rule based only on that string may not match traffic using another identity. |
| Published crawler identity | Provides a reference for comparing a declared user agent and IP range with incoming traffic. Perplexity publishes crawler details and WAF guidance. | A mismatch does not, by itself, establish who operated the request; Cloudflare and Perplexity disputed attribution in this case. |
| Behavior-based bot controls | Cloudflare’s newer AI traffic controls classify traffic by purpose categories such as Search, Agent, and Training, enabling different policies for different classes. | Classification is not the same as a universally correct answer about a crawler’s identity or intent; operators still need to choose a policy appropriate to their site. |
How to choose a policy for Perplexity and other AI crawlers
Decide separately whether you want pages available to search-oriented bots, user-directed agents, or training crawlers. A single “allow AI” or “block AI” choice can be too broad if you want search discovery but do not want training access.
- Allow search access: Identify the relevant Search category or declared search crawler and permit it under your chosen policy.
- Control user-directed retrieval: Treat Agent traffic as its own category if your preference differs from search crawling or training.
- Restrict training access: Apply a Training-specific rule rather than assuming a block on PerplexityBot will cover every crawler or every purpose.
- Use layered controls: Publish robots.txt instructions for compliant crawlers and use edge/WAF controls when you need enforcement at the request layer.
- Review scope: Determine whether a policy should apply site-wide, to selected paths, or to a particular class of pages. Check Cloudflare’s current AI Crawl Control and bot-reference documentation for the controls available to your account.
Cloudflare’s July 2026 changelog says that domains added beginning September 15, 2026 receive defaults that block Training and Agent bots on pages displaying ads while leaving Search allowed. This is a documented default for those new domains, not a claim that every existing Cloudflare zone has the same settings. Check the actual policy applied to your site rather than relying on a default described for a particular group of domains.
Rank #3
- 【Flexible Port Configuration】1 2.5Gigabit WAN Port + 1 2.5Gigabit WAN/LAN Ports + 4 Gigabit WAN/LAN Port + 1 Gigabit SFP WAN/LAN Port + 1 USB 2.0 Port (Supports USB storage and LTE backup with LTE dongle) provide high-bandwidth aggregation connectivity.
- 【High-Performace Network Capacity】Maximum number of concurrent sessions – 500,000. Maximum number of clients – 1000+.
- 【Cloud Access】Remote Cloud access and Omada app brings centralized cloud management of the whole network from different sites—all controlled from a single interface anywhere, anytime.
- 【Highly Secure VPN】Supports up to 100× LAN-to-LAN IPsec, 66× OpenVPN, 60× L2TP, and 60× PPTP VPN connections.
- 【5 Years Warranty】Backed by our 5-years warranty and free technical support from 6am to 6pm PST Monday to Fridays
Allowing Perplexity without opening the door to training crawlers
Start by deciding which Perplexity activity you intend to allow. Perplexity’s published crawler material distinguishes its search crawler from Perplexity-User, which relates to user-directed retrieval; Cloudflare’s newer controls offer the broader Search, Agent, and Training categories. The categories are not interchangeable labels, so use Cloudflare’s current bot reference and AI Crawl Control documentation to map the available identities and traffic classes to your own policy.
- Set the purpose policy. Decide whether Search, Agent, and Training traffic should be allowed, blocked, monitored, or selectively challenged. If you want search visibility but not model-training crawling, express those as separate decisions.
- Check the declared crawler details. Compare Perplexity’s documented user agents and IP ranges with the identity signals your controls use. Perplexity also provides AWS WAF allowlisting advice for site owners seeking reliable access for its declared crawlers.
- Use robots.txt for instructions. Set rules for the crawlers you choose to address, but do not treat robots.txt as a technical block against a crawler that disregards it.
- Apply edge rules for enforcement. Where you need requests blocked or handled at the network layer, configure the relevant Cloudflare or AWS WAF policy for the intended bot class or identity, rather than relying on robots.txt alone.
- Validate the result. Review the traffic and rule outcomes after deploying the policy. A generic browser-like user agent or an IP outside a published range should not be treated as conclusive proof of Perplexity attribution; investigate the signals available to you before assigning responsibility.
What the incident does—and does not—prove
Cloudflare’s report describes a specific pattern it said its systems detected and says it added signatures to a managed rule. Perplexity offered a competing explanation involving BrowserBase and rejected the inference that the attributed traffic was its crawler activity. The evidence summarized here does not establish a court ruling, independent audit, or definitive resolution of that dispute.
Rank #4
- Compact and Efficient Design: The FortiGate 40F is designed for small to mid-sized businesses and enterprise branch offices, featuring a compact, fanless desktop form factor that ensures quiet operation and minimizes space usage.
- Robust Connectivity Options: Equipped with 5 GE RJ45 ports, including 1 WAN port and 4 internal ports, this model provides essential connectivity and flexibility for various network configurations in a small-scale environment.
- High-Performance Security: Offers up to 1 Gbps IPS throughput and 600 Mbps threat protection throughput, using Fortinet’s purpose-built security processor technology to deliver industry-leading performance and protection for SSL encrypted traffic.
- Advanced Threat Protection: Integrated with Fortinet’s AI-powered FortiGuard Labs, the FortiGate 40F offers comprehensive cybersecurity, identifying and mitigating both known and unknown threats to maintain robust security across your network.
- Simplified Management and Deployment: Features a user-friendly management console that provides comprehensive network automation and visibility, coupled with Zero Touch Integration with Fortinet’s Security Fabric for easy deployment.
For practical purposes, site owners do not need to settle the attribution dispute before setting a policy. Robots.txt communicates preferences to compliant crawlers; WAF and edge rules control requests at a different layer; published identities help with matching; and behavior-based controls provide another signal. Use each for the job it can perform, and avoid treating any single user-agent label as proof of who is behind a request.
Quick Recap
Best Value
- 【Flexible Port Configuration】1 Gigabit SFP WAN Port + 1 Gigabit WAN Port + 2 Gigabit WAN/LAN Ports plus1 Gigabit LAN Port. Up to four WAN ports optimize bandwidth usage through one device.
- 【Increased Network Capacity】Maximum number of associated client devices – 150,000. Maximum number of clients – Up to 700.
- 【Integrated into Omada SDN】Omada’s Software Defined Networking (SDN) platform integrates network devices including gateways, access points & switches with multiple control options offered – Omada Hardware controller, Omada Software Controller or Omada cloud-based controller(Contact TP-Link for Cloud-Based Controller Plan Details). Standalone mode also applies.
- 【Cloud Access】Remote Cloud access and Omada app brings centralized cloud management of the whole network from different sites—all controlled from a single interface anywhere, anytime.
- 【SDN Compatibility】For SDN usage, make sure your devices/controllers are either equipped with or can be upgraded to SDN version. SDN controllers work only with SDN Gateways, Access Points & Switches. Non-SDN controllers work only with non-SDN APs. For devices that are compatible with SDN firmware, please visit TP-Link website.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

