Cloudflare alleged on August 4, 2025, that Perplexity used undeclared traffic to access sites after the company’s declared crawlers were blocked. Perplexity disputed Cloudflare’s interpretation, saying the traffic may have been misattributed and describing its fetching as user-triggered and real-time. The sources establish the accusation and rebuttal, not an independent finding of who was right.
What did Cloudflare say its test showed?
Cloudflare said customers had reported that Perplexity could reach their sites despite robots.txt restrictions and web application firewall (WAF) rules blocking Perplexity’s publicly declared crawlers. In response, Cloudflare said it set up newly purchased domains that were not indexed or otherwise publicly discoverable, placed disallow directives in robots.txt, and added WAF rules blocking the declared crawlers. It then asked questions about specific content on those test sites and said Perplexity answered them. Cloudflare’s August 4 report describes this methodology; it is Cloudflare’s account, not an independently reproduced test.
Cloudflare said it observed requests using both Perplexity’s declared user agent and a generic browser string that appeared to identify itself as Chrome on macOS. It attributed the latter traffic to Perplexity based on machine-learning and network signals, and said the requests came from IP addresses outside Perplexity’s published range and shifted among IP addresses and autonomous systems after blocks. Those identification and traffic claims are also Cloudflare’s.
Cloudflare characterized the behavior this way: “Although Perplexity initially crawls from their declared user agent, when they are presented with a network block, they appear to obscure their crawling identity in an attempt to circumvent the website’s preferences.” The statement appears in the company’s August 4, 2025 report and should be read as its allegation, not a neutral technical finding.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
What numbers did Cloudflare report?
Cloudflare’s report gave these estimates and product-use figures. They are company-reported observations, not independently verified industry totals.
| Figure | What Cloudflare said it represented |
|---|---|
| 20–25 million daily requests | Cloudflare’s 2025 estimate associated with the declared Perplexity-User user agent in its August 4 report. |
| 3–6 million daily requests | Cloudflare’s 2025 estimate associated with the Chrome-like user agent it labeled stealth traffic in that report. |
| Tens of thousands of domains | Cloudflare’s description of the scale of domains where it observed the activity. |
| More than 2.5 million websites | Cloudflare’s August 2025 count of sites that had chosen to disallow AI training through its managed robots.txt feature or managed AI crawler blocking rule. |
The figures describe different things: two estimates of daily requests, a domain-scale description, and a count of sites using specified Cloudflare controls. They should not be added together or treated as measurements by an independent monitor.
Rank #2
- Protects against known exploits, malware and malicious websites; detects unknown attacks; identify thousands of applications
How did Perplexity respond?
Perplexity disputed Cloudflare’s account. In a statement reproduced by Daring Fireball, the company said Cloudflare had either sought publicity or misattributed 3–6 million daily requests from BrowserBase’s automated browser service. Search Engine Land’s summary described Perplexity’s position as saying the requests were user-initiated, real-time fetches rather than preemptive crawling.
Perplexity’s response, as reproduced by Daring Fireball, said: “When you misattribute millions of requests, publish completely inaccurate technical diagrams, and demonstrate a fundamental misunderstanding of how modern AI assistants work, you’ve forfeited any claim to expertise in this space.” This is the company’s advocacy in a disputed exchange, not independent evidence resolving the attribution.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #3
- Fortinet Web Application Firewall - virtual appliance for all supported platforms. Supports up to 1 x vCPU core
- Fortinet HW FWB-VM01
- Manufacturer Part: FWB-VM01
The accounts differ on key technical questions: whether Cloudflare’s test traffic was correctly attributed, whether the requests were triggered by users or initiated preemptively, and whether another party’s browser service explains the observed traffic. The sources reviewed do not provide an independent reproduction that settles those questions. Ars Technica also covered Cloudflare’s accusation and Perplexity’s denial in its August 4, 2025 report.
Does robots.txt actually stop AI bots?
No. robots.txt is a published, machine-readable request telling crawlers which parts of a site they should avoid. It is not authentication, and it does not technically prevent a client from requesting a public URL. Whether a crawler respects the directive is separate from whether it can technically fetch the page.
Rank #4
A WAF operates differently: it can block or challenge requests at the network or application layer. Cloudflare said its test used both robots.txt directives and WAF restrictions, so the allegation was not simply that a bot ignored robots.txt. A fetch after a published disallow directive, a bypass of a configured block, the identity of the requester, and legal or contractual permission are distinct questions. The cited reports do not determine the broader legal questions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Can a website block AI crawlers with a firewall?
A site operator can use network and application controls such as WAF rules, bot-management rules, or challenges to restrict requests, rather than relying on robots.txt alone. Effectiveness depends on the site’s configuration and the traffic signals available to the service; no control should be treated as a guarantee against every changing crawler or request pattern.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Cloudflare said that, at the time of its August 4, 2025 post, it had removed Perplexity from its verified-bot list and added signatures for the traffic it described to a managed rule intended to block AI crawling. It also said customers with existing bot-management block rules were protected and could use challenge rules. These are descriptions of Cloudflare’s response and product behavior at that time, not assurance about current settings or a claim that Cloudflare is the only way to manage crawler traffic.
In a separate Cloudflare explainer on robots.txt controls, the company says its managed feature can publish directives against AI training crawlers and is designed to update as the crawler landscape changes. Such directives express site preferences; technical enforcement requires access controls. Cloudflare said its managed robots.txt feature or managed AI crawler blocking rule was being used by more than 2.5 million websites to disallow AI training in August 2025.
Quick Recap
What can publishers conclude from the dispute?
- Cloudflare made a specific allegation based on tests it says it conducted; its published account is not a court finding or an independent adjudication.
- Perplexity publicly disputed the attribution and characterized the requests as user-connected real-time fetching, but those claims are not independently established in the cited coverage.
- Robots.txt communicates a crawler preference. A firewall or bot-management control is the mechanism for blocking or challenging requests, subject to configuration and changing traffic.
- A crawler’s declared identity, the purpose of a request, compliance with site preferences, technical access, and permission under law or contract are related but separate issues.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

