Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →To allow or ask an AI crawler to stay away, add a User-agent group for its documented crawler token to your site’s root-level /robots.txt file, then use Allow or Disallow for the paths you want it to access. Use separate groups when you want different rules for search, user-requested retrieval, and training crawlers. These rules are requests, not access controls: enforce a block at your server, firewall, or CDN if the requests must be stopped.
How to allow or block AI crawlers with robots.txt
Edit the plain-text file at the root of the canonical host, such as https://example.com/robots.txt. RFC 9309 specifies that the file is served at the top-level /robots.txt URI and uses UTF-8 text. Put each crawler’s product token after User-agent:, then list the path rules that apply to it.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
GoolRC 939A Pocket Robot Talking Interactive Dialogue Voice Recognition Record Singing Dancing... | $22.99 | Buy on Amazon |
Ask a named crawler to stay away from the whole site
User-agent: GPTBot
Disallow: /
This asks a crawler identifying as GPTBot not to fetch paths on the site.
Allow a named crawler across the site
User-agent: GPTBot
Allow: /
This explicitly allows the named crawler under the robots rules. It does not grant access to content otherwise protected by authentication or other controls.
#1 Best Overall
- Function: Interactive communication, singing, dancing, LED light, telling story, decoration
- Smart Appearance: Robot is mini sized 85mm that you can hold it in hands
- Robot's eyes flash happily when got different commands, the arms of the robot can rotate flexibly
- Repeat Mode: pocket robot can record your voice and repeat to you with robotic sound effect, not noisy
- Conversation Mode: just talk to him, cute robot could recognize voice and reply to you, a good companion when alone.
Use a catch-all rule for other crawlers
User-agent: *
Disallow: /private/
User-agent: * applies when no more specifically named group matches. It is not an AI-only rule: it can affect any crawler that falls under the wildcard group. Choose paths carefully, especially if search-engine crawling should continue.
How to allow search while blocking training crawlers
Do not treat “AI crawlers” as a single category. An operator may use distinct crawler tokens for search and training, and a user-triggered fetch may use another identity. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for crawling related to model training, with independent settings.
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
With this configuration, the file requests access for OpenAI’s search crawler while asking its training-related crawler to stay away. OpenAI says that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, although they may still appear as navigational links. OpenAI also says its systems may take about 24 hours after a robots.txt change to adjust for search results; that interval is specific to OpenAI, not a general refresh guarantee for every crawler.
For other operators, verify the current product token and purpose in that operator’s documentation before adding a rule. Cloudflare’s crawler reference lists examples including ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, and Google-CloudVertexBot. That reference is an example inventory, not a complete or definitive registry; a token’s presence there alone does not establish every detail of its use.
How robots.txt decides which rule applies
RFC 9309 defines how compliant crawlers interpret groups and paths. Matching of the crawler product token is case-insensitive, and matching groups are combined. A duplicate named section is therefore not necessarily an isolated override. If no named group matches, a wildcard group applies if one is present.
For paths, the crawler compares rules from the beginning of the URL path and follows the most specific matching rule. If equally specific Allow and Disallow rules match, RFC 9309 resolves the tie in favor of Allow. Keep rules simple to reduce parser surprises. Google documents support for * and $ in path patterns, but do not assume every crawler implements those wildcard details the same way.
For example, the following rules allow the public section while asking the crawler not to fetch the more specific archive path:
User-agent: ExampleBot
Allow: /articles/
Disallow: /articles/archive/
Replace ExampleBot with a documented token. The example illustrates path specificity; it does not claim that an actual operator uses that token.
Which AI crawlers should you name?
Name the specific crawler whose activity you want to control, rather than assuming one rule covers every AI-related request. The purposes below reflect the available operator documentation and Cloudflare’s reference; check the current documentation for the service you care about.
| Token | Purpose or qualification |
|---|---|
OAI-SearchBot |
OpenAI identifies it with ChatGPT search. |
GPTBot |
OpenAI identifies it with crawling related to model training. |
ChatGPT-User |
Listed in Cloudflare’s crawler reference as a user-triggered request identity. |
ClaudeBot, Claude-SearchBot, Claude-User |
Listed in Cloudflare’s crawler reference; confirm current purposes and policies with Anthropic. |
PerplexityBot, Perplexity-User |
Listed in Cloudflare’s crawler reference; confirm current purposes and policies with Perplexity. |
Googlebot |
Google’s search crawler. |
Google-CloudVertexBot |
Listed as an AI crawler in Cloudflare’s reference; confirm the relevant Google documentation and purpose. |
Cloudflare’s reference also includes crawlers associated with Microsoft/Bing, Meta, Apple, Amazon, ByteDance, and Common Crawl. Treat such lists as useful starting points, not exhaustive guarantees. Crawler identities and policies can change.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose rules based on the access you want
- Search and referral visibility: Decide whether the service’s search crawler should fetch your pages. For OpenAI, opting out of
OAI-SearchBotaffects inclusion in ChatGPT search answers. - User-requested retrieval: Decide separately whether a service should fetch a page in response to an individual user’s request. Confirm the operator’s current token and behavior.
- Training-related crawling: Use the training-related product token documented by the operator; do not infer that blocking a search crawler also blocks training activity, or vice versa.
- Enforcement: Decide whether a request in robots.txt is sufficient. If not, block requests through infrastructure controls and verify that policy independently.
Does robots.txt actually stop AI bots?
No. RFC 9309 says robots rules are not access authorization. They communicate a request for compliant crawlers to follow; a crawler can ignore them. Do not use robots.txt to protect confidential, private, or otherwise restricted content. Use authentication or technical request blocking when access must be denied.
Cloudflare provides managed robots.txt functionality and separate AI Crawl Control enforcement. Those are distinct controls: a robots.txt rule expresses the site’s preference, while enforcement depends on the applicable service and configuration.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Check that the rules are live and consistent
- Fetch the root file. Request
https://example.com/robots.txt, substituting the site’s canonical host. Confirm the response contains the intended current rules, rather than only checking a local draft. - Review every matching group. Look for duplicate or conflicting sections from a CDN, hosting platform, or plugin. Matching groups may be combined, so do not assume a later duplicate group overrides an earlier one.
- Check infrastructure separately. Inspect server, CDN, firewall, and bot-management policies, along with access logs. A correct robots.txt response does not prove those systems permit or block the same requests.
- Recheck after edits. Confirm the served file again after deployment. Allow for operator-specific processing time; OpenAI’s stated approximate 24-hour adjustment period for its search results should not be applied to other crawlers.
- Review the policy periodically. Reconfirm tokens and infrastructure settings with the operators and services involved, since crawler identities and vendor policies can change.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

