Free tools Windows power users keep installed
One-click scans. No signup required.
Robots.txt can tell compliant crawlers what you prefer them to do; it cannot secure a page or guarantee that every crawler will comply. For publishers, the practical choice is not simply whether to “block AI.” Decide separately whether each provider may collect content for model training, discover pages for search, or retrieve content in response to a user. Keep those crawler preferences separate from search-indexing controls and from security for private material.
What robots.txt does—and what it cannot do
Robots.txt is a publicly readable instruction file for web crawlers, not a login system or a firewall. The Internet Engineering Task Force’s RFC 9309, published in September 2022, standardizes the Robots Exclusion Protocol: a crawler checks a UTF-8 plain-text file at a site’s top-level /robots.txt and uses its product token to find applicable groups and parseable rules. The standard says that crawlers must follow parseable rules when they successfully fetch the file, but explicitly warns: “These rules are not a form of access authorization.”
A disallowed URL can still be requested directly, and listing a path in robots.txt makes that path visible to anyone who reads the file. Put confidential or restricted content behind authentication and enforce authorization on the server. Use robots.txt to express crawler preferences, not to protect data.
Rules apply to a specific origin
The policy is served from the top-level /robots.txt for a particular origin. Google’s robots.txt documentation defines that scope as the same host, protocol, and port. For example, a file served over HTTPS on www.example.com does not automatically govern HTTP, a different port, example.com, or another subdomain. Publish and check a policy separately for every hostname you intend to cover.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Fetch failures can change crawler behavior
RFC 9309 distinguishes an unavailable file from one that cannot be reached because of server or network errors. Under the standard’s default behavior, an unavailable response such as a 4xx may permit crawling, while an unreachable file is treated as a complete disallow. Crawlers may cache the file; the RFC says a cached copy generally should not be used for more than 24 hours unless the file is unreachable.
Do not assume every crawler handles failures alike. Google says it generally treats most 4xx responses as if no robots.txt restrictions exist, while 5xx errors trigger different retry and cached-file behavior; it generally caches robots.txt for up to 24 hours and may use it longer if it cannot refresh. These are Google’s documented behaviors, not a universal rule. Check the response actually served and the relevant provider’s current documentation.
Rank #2
AI crawler controls are about different purposes
Some providers document separate crawlers for model-training collection, search discovery, and access initiated by a user. Those distinctions matter: blocking one crawler does not necessarily block the others, and the effect of a rule depends on the provider’s implementation. The examples below reflect the providers’ documentation cited here; they are not a complete list of AI crawlers or a guarantee of future behavior.
| Provider and user agent | Documented role | What blocking it means |
|---|---|---|
OpenAI GPTBot |
May collect content for use in training OpenAI foundation models. | OpenAI documents this as independent from its ChatGPT search control, so a publisher can disallow GPTBot while allowing OAI-SearchBot. |
OpenAI OAI-SearchBot |
Surfaces websites in ChatGPT search results. | OpenAI says disallowing it removes the site from ChatGPT Search answers, though pages may still appear as navigational links. |
OpenAI ChatGPT-User |
Can access websites for certain user actions; OpenAI says it is not an automatic web crawler. | Do not treat it as the ChatGPT Search opt-out. OpenAI says robots.txt rules may not apply to these user-initiated actions. |
Anthropic ClaudeBot |
Collects web content that could potentially contribute to model training. | Anthropic says restricting it signals that future materials should be excluded from its model-training datasets. |
Anthropic Claude-SearchBot |
Navigates the web to improve search-result quality. | Disabling it prevents indexing for search optimization and may reduce visibility and accuracy in user search results. |
Anthropic Claude-User |
Accesses websites in response to user queries. | Disabling it prevents retrieval in response to user questions and may reduce visibility for user-directed search. |
| Google Search crawlers | Google’s robots.txt documentation describes crawling restrictions. | Do not assume a Google robots.txt rule is a distinct control for AI model training; check current Google product documentation for the use you mean to control. |
OpenAI’s crawler documentation says a robots.txt change may take about 24 hours to affect ChatGPT search results. That is provider guidance, not a propagation guarantee for other services. Anthropic’s crawler help article, dated April 7, 2026, says Anthropic honors robots.txt directives and supports the non-standard Crawl-delay extension. It also says rules should be placed in the top-level file for each subdomain the publisher wants covered. Crawl-delay is not part of RFC 9309, so do not assume all crawlers recognize it.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsChoose what to allow before writing rules
Make the choice by provider and purpose rather than using a blanket assumption that every “AI bot” behaves the same way. For each provider, decide whether you want to allow model-training collection, search discovery, user-directed retrieval, or none of those forms of crawler access. Then consider the trade-off: restricting a search bot may reduce visibility in that provider’s search experience, while restricting a user-directed bot may prevent retrieval when someone asks a question that could surface your content.
OpenAI documents its GPTBot and OAI-SearchBot choices as independent. Anthropic likewise describes separate training, search, and user-query crawlers. That gives publishers more granular choices than a single all-or-nothing “AI” switch, but only for the user agents and behaviors each provider documents. A rule for one token is not evidence that another provider—or another bot from the same provider—will follow the same policy.
Rank #4
How to configure and verify your policy
- Set the purpose. Write down whether the goal is to limit training collection, search discovery, user-triggered retrieval, or all crawler access. Do not use a search bot’s rule as a substitute for a training control.
- Identify the exact user agent. Consult the provider’s current crawler documentation and choose the documented product token for the use you want to control. Avoid relying on broad labels such as “AI bots.”
- Check every origin. Request the top-level
/robots.txtfor each relevant hostname, protocol, and port. Confirm the file is reachable and returns the intended contents; one subdomain’s file does not cover another. - Inspect the effective rules. Look for overlapping user-agent groups, wildcard rules, CMS- or hosting-generated directives, and CDN-level blocks. Confirm the final served file—not only the setting in an administrative interface—matches your intent.
- Keep access control separate. If a page must not be available to the public, require authentication and enforce permissions server-side. A disallow rule does not stop direct requests.
- Use indexing controls for search-removal goals. If you want a URL removed from Google Search, use Google’s documented indexing controls rather than relying on robots.txt alone.
- Recheck over time. Provider policies and crawler tokens can change. Review the provider documentation and the live files for each origin when you update your policy.
Blocking a crawler does not necessarily remove a page from Google
Google warns that a URL disallowed in robots.txt may still appear in search results if Google discovers it through links. Disallowing crawling is therefore not the same as preventing indexing or removing an existing result. Choose the control that matches the outcome: use Google’s documented robots.txt guidance and indexing methods for search visibility, and use password protection when the material must be private.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.There is no universal opt-out that guarantees compliance
Robots.txt is a voluntary crawler protocol, and provider behavior is not coordinated across the AI ecosystem. The Internet Architecture Board’s AI-CONTROL workshop report, published as RFC 9969, says: “In particular, the emerging practice of using the Robots Exclusion Protocol [RFC9309] — also known as “robots.txt” — has not been coordinated between AI crawlers, resulting in considerable differences in how they treat it.”
Best Value
That makes robots.txt useful as a clearly published preference for crawlers that honor it, but not a universal enforcement mechanism. It also does not settle copyright, licensing, or other legal questions; those issues are separate from the operational behavior described here.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

