Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can ask individual AI providers not to use your public website content for training by adding their documented crawler tokens to your site’s root robots.txt. There is no single rule that opts you out of every AI service, and training crawlers may be separate from crawlers used for search discovery or user-requested retrieval. Choose a policy for each provider and purpose, then check the provider’s current documentation and your live file.

How to stop AI bots from scraping your website for training

Use a separate User-agent group for each provider-specific crawler you want to restrict. For example, this illustrative file requests that several named crawlers not access the site while allowing one search crawler:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

This is not a universal AI opt-out or a recommended setting for every site. It shows the format only; confirm each token, supported rules, paths, and effects in the operator’s current documentation. Put the groups in the root-level robots.txt, such as https://example.com/robots.txt, and avoid assuming a generic User-agent: * rule expresses the same purpose-specific choice.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide what you want to block for each provider

Before adding a rule, identify whether the crawler is associated with model training, search discovery, or fetching content in response to a user. A provider may use distinct crawlers for these purposes. Blocking a training-oriented crawler does not necessarily block the provider’s search crawler; blocking the latter may reduce whether your pages can appear in that search product.

#1 Best Overall
Notary Privacy Guard Suitable for Journal of Notarial Events
  • No more exposed information in unprotected notary journals. This product shields clients' confidential information from prying eyes. It allows the Notary Public to keep the journal open during the transaction, as NO prior client information is viewable.
  • Shields clients' AND Notaries Public' confidential information
  • GLBA and HIPAA require strict confidentiality policies and procedures. Notary Privacy Guard is a compliance tool for the professional Notary Public.
  • Decreases Notary Public's liability from exposing client information
  • Journal column headers are printed on the Notary Privacy Guard, no having to peek underneath to complete the journal entry. Becomes part of the journal and also acts as a place marker.

Robots rules express a crawler preference, not a guarantee that content will never be collected or used. They also do not establish that content collected earlier has been deleted or that a model has been retrained. The documented controls here concern crawler behavior going forward, not retroactive remedies.

Which AI crawler tokens should you use?

Use only the token and policy documented by the relevant operator. The following distinctions are important when deciding what to allow or disallow:

Rank #2
Notary Privacy Guard Suitable for Dome Notary Journal
  • No more exposed information in unprotected notary journals. This product shields clients' confidential information from prying eyes. It allows the Notary Public to keep the journal open during the transaction, as NO prior client information is viewable.
  • Shields clients' AND Notary Publics' confidential information
  • GLBA and HIPAA require non-disclosure policies and procedures. Notary Privacy Guard is a compliance tool for the professional Notary Public.
  • Decreases Notary Public's liability from exposing client information
  • Journal column headers are printed on the Notary Privacy Guard, no having to peek underneath to complete the journal entry. Becomes part of the journal and also acts as a place marker.
Operator Documented token or distinction What the setting is for Search visibility implication
OpenAI GPTBot and OAI-SearchBot are separate crawlers. See OpenAI’s crawler overview and publisher and developer FAQ. GPTBot relates to potential training use; OAI-SearchBot supports surfacing websites in ChatGPT search. OpenAI says the settings are independent. Blocking OAI-SearchBot can affect appearance in ChatGPT search. Blocking GPTBot alone is a separate choice.
Google Google-Extended is a standalone robots.txt token. See Google’s crawler documentation. It controls whether content Google crawls may be used for certain Gemini model training and grounding. Google says Google-Extended does not affect inclusion in Google Search or act as a Search ranking signal.
Anthropic ClaudeBot. See Anthropic’s crawler guidance. Anthropic documents it as a potential training crawler and provides a robots.txt method to block it. The cited guidance establishes the training-crawler control; check Anthropic’s current documentation for any other crawler identities and product effects.

These names and uses can change. Cloudflare’s bot reference separates AI crawlers, AI search crawlers, and assistant categories, illustrating why a copied blocklist may become incomplete or misclassify a crawler. Check provider documentation and, where available, your server logs rather than relying on an old list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can you block AI training but stay in ChatGPT search?

OpenAI documents separate controls for the two purposes: GPTBot relates to potential training use, while OAI-SearchBot supports ChatGPT search discovery. If you want to request that training crawler not access your site while allowing search discovery, use a GPTBot disallow group and do not disallow OAI-SearchBot. OpenAI describes these settings as independent in its crawler overview.

Allowing a crawler in robots.txt does not guarantee that your pages will appear in search: it only avoids that particular robots.txt restriction. OpenAI also notes that updated robots.txt settings may take time to be reflected in search systems, so do not expect an immediate change.

Does Google-Extended block your site from Google Search?

No. Google describes Google-Extended as a standalone robots.txt token concerning use of crawled content for certain Gemini model training and grounding. Its crawler documentation says, “It has no effect on Google Search or other products.” That statement refers to Google-Extended’s effect on Search and other products; it does not mean that every Google crawler or Search control is covered by this token.

Rank #4
BTSFTOGET Password Book Refill Pages 212 Replacement Pages Internet Log Book, 8.2x5.6in, Large Print 576 Entries Durable Divider with Alphabetical Tabs, For Men Women Seniors Home Office Use
  • Value Pack: Our password keeper refill comes with 216 pages 80gsm paper and 12 durable laminated dividers with alphabetical tabs. for password organizer section each page has 3 entries, total allows 576 records of website, username/ID, password/hint, name, phone, email, security questions/notes etc., 12 pages/72 records of Software. license number and purchase date etc., 6 lined pages for important things to remember, 1 page for emergency information and 1 PVC protect film.
  • Premium Quality: 80gsm off-white paper which will protect your eyes from strong lights and viewing strain and allows smooth writing and reducing ink leakage, erase fraying and shade issue. 12 film laminating durable dividers with alphabetical tabs for easy scrolling of your search.1pc PVC sheet protects all inner pages from wetting.
  • Fits A5 6 Ring binder: Fits binder cover No smaller than 6.7" W x 9.25" H x1" Thick. Divider is 5.6” W x 8.19” H, inner page is 5.2" W x 8.2" H, 212 pages/106 sheets, both sides printing, 6 holes punched (hole space is 0.75in/19mm, hole space between 3rd & 4th holes is 2.76in/70mm, dia 0.197in/5mm). Suggest match this large print passwords book refills with A5 lockable binder whose size is large than 9.25"x6.7"x1" for perfect combination.
  • Pairs perfectly with our hardcover refillable password book with lock B09FGZ1CDF. Keep your important internet passwords, website, username/ID, password/hint, name, phone, email, security questions/notes. license and purchasing date with this pack of refill pages for perfect internet password keeper,huge space to store all your passwords and account & website login details in one place ,fully protect your personal privacy and keep online website account information & user data safe.
  • 100% money back if you're not satisfied with our products. Any questions, don't hesitate, just contact us!

If your goal is to change how a page appears in Google Search or prevent its Search indexing, use the relevant Google Search controls rather than treating Google-Extended as a Search-removal setting. Google’s AI features guidance says those Search features follow Googlebot controls and discusses preview controls including nosnippet, data-nosnippet, max-snippet, and noindex. Choose a control for the result you want; Search presentation and AI-training preferences are different decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does robots.txt stop AI training?

No. robots.txt is a request to compliant crawlers, not an access-control mechanism. Google notes that some crawlers may not obey it, and disallowing a URL does not itself keep that page out of Search. A publicly accessible page may still be reached by systems that ignore the request, or by other means.

Best Value
Privacy Notary Journal - All States - 200 Entries with Privacy Guard
  • 8 ¾ x 11 inches, spiral bound soft cover.
  • Block style entry, 200 entries per journal
  • Privacy guard protects client information
  • For use in any state
  • Electronic & remote notarization option

For confidential material, require authentication or remove it from public access. If your concern is search indexing or snippets, use the search engine’s supported indexing and preview controls; a robots.txt disallow rule is not a substitute for those controls.

How to add and verify crawler rules

  1. Choose the desired outcome. For each provider, decide separately whether to allow training-related crawling, search discovery, or another documented retrieval use.
  2. Check the operator’s current instructions. Confirm the exact user-agent token, applicable paths, and stated consequences. Do not infer a provider’s entire crawler policy from one bot name.
  3. Edit the root robots.txt. Add a distinct group for each token you are configuring, using the documented directives. Review existing groups so a broad rule does not undermine the decision you intended to make.
  4. Check hosting and edge rules. Review CDN, firewall, host, and bot-management settings as well. An allow rule in robots.txt does not force an edge firewall to serve a crawler. Cloudflare describes tools for managing crawler policies and bot categories in its bot-management documentation.
  5. Verify the live response and monitor logs. Fetch the public /robots.txt URL after deployment and confirm it returns the intended rules. Check server or CDN logs for requests from the relevant crawler where those logs identify it. Recheck after site migrations or policy changes.
  6. Allow for propagation. Do not assume a changed file immediately alters a provider’s search results or other systems. OpenAI says updated robots.txt settings may take time to be reflected in search.

Keep training controls separate from search controls

Use crawler-specific robots rules when the goal is to communicate a preference to a named compliant crawler. Use authentication for private content, and use a search engine’s supported indexing or preview controls when the goal is to change Search results. These mechanisms solve different problems; one should not be treated as a substitute for another.

Quick Recap

Bestseller No. 1
Notary Privacy Guard Suitable for Journal of Notarial Events
Notary Privacy Guard Suitable for Journal of Notarial Events
Shields clients' AND Notaries Public' confidential information; Decreases Notary Public's liability from exposing client information
$9.95
Bestseller No. 2
Notary Privacy Guard Suitable for Dome Notary Journal
Notary Privacy Guard Suitable for Dome Notary Journal
Shields clients' AND Notary Publics' confidential information; Decreases Notary Public's liability from exposing client information
$9.95
Bestseller No. 5
Privacy Notary Journal - All States - 200 Entries with Privacy Guard
Privacy Notary Journal - All States - 200 Entries with Privacy Guard
8 ¾ x 11 inches, spiral bound soft cover.; Block style entry, 200 entries per journal; Privacy guard protects client information
$30.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.