What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes—often. A news publisher can block a crawler used for AI training while allowing a separate crawler used for search, but the rules are service-specific. Blocking Googlebot is different: Google says it affects Search, Discover, Google News, and other Google products. Decide whether you want to restrict training use, search discovery, page previews, or indexing before changing robots.txt.

What to decide before blocking a crawler

“AI crawler” can refer to bots with different purposes. A crawler may fetch pages to support a search service, to collect content for potential model training, or for other uses. Blocking one does not automatically block the others.

There is also a separate question: even when a search crawler can access a page, what may the search service index or display? Crawl access, search-result inclusion, and preview text are distinct controls. Choose the outcome you want before editing a rule.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Stop potential training use: target the service’s documented training-related crawler, if it offers one.
  • Stop a search service from discovering or answering from your pages: consider its search crawler, understanding that blocking it may reduce visibility in that service.
  • Remove a page from Google results: use an indexing control such as noindex, not a robots.txt block that prevents Google from reading the directive.
  • Limit Google previews: use Google’s snippet controls while allowing Googlebot to crawl the page.

How the main crawler choices affect visibility

Choice What it controls Search-visibility consideration
Block Googlebot Google’s crawling of the site or matched paths Google says this affects Google Search, Discover, Search features, Google Images, Google Video, and Google News. Avoid this if your aim is to preserve Google visibility. Google’s robots.txt documentation also notes that a robots.txt block does not necessarily keep a URL out of results.
Block Google-Extended Specified Gemini training and grounding uses Google says Google-Extended is a separate token from Googlebot and does not affect Google Search or Search ranking. Google-Extended documentation
Block OAI-SearchBot OpenAI’s crawler for ChatGPT search OpenAI says blocking it may prevent content from appearing in ChatGPT search answers. Other discovery may still produce a navigational link. OpenAI crawler documentation and ChatGPT search help
Block GPTBot OpenAI’s crawler associated with potential model training use OpenAI documents this separately from OAI-SearchBot, so a publisher can block GPTBot without using that rule to block ChatGPT search crawling. OpenAI crawler documentation
Apply noindex Google’s indexing and search-result inclusion for a page Google must be able to crawl the page to read the directive. Blocking the same page in robots.txt can prevent Google from seeing it. Google’s indexing documentation
Apply snippet limits How much page content Google may show in Search and AI features Google documents nosnippet, data-nosnippet, and max-snippet. Changes take time to be recrawled and processed. Google’s snippet documentation

Google: keep Googlebot separate from Google-Extended

If maintaining Google Search visibility matters, do not use a broad block against Googlebot as a proxy for restricting AI-related use. Google says: “Blocking Googlebot affects Google Search (including Discover and all Google Search features), as well as other products such as Google Images, Google Video, and Google News.” See Google’s robots.txt guidance.

Google-Extended is a separate robots.txt product token. Google says publishers can use it to control whether content is used in specified Gemini training and grounding contexts, and that this does not affect Google Search or Search ranking. See Google’s Google-Extended documentation. This is the more targeted choice when the goal is to limit those uses without blocking Google Search crawling.

OpenAI: distinguish search crawling from training-related crawling

OpenAI identifies OAI-SearchBot as supporting ChatGPT search and GPTBot as indicating that content should not be used for potential model training. Its documentation treats the controls independently. Blocking OAI-SearchBot can reduce the chance that a publisher’s content appears in ChatGPT search answers; blocking GPTBot alone need not block that search crawler. Check OpenAI’s current crawler documentation before deployment because crawler names and behavior can change.

Rank #2
FORTINET | FG-100E | FortiGate-100E Network Security Appliance
  • Protects against known exploits, malware and malicious websites; detects unknown attacks; identify thousands of applications

Robots.txt, noindex, and snippet limits do different jobs

Use robots.txt to manage crawling

A robots.txt rule asks compliant crawlers not to fetch matched URLs. It does not guarantee that a URL will disappear from search results: Google may still know the URL from other sources, and a blocked page cannot communicate a page-level directive to Google through its content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use noindex to request exclusion from Google results

For a page that should not appear in Google Search, Google must be able to crawl it and read the noindex directive. Do not disallow that URL in robots.txt at the same time if you need Google to process the directive. See Google’s documentation on blocking indexing.

Rank #3
Fortinet Web Application Firewall - Virtual Appliance for All Supported Platforms. Supports up to 1 x vCPU core FWB-VM01
  • Fortinet Web Application Firewall - virtual appliance for all supported platforms. Supports up to 1 x vCPU core
  • Fortinet HW FWB-VM01
  • Manufacturer Part: FWB-VM01

Use snippet controls to limit previews

If the page should remain crawlable and eligible for Search but show less or no text in previews, Google documents nosnippet, data-nosnippet, and max-snippet. These controls apply to Google Search and AI features. Google says recrawling and processing can take several days to several months. See Google’s snippet controls.

Deploy targeted rules and verify actual access

  1. Write down the desired outcome. Decide whether the aim is to limit potential training use, prevent an AI search product from crawling pages, reduce crawl load, limit previews, or remove pages from Google’s index.
  2. Check the current official crawler instructions. Confirm the exact user-agent token and what the service says the token controls. Do not assume that blocking one service’s bot blocks another.
  3. Use the narrowest applicable rule. Keep Googlebot accessible if Google visibility is important; use a separate training-related token such as Google-Extended where appropriate. For OpenAI, make an explicit choice between OAI-SearchBot and GPTBot.
  4. Review the live robots.txt and affected paths. Verify that the rules on the public site match the intended paths and do not accidentally cover Googlebot or other search crawlers.
  5. Check infrastructure rules as well. A CDN, web application firewall (WAF), bot-management system, CAPTCHA, authentication layer, or application check can block a crawler even when robots.txt permits it. OpenAI discusses these access layers in its crawler documentation.
  6. Validate with logs and publisher metrics. Review server requests and any available Search Console data. Google warns that its user-agent can be spoofed, so do not treat a matching user-agent string alone as proof of crawler identity; follow its Googlebot verification guidance.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to measure after a change

There is no universal traffic-loss figure for a news publisher that blocks a particular AI crawler. Official documentation describes crawler behavior and controls, not a measured causal impact on an individual publisher. Track your own results over time: Google Search Console impressions and clicks, news referral traffic, server request volume, and referrals from AI search products. Google says AI-feature traffic is included in overall Search traffic in Search Console; its AI features guidance explains how those features relate to Search.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.