What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Yes—often. A news publisher can block a crawler used for AI training while allowing a separate crawler used for search, but the rules are service-specific. Blocking Googlebot is different: Google says it affects Search, Discover, Google News, and other Google products. Decide whether you want to restrict training use, search discovery, page previews, or indexing before changing robots.txt.
What to decide before blocking a crawler
“AI crawler” can refer to bots with different purposes. A crawler may fetch pages to support a search service, to collect content for potential model training, or for other uses. Blocking one does not automatically block the others.
There is also a separate question: even when a search crawler can access a page, what may the search service index or display? Crawl access, search-result inclusion, and preview text are distinct controls. Choose the outcome you want before editing a rule.
- Stop potential training use: target the service’s documented training-related crawler, if it offers one.
- Stop a search service from discovering or answering from your pages: consider its search crawler, understanding that blocking it may reduce visibility in that service.
- Remove a page from Google results: use an indexing control such as
noindex, not a robots.txt block that prevents Google from reading the directive. - Limit Google previews: use Google’s snippet controls while allowing Googlebot to crawl the page.
How the main crawler choices affect visibility
| Choice | What it controls | Search-visibility consideration |
|---|---|---|
| Block Googlebot | Google’s crawling of the site or matched paths | Google says this affects Google Search, Discover, Search features, Google Images, Google Video, and Google News. Avoid this if your aim is to preserve Google visibility. Google’s robots.txt documentation also notes that a robots.txt block does not necessarily keep a URL out of results. |
| Block Google-Extended | Specified Gemini training and grounding uses | Google says Google-Extended is a separate token from Googlebot and does not affect Google Search or Search ranking. Google-Extended documentation |
| Block OAI-SearchBot | OpenAI’s crawler for ChatGPT search | OpenAI says blocking it may prevent content from appearing in ChatGPT search answers. Other discovery may still produce a navigational link. OpenAI crawler documentation and ChatGPT search help |
| Block GPTBot | OpenAI’s crawler associated with potential model training use | OpenAI documents this separately from OAI-SearchBot, so a publisher can block GPTBot without using that rule to block ChatGPT search crawling. OpenAI crawler documentation |
Apply noindex |
Google’s indexing and search-result inclusion for a page | Google must be able to crawl the page to read the directive. Blocking the same page in robots.txt can prevent Google from seeing it. Google’s indexing documentation |
| Apply snippet limits | How much page content Google may show in Search and AI features | Google documents nosnippet, data-nosnippet, and max-snippet. Changes take time to be recrawled and processed. Google’s snippet documentation |
Google: keep Googlebot separate from Google-Extended
If maintaining Google Search visibility matters, do not use a broad block against Googlebot as a proxy for restricting AI-related use. Google says: “Blocking Googlebot affects Google Search (including Discover and all Google Search features), as well as other products such as Google Images, Google Video, and Google News.” See Google’s robots.txt guidance.
#1 Best Overall
Google-Extended is a separate robots.txt product token. Google says publishers can use it to control whether content is used in specified Gemini training and grounding contexts, and that this does not affect Google Search or Search ranking. See Google’s Google-Extended documentation. This is the more targeted choice when the goal is to limit those uses without blocking Google Search crawling.
OpenAI: distinguish search crawling from training-related crawling
OpenAI identifies OAI-SearchBot as supporting ChatGPT search and GPTBot as indicating that content should not be used for potential model training. Its documentation treats the controls independently. Blocking OAI-SearchBot can reduce the chance that a publisher’s content appears in ChatGPT search answers; blocking GPTBot alone need not block that search crawler. Check OpenAI’s current crawler documentation before deployment because crawler names and behavior can change.
Rank #2
- Protects against known exploits, malware and malicious websites; detects unknown attacks; identify thousands of applications
Robots.txt, noindex, and snippet limits do different jobs
Use robots.txt to manage crawling
A robots.txt rule asks compliant crawlers not to fetch matched URLs. It does not guarantee that a URL will disappear from search results: Google may still know the URL from other sources, and a blocked page cannot communicate a page-level directive to Google through its content.
Use noindex to request exclusion from Google results
For a page that should not appear in Google Search, Google must be able to crawl it and read the noindex directive. Do not disallow that URL in robots.txt at the same time if you need Google to process the directive. See Google’s documentation on blocking indexing.
Rank #3
- Fortinet Web Application Firewall - virtual appliance for all supported platforms. Supports up to 1 x vCPU core
- Fortinet HW FWB-VM01
- Manufacturer Part: FWB-VM01
Use snippet controls to limit previews
If the page should remain crawlable and eligible for Search but show less or no text in previews, Google documents nosnippet, data-nosnippet, and max-snippet. These controls apply to Google Search and AI features. Google says recrawling and processing can take several days to several months. See Google’s snippet controls.
Deploy targeted rules and verify actual access
- Write down the desired outcome. Decide whether the aim is to limit potential training use, prevent an AI search product from crawling pages, reduce crawl load, limit previews, or remove pages from Google’s index.
- Check the current official crawler instructions. Confirm the exact user-agent token and what the service says the token controls. Do not assume that blocking one service’s bot blocks another.
- Use the narrowest applicable rule. Keep Googlebot accessible if Google visibility is important; use a separate training-related token such as Google-Extended where appropriate. For OpenAI, make an explicit choice between OAI-SearchBot and GPTBot.
- Review the live robots.txt and affected paths. Verify that the rules on the public site match the intended paths and do not accidentally cover Googlebot or other search crawlers.
- Check infrastructure rules as well. A CDN, web application firewall (WAF), bot-management system, CAPTCHA, authentication layer, or application check can block a crawler even when robots.txt permits it. OpenAI discusses these access layers in its crawler documentation.
- Validate with logs and publisher metrics. Review server requests and any available Search Console data. Google warns that its user-agent can be spoofed, so do not treat a matching user-agent string alone as proof of crawler identity; follow its Googlebot verification guidance.
What to measure after a change
There is no universal traffic-loss figure for a news publisher that blocks a particular AI crawler. Official documentation describes crawler behavior and controls, not a measured causal impact on an individual publisher. Track your own results over time: Google Search Console impressions and clicks, news referral traffic, server request volume, and referrals from AI search products. Google says AI-feature traffic is included in overall Search traffic in Search Console; its AI features guidance explains how those features relate to Search.
Quick Recap
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

