Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—but Google itself is not one crawler. Google Search is an automated search service that uses software called Googlebot to fetch pages. Crawling is only the discovery and fetching stage; Google must separately process a page for indexing and then decide whether to show it in search results.

What “Google” means in this question

The word Google can refer to the company, Google Search, or the automated clients that request web pages. Those are not interchangeable.

  • Google: the company and its products.
  • Google Search: the search engine that discovers, processes and serves information.
  • Googlebot: Google’s documented crawler software for ordinary Search crawling.

Google describes Search as a fully automated search engine that uses web crawlers to explore the web and find pages for its index. Therefore, it is accurate to say that Google Search uses a web crawler, but imprecise to call the entire company or service “a crawler.”

How Google Search works: three separate stages

A request from Googlebot does not mean that a page will rank, appear in Search, or even enter Google’s index. Google explains Search as three related stages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. Crawling

Googlebot requests resources it has discovered, including HTML, images and video. It finds URLs mainly through links on pages it already knows about. A sitemap submitted by a site owner is another discovery path. Google chooses algorithmically which sites to visit, how often to return and how many URLs to request.

2. Indexing

After fetching a page, Google analyzes its text, media, metadata and relationships to other pages. It may render the page and execute JavaScript with a recent version of Chrome. Google can then decide whether to store information about the page in its index. A successful fetch is not an indexing guarantee.

3. Serving results

When someone searches, Google matches the query against information in its index and selects results to display. A crawled and indexed page can still be omitted for a particular query. Google says it does not guarantee crawling, indexing or serving, even when a site follows its Search guidance.

Stage What happens What it does not guarantee
Crawling Googlebot requests a discovered URL and its resources. Indexing or visibility in results.
Indexing Google analyzes and may store information about the fetched content. A ranking or a result for every query.
Serving Google selects indexed content for a user’s search. That every indexed page will be shown.

Which Googlebot visits a site?

Googlebot Smartphone

Google documents a smartphone crawler and says most Search crawling for most sites uses the smartphone version. This reflects Google’s mobile-first indexing approach: the mobile version is generally the version it primarily evaluates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Googlebot Desktop

The desktop crawler simulates a desktop user. Both smartphone and desktop Googlebot variants use the same Googlebot product token in robots.txt. You therefore cannot use that token to permit one variant while blocking the other.

Other Google clients

Google also documents separate categories for common crawlers, special-case crawlers and fetchers used by particular products or actions. Their behavior and robots.txt arrangements can differ from ordinary Search crawling.

Google-Extended is not a crawler

Google-Extended is a standalone robots.txt product token, not an HTTP user-agent string identifying a separate bot. Publishers can use it to control whether content Google crawls may be used for training future Gemini models or grounding in certain Gemini products. Google says this token does not affect inclusion in Search and is not a Search ranking signal.

How Google discovers URLs

Links

Links from pages Google already knows about are the principal discovery mechanism. A page with no discoverable link can take longer to find, even if it is publicly accessible.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sitemaps

A sitemap gives Google a list of URLs that you consider important. It is a discovery signal, not a command. Submitting one does not guarantee that Googlebot will request every URL or that those URLs will be indexed.

Scheduling and server conditions

Google determines crawl frequency and volume algorithmically. Googlebot tries not to crawl too quickly and can slow down when a server returns conditions such as HTTP 500 errors. Temporary downtime, slow responses and resource limits can therefore affect the rate of fetching.

robots.txt, noindex and access control are different

These controls solve different problems. Choosing the wrong one can leave a URL visible when you expected it to disappear.

robots.txt controls requests

A file at the site’s root can tell supported crawlers which paths they may request. Google’s documented fields include User-agent, Allow, Disallow and Sitemap. Google does not support crawl-delay.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example:

User-agent: Googlebot
Disallow: /private/
Sitemap: https://example.com/sitemap.xml

This asks Googlebot not to fetch URLs under /private/. It is not an instruction to remove those URLs from Search. Google may learn a blocked URL from links and show the URL without a snippet.

noindex controls index inclusion

A noindex meta directive or HTTP response header tells Google not to include a page in its index. Googlebot must be able to fetch the page and read that directive. Blocking the same URL in robots.txt can prevent Google from seeing the noindex rule.

HTML example:

<meta name="robots" content="noindex">

HTTP-header example:

X-Robots-Tag: noindex

Authentication controls people and bots

Password protection or another access-control system prevents the public, including crawlers, from retrieving the content. Use it when material must not be accessible to visitors as well as excluded from indexing. Neither robots.txt nor noindex is an access-control mechanism.

Goal Appropriate control Important limitation
Reduce crawler requests to a path robots.txt The URL can remain known and appear without a snippet.
Keep a fetchable page out of the index noindex meta tag or HTTP header Google must be allowed to fetch and read it.
Keep content inaccessible to the public Password protection or equivalent access control Visitors and crawlers cannot retrieve the protected response.

How to tell whether a request really came from Googlebot

An HTTP user-agent string is only a claim and can be spoofed by any client. Do not treat a request as genuine solely because its header says Googlebot.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reverse DNS verification

Perform a reverse DNS lookup on the source IP and confirm that the resulting hostname belongs to a Google domain. Then perform a forward lookup on that hostname and check that it resolves back to the original IP. This two-way check helps reject a forged hostname.

Published IP-range verification

Alternatively, compare the source address with Google’s published crawler IP ranges. Keep the ranges current; an old allowlist can produce false negatives.

Apply verification before granting special treatment such as bypassing a rate limit, allowing an otherwise restricted route or labeling analytics traffic as Googlebot.

Practical checks for site owners

When a new page is not found

  1. Confirm that the URL returns a successful response to ordinary visitors.
  2. Check that a link or sitemap exposes the URL.
  3. Inspect robots.txt for a matching Disallow rule.
  4. Look for accidental noindex meta or X-Robots-Tag headers.
  5. Ensure important content is present in the rendered page, not only after an interaction Google cannot perform.
  6. Allow time for Google’s scheduling and processing; discovery, crawling and indexing are separate decisions.

When traffic claiming to be Googlebot is excessive

  1. Record the complete request, including source IP, timestamp, path and user-agent.
  2. Verify the IP with reverse and forward DNS or Google’s current published ranges.
  3. Return appropriate server errors rather than silently serving corrupted content during overload.
  4. Use rate controls that protect availability without assuming every user-agent label is authentic.

When a page must disappear from Search

Do not rely on robots.txt alone. Make the page fetchable so Google can read a noindex directive, or require authentication if the content must also be inaccessible. Existing results can take time to change after the rule is corrected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

See the rendered page without building a crawler

Googlebot can render JavaScript, but a visual check is useful when diagnosing consent banners, overlays, responsive layouts or lazy-loaded content. ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF output; its pre-capture steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets.

Or skip the browser setup

Use the API instead of maintaining a browser:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for parameters. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.

Common misconceptions

  • “Google is one bot.” Google operates multiple crawlers and fetchers; Googlebot is the ordinary Search crawler name.
  • “A crawl means a ranking.” Fetching is only the first stage.
  • “A sitemap forces indexing.” It supplies URLs for discovery but does not compel crawling or inclusion.
  • “robots.txt removes a page.” It controls fetching and can leave a known URL eligible to appear.
  • “The user-agent proves identity.” Headers are spoofable; verify the source IP.
  • “Googlebot only downloads HTML.” Google fetches resources such as images and video and can render JavaScript.

Frequently Asked Questions

Is Googlebot the same as Google Search?

No. Google Search is the service; Googlebot is software that fetches pages for its crawling stage.

Can I block only Googlebot Smartphone?

Not with the shared Googlebot robots.txt product token. Smartphone and desktop Googlebot use that same token.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does robots.txt prevent a URL from appearing in Google?

No. A blocked URL can remain known through links and may appear without a snippet. Use a readable noindex directive when index exclusion is the goal.

How can I verify a Googlebot IP?

Use reverse DNS followed by a forward lookup, or compare the address with Google’s published crawler IP ranges. Never rely on the user-agent string alone.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.