Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Googlebot’s User-Agent contains the stable Googlebot product token, but that text alone does not prove a request came from Google. To identify it in logs, match the token and record the source IP. To verify the crawler, reverse-resolve the IP to a Googlebot hostname, forward-resolve that hostname back to the same IP, or compare the IP with Google’s current crawler ranges.

What the Googlebot User-Agent string means

Googlebot is Google’s generic name for the crawlers used by Google Search. The two common types are Googlebot Smartphone and Googlebot Desktop. Their HTTP User-Agent headers identify the subtype, while both use the single Googlebot token in robots.txt.

Crawler Example User-Agent What to look for
Googlebot Smartphone Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) Googlebot/2.1 plus Mobile
Googlebot Desktop Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36 Googlebot/2.1 without the mobile marker
Less-common forms Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
Googlebot/2.1 (+http://www.google.com/bot.html)
The stable product token, even when browser details are absent

W.X.Y.Z is a placeholder for the Chromium version Googlebot uses. Google changes that portion as Chromium changes, so filters should match Googlebot (and, when needed, Mobile) rather than one hard-coded Chrome version. Google says most sites are primarily indexed with the mobile version, so a majority of ordinary Search crawler requests are expected to be Smartphone requests.

How to find Googlebot requests in server logs

1. Preserve the fields you need

For every candidate request, retain the timestamp, method, requested URL, status code, response bytes, source IP, and complete User-Agent. A shortened or normalized User-Agent can remove the evidence needed to distinguish variants or investigate spoofing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Google Pixel 11 Pro - Unlocked Smartphone, Gemini - 256 GB - Obsidian
  • Attention-grabbing design meets the latest evolution of the Google Pixel Camera on the new Google Pixel 11 Pro; Gemini Intelligence helps manage details so you can live in the moment[1]; and the phone is available in two sizes
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan: Works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers[2]
  • Stay informed without looking at your screen: When your phone is face down, Pixel HiLight gently alerts you with subtle glowing lights when your favorite contacts are calling or you’re talking with Gemini; exclusive to Google Pixel 11 Pro phones
  • Magic Capture catches the moment as you live it: With just one tap, Pixel 11 Pro captures video and photos, and automatically edits, crops, and unblurs a curated collection, ready to share – and you get the memory of how it felt to be in the moment
  • Two new cameras for more brilliant photos: A larger telephoto sensor captures 30% more light for clear, beautiful photos and videos, even in the dark[3]; Pixel’s longest zoom ever helps you capture details from impressive distances[4]

2. Search common access-log formats

In an Nginx or Apache combined log where the User-Agent is the final quoted field, a simple search is:

grep -i 'googlebot' /var/log/nginx/access.log

To search rotated and compressed logs:

zgrep -hi 'googlebot' /var/log/nginx/access.log*

Case-insensitive matching avoids missing an unusual capitalization. Treat every match as a candidate, not as authenticated Google traffic.

3. Separate Smartphone and Desktop requests

After finding the token, inspect the same User-Agent for Mobile. For a quick count in a standard combined log:

grep -i 'googlebot' access.log | grep -i 'mobile' | wc -l
grep -i 'googlebot' access.log | grep -iv 'mobile' | wc -l

This is a reporting distinction only. In robots.txt, the Googlebot record applies to both common crawler types; there is no separate Smartphone token for targeting rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Google Pixel 10a - 30+ Hours Battery, Camera Coach, Gemini - Obsidian 128GB
  • Google Pixel 10a is a durable, everyday phone with more[1]; snap brilliant photography on a simple, powerful camera, get 30+ hours out of a full charge[2], and do more with helpful AI like Gemini[3]
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan; it works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Pixel 10a is sleek and durable, with a super smooth finish, scratch-resistant Corning Gorilla Glass 7i display, and IP68 water and dust protection[4]
  • The Actua display with 3,000-nit peak brightness shows up clear as day, even in direct sunlight[5]
  • Plan, create, and get more done with help from Gemini, your built-in AI assistant[3]; have it screen spam calls while you focus[6]; chat with Gemini to brainstorm your meal plan[7], or bring your ideas to life with Nano Banana[8]

4. Parse the records without trusting the header

This small Python example extracts candidate IPs and User-Agents from a conventional combined log. Adjust the regular expression if your proxy or web server uses a different format.

import re
from collections import Counter

pattern = re.compile(r'^(?P<ip>S+) .*? "(?P<method>[A-Z]+) (?P<url>[^ ]+) [^"]+" (?P<status>d{3}) S+ "[^"]*" "(?P<ua>[^"]*)"')
counts = Counter()

with open("access.log", encoding="utf-8", errors="replace") as log:
    for line in log:
        match = pattern.match(line)
        if not match or "googlebot" not in match.group("ua").lower():
            continue
        ip = match.group("ip")
        kind = "smartphone" if "mobile" in match.group("ua").lower() else "desktop-or-other"
        counts[(ip, kind)] += 1

for (ip, kind), total in counts.most_common():
    print(f"{ip}t{kind}t{total}")

Do not use this script to whitelist traffic. It only finds and summarizes candidates; identity verification is a separate DNS and IP-range check.

How to verify that a request is really from Googlebot

The User-Agent is self-reported and easy for any client to copy. Google’s verification method uses the source IP and DNS in addition to the header.

  1. Start with the source IP in your log. Do not verify an address supplied in a forwarded header unless your infrastructure explicitly trusts the proxy that set it.
  2. Run a reverse DNS lookup. Google’s examples use hostnames such as crawl-66-249-66-1.googlebot.com or geo-crawl-66-249-66-1.geo.googlebot.com.
  3. Check the hostname suffix. A result should use the relevant Google crawler domain, not merely contain the word “google.”
  4. Run a forward lookup on that hostname. It must resolve back to the original source IP. This prevents an unrelated hostname from being accepted on the strength of reverse DNS alone.
  5. For bulk checks, compare the IP with Google’s maintained crawler ranges. Google publishes separate files for common crawlers, special-case crawlers, and other fetchers. Select the list that corresponds to the traffic you are investigating and refresh it from Google rather than embedding a static list in code.

Command-line example

Google’s documented example uses host:

host 66.249.66.1
host crawl-66-249-66-1.googlebot.com

The first command should return a Googlebot hostname; the second should return the original 66.249.66.1 address. Replace the example address with the IP from your own log.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why all three signals matter

Google identifies crawlers and fetchers through the HTTP User-Agent, source IP, and reverse-DNS hostname. A request can contain a convincing Googlebot/2.1 string and still be a scraper. Conversely, Google has common crawlers, special-case crawlers, and user-triggered fetchers; not every Google-branded request should be treated as ordinary Googlebot Search traffic.

Automating verification safely

Run DNS checks asynchronously from a log-processing worker rather than inside the request path. DNS can be slow or temporarily unavailable, and blocking a live request while checking identity can increase latency.

import ipaddress
import socket

def verify_candidate(ip_text: str) -> tuple[bool, str]:
    try:
        ipaddress.ip_address(ip_text)
        names, _, addresses = socket.gethostbyaddr(ip_text)
    except (ValueError, socket.herror, socket.gaierror):
        return False, "reverse lookup failed"

    hostname = names[0].lower().rstrip(".")
    allowed_suffix = hostname.endswith(".googlebot.com") or hostname.endswith(".geo.googlebot.com")
    if not allowed_suffix:
        return False, f"unexpected reverse hostname: {hostname}"

    try:
        forward_addresses = {
            item[4][0] for item in socket.getaddrinfo(hostname, None, socket.AF_UNSPEC)
        }
    except socket.gaierror:
        return False, "forward lookup failed"

    if ip_text not in forward_addresses:
        return False, "forward lookup did not return the source IP"
    return True, hostname

print(verify_candidate("66.249.66.1"))

This demonstrates the DNS logic, not Google’s complete range-file policy. In production, add IPv6 handling, DNS timeouts, caching, logging of the raw result, and a regularly updated comparison with the appropriate Google IP-range file.

Googlebot, robots.txt, and indexing are different controls

Google’s common crawlers obey robots.txt for automatic crawling. A rule addressed to Googlebot covers both Smartphone and Desktop; you cannot use that file to target only one of them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Google Pixel 10 Pro - Unlocked Smartphone with Gemini - Obsidian - 128 GB
  • Google Pixel 10 Pro is the ultimate Pixel experience, featuring advanced AI with Gemini, unbelievable camera quality, impeccable design in two sizes, and the next-gen Google Tensor G5 chip[1]
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan[2]; it works - Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Get a head start on syncing your data before it even arrives: After you purchase your new Pixel, look for an email that explains how to transfer your photos, videos, passwords, and more in just a few quick steps[11]
  • Pixel’s pro camera system makes everything look amazing, even in low light; capture more of the scene with advanced Google AI models, and bring out incredible details with 100x Pro Res Zoom, stunning 50 MP images, and super steady videos in 8K[10]
  • Pixel 10 Pro is built with durable aluminum and Corning Gorilla Glass Victus 2 for scratch and drop resistance; the 6.3-inch Super Actua display with 3,300-nit peak brightness is easy on the eyes, even in direct sunlight[3,13,18]

Blocking a URL from crawling does not guarantee that its address will stay out of Google Search. Google may know the URL from links or other signals without fetching the page. If the objective is to keep content out of the index, use an indexing directive such as noindex where Google can access the response. If the objective is to prevent both crawlers and people from accessing content, use access control such as authentication or a password.

Fetch limits and locale behavior

Maximum response Googlebot fetches

Google Search Central stated in March 2026 that Googlebot fetches up to 2 MB for an individual URL, excluding PDFs. The stated PDF limit is 64 MB. These are download limits, not a promise that every byte is indexed; processing considers only the downloaded portion after a limit is reached.

Geo-distributed crawling

For locale-adaptive pages, Google says Googlebot uses the same User-Agent across crawling configurations, including geo-distributed crawling. Its source IP may be outside the United States. Do not serve critical language or product variants only by guessing a visitor’s country. Use separate locale URLs and hreflang annotations so each version can be discovered and selected explicitly.

Troubleshooting Googlebot identification

Symptom Likely cause Fix
The log contains Googlebot, but DNS points elsewhere User-Agent spoofing or an untrusted log field Reject the candidate; verify the actual source IP and check the maintained Google ranges.
Reverse DNS returns no name Temporary DNS failure, a non-Google client, or an address not configured for Googlebot Record the failure, retry asynchronously, and do not whitelist the request from its header.
Reverse DNS looks valid, but forward DNS returns a different IP The hostname does not prove control of the source address Fail verification and investigate the original connection and proxy chain.
A filter misses newer Googlebot requests It hard-codes one Chrome version or an entire sample string Match the stable Googlebot token and test the optional Mobile marker separately.
Blocking Googlebot did not remove a URL from Search Crawling and indexing controls were treated as the same thing Use an appropriate indexing directive or access control for the desired outcome.
Different countries receive different content from the same agent Locale adaptation based only on IP geolocation Publish distinct locale URLs and annotate them with hreflang.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you are documenting how a page appears during a crawler investigation, you can capture it through ScreenshotNeo instead of installing a browser and writing capture code. This does not authenticate Googlebot; DNS and IP verification above remain the identity test. ScreenshotNeo accepts a URL and returns a PNG, JPEG, WebP, or PDF, while removing cookie-consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

The ScreenshotNeo documentation covers the same API plus full-page and element captures, device and viewport settings, custom headers and cookies, waits, blocking rules, PDFs, caching, asynchronous jobs, bulk capture, and its MCP server for AI clients such as Claude and Cursor. There is a free allowance of 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.

Best Value
Google Pixel 7-5G Android Phone - Unlocked Smartphone with Wide Angle Lens and 24-Hour Battery - 256GB - Lemongrass
  • Google Pixel 7 is powered by Google Tensor G2; it’s faster, more efficient, and more secure, with the best photo and video quality yet on Pixel[1].Other camera description:Front,Rear.Bluetooth Version 5.2 with dual antennas for enhanced quality and connection.
  • Unlocked Android 5G phone gives you the flexibility to change carriers and choose your own data plan[2]; works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Pixel’s Adaptive Battery can last over 24 hours; when Extreme Battery Saver is turned on, it can last up to 72 hours[3]
  • The 6.3-inch Pixel 7 display is super sharp, with rich, vivid colors; it’s fast and responsive for smoother gaming, scrolling, and moving between apps[4]
  • Google Pixel 7 has wide and ultrawide lenses with up to 8x Super Res Zoom[5]; and Cinematic Blur brings more drama to your videos

FAQ

Should DNS verification run on every request?

No. Store the raw request evidence and process candidates in a worker with bounded retries and short-lived caching. This keeps DNS failures from slowing page delivery while preserving an auditable result.

What should I save when investigating a suspected spoof?

Keep the timestamp, source socket IP, complete User-Agent, requested path, response status, proxy metadata you trust, and the reverse/forward DNS results. Those fields let you recheck the decision after Google changes crawler ranges or DNS records.

Frequently Asked Questions

Should DNS verification run on every request?

No. Process candidates asynchronously with bounded retries and short-lived caching so DNS issues cannot slow normal responses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I save when investigating a suspected spoof?

Retain the timestamp, source IP, complete User-Agent, requested path, status, trusted proxy metadata, and both DNS lookup results.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.