Free tools Windows power users keep installed
One-click scans. No signup required.
A search engine crawler is software that automatically discovers URLs and requests web pages and related resources so a search engine can process their content. Crawling is an early step—not a guarantee that a page will be indexed, appear in search results, or rank for a query. For Google Search, the crawler is called Googlebot.
What is a search engine crawler?
A search engine crawler is an automated program that visits web addresses and fetches their content. It may also be called a bot, robot, or spider. Google Search Central describes its own fetching program this way: “The program that does the fetching is called Googlebot (also known as a crawler, robot, bot, or spider).” Google Search Central: How Search Works
A crawler is software, not a person or usually a single physical machine. Search engines use automated systems to find pages across the web; the details and names of those systems vary by search engine. Googlebot is Google’s crawler, not a universal name for every search engine’s crawler.
How do search engines find and crawl pages?
- A URL becomes known. A search engine may already know an address, discover it by following a link from another known page, or learn about it from a sitemap.
- A crawler may request it. The search engine decides whether and when to fetch the URL. Google says its crawler chooses which sites to visit, how often, and how many pages to request, while attempting to avoid overloading sites. Server responses, including HTTP 500 errors, can affect its crawl activity.
- The search engine processes the fetched content. For Google Search, this may include rendering a page and running JavaScript with a recent version of Chrome. Other search engines may work differently.
There is no central register of every page on the web, and a sitemap or link can help a crawler discover a URL without guaranteeing that it will be fetched. Google also says it does not guarantee that it will crawl, index, or serve a page, even if the page follows its guidance. Google Search Central: How Search Works
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
Crawling, indexing, and search results are different
| Stage | What it means |
|---|---|
| Crawling | A crawler discovers a URL and requests its content. |
| Indexing | The search engine analyzes content and signals, then may store information about the page in its index. |
| Serving results | The search engine selects information it considers relevant to a user’s query. |
A page can fail to advance at any of these stages. Crawling makes content available for processing; it does not put the page in the index by itself, and indexing does not guarantee that the page will be shown for a particular search. Google Search Central: How Search Works
What is Googlebot?
Googlebot is Google’s crawler for Google Search. Google describes two general types: Googlebot Smartphone and Googlebot Desktop, which simulate mobile and desktop users. Both use the same Googlebot product token in robots.txt, so site owners cannot target them separately with that file. Google says mobile crawling makes up most Googlebot requests for most sites. These are Google-specific details, not rules for every search engine. Google Search Central: Googlebot
Rank #2
Google’s March 31, 2026 post describes Googlebot as one client of shared crawling infrastructure. It states that Googlebot fetches up to 2 MB from an individual URL, excluding PDFs, and up to 64 MB for a PDF; the stated limit includes the HTTP header. These limits describe Google’s implementation, not a general crawler limit. Google Search Central: Googlebot and shared crawling infrastructure
Robots.txt, noindex, and private pages
These controls solve different problems. Use the one that matches your goal:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #3
| Goal | Approach | Important distinction |
|---|---|---|
| Ask crawlers not to fetch specified paths | Set rules in robots.txt. | Rules apply to the host, protocol, and port where that robots.txt file is hosted. Blocking a URL does not reliably keep it out of search results; the URL may still appear without a description based on fetched content. |
| Tell Google not to index a page | Allow crawling and provide a noindex directive. | If crawling is blocked, Google may not be able to fetch the page and see the noindex instruction. |
| Keep content private | Restrict access, for example by requiring authentication. | Robots.txt is not an access-control or security boundary. |
Google documents these distinctions in its guidance on robots.txt and blocking search indexing. Google says Bing and other major search engines also support the sitemap field in robots.txt, but a sitemap entry is a discovery aid, not an indexing guarantee.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to verify whether a request is really from Googlebot
Do not rely on a request’s user-agent string alone: it can be spoofed. Google recommends verifying a claimed Google crawler by performing a reverse-DNS lookup and checking the result, or by comparing the source IP address with Google’s published crawler IP ranges. Google Search Central: Verify Googlebot and other Google crawlers
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

