Google “scrapes” the web through automated crawling: Googlebot discovers URLs, requests pages and resources, and may render pages before Google evaluates what—if anything—to include in its Search index. These are separate stages. A page can be crawled without being indexed, and being indexed does not guarantee a ranking or a particular search result.
Understanding those distinctions makes it easier to diagnose visibility problems and choose the right control. A sitemap can help Google find a URL; robots.txt can restrict crawling; and a noindex directive can request exclusion from the index—but only if Google can fetch the page and see it.
What Google means by crawling and indexing
In ordinary conversation, “scraping” can mean collecting information from websites. Google describes its Search process in terms of crawling and indexing. Googlebot is Google’s crawler: software that requests publicly accessible URLs. Google’s larger process also includes discovery, rendering, index processing, and deciding what to show in Search. Each stage has a different outcome, so a successful request is not proof that a URL has made it into the index. Google’s overview of crawling and indexing explains the distinction.
This article concerns Google Search’s web crawling, not every Google service or every crawler that uses a Googlebot-like name. Google operates different crawlers for different purposes. Rules and behavior can depend on the crawler’s product token; Google’s Search documentation describes the relevant Search crawlers and how to identify them. Google’s Googlebot documentation is the reference for those details.
#1 Best Overall
How Google gets from a URL to a search result
- Discovery: Google finds URLs through links on pages it has already crawled and through submitted sitemaps, among other sources. Discovery only puts a URL in the set Google may consider; it does not order an immediate visit.
- Crawl scheduling and fetch: Google’s systems choose whether and when to request a URL. They account for Google’s demand for the URL and the site’s ability to respond. Google says its crawlers try not to overload websites; server errors and other capacity problems can lead Googlebot to slow down.
- Rendering: After fetching a page, Google may render it and run JavaScript using a recent version of Chrome. The browser may need to fetch CSS, JavaScript, images, and other resources separately. If important resources are blocked or fail to load, Google’s rendered view may differ from what a visitor sees.
- Index processing: Google analyzes the content and signals it can access, including text and key metadata. It can assess duplicate pages and canonical versions, then decide whether a page is suitable for the index. Crawling is not a promise of inclusion.
- Search presentation: Being indexed is not a promise of ranking, a particular position, or a particular appearance. Google’s troubleshooting guidance notes that a page may not appear even after it has been crawled if Google considers its value or user demand insufficient.
For most sites, Google Search primarily indexes the mobile version. Googlebot Smartphone and Googlebot Desktop share the same robots.txt product token, so robots.txt cannot be used to allow one of those subtypes while disallowing the other. Google’s crawler documentation covers this distinction.
How Google discovers URLs—and what a sitemap can do
Links help Google move from known pages to other pages. Make important pages reachable through standard crawlable links rather than relying only on a visitor action Google may not perform. An accurate sitemap is another way to submit a URL list and associated metadata, and can be especially useful for a large or complex site. Neither links nor sitemap entries compel Google to crawl or index a URL.
Google calls sitemap submission a hint: it does not guarantee Google will download the sitemap or use it to crawl the listed URLs. Keep the sitemap current and its last-modified information honest. A sitemap file can contain up to 50 MB uncompressed or 50,000 URLs, according to Google’s sitemap guidance; larger inventories can be divided among multiple sitemap files and referenced from a sitemap index. Google’s sitemap documentation has the submission and format details.
A sitemap is not a substitute for navigation, internal links, or a working site. If a URL is important, link to it from relevant pages as well as including it in a sitemap. If the sitemap reports URLs that redirect, are blocked, return errors, or are duplicates of preferred URLs, investigate the underlying URL structure rather than assuming submission will resolve the issue.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
How crawl scheduling and crawl budget work
Google does not give every site a single fixed crawl quota that should be maximized. Its crawl-budget guidance frames crawling around two factors: crawl capacity (how much crawling a site can handle) and crawl demand (how much Google wants to crawl particular URLs). Large or frequently changing sites may have reason to examine crawl budget more closely; for most sites, Google says that keeping a sitemap current and checking the Page Indexing report is adequate. The guidance was updated July 22, 2026. Google’s crawl-budget guide describes when additional work is relevant.
Repeatedly adding and removing robots.txt rules is not a general way to shift crawling onto preferred pages. Google says that blocking some URLs does not cause a general reallocation of crawl budget to other URLs. For a large site, begin with an accurate URL inventory, server health, and the reasons for generating so many URLs. Consolidate duplicate URL variants where appropriate and look for crawl traps, such as faceted navigation that can produce an unbounded number of combinations.
Google also publishes file-size limits for Googlebot: for supported file types it fetches the first 2 MB, and for PDF files the first 64 MB. Google applies these limits to uncompressed data; resources referenced by a page are fetched separately. These limits matter when investigating unusually large documents, but do not explain every failure to index a page. See Google’s current Googlebot guidance for the qualifications.
Robots.txt, noindex, and private content are different controls
| Control | What it affects | What Google must be able to do | Use it when |
|---|---|---|---|
| robots.txt | Whether a crawler is allowed to request a URL or resource | Google can read the applicable robots.txt rules, but a disallowed page cannot reliably deliver its own page-level directives | You want to restrict crawling, for example of selected resources or URL areas |
| noindex meta tag or HTTP header | Whether a crawlable page should be excluded from Search indexing | Google must fetch the page and see the directive | The page can be fetched, but should not be in Search |
| Authentication or password protection | Access to the content itself | A user or crawler must have valid credentials to access protected content | The content should not be publicly accessible |
These controls are not interchangeable. robots.txt is a crawl control, not a reliable removal instruction. Google warns that a known URL can still appear in results even when Googlebot is blocked from crawling it—for example, if other pages link to it. Blocking the page can also prevent Google from seeing a noindex directive placed on that page. If the goal is to keep publicly accessible content out of the index, allow Google to crawl it and serve a noindex instruction. If the goal is to make the content private, use access protection instead. Google’s noindex guidance and developer SEO guide explain these distinctions.
Rank #3
Google supports noindex through a robots meta tag in a page’s HTML or through an HTTP response header. Choose the method that fits the content and delivery stack, and verify the directive on the response Google can fetch. A rule in robots.txt does not itself tell Google to remove a URL from its index. Google’s robots meta tag and X-Robots-Tag reference documents the syntax and scope.
Why a page can be crawled but not indexed
“Crawled – currently not indexed” and similar status descriptions can be frustrating, but they do not mean Google has promised to index a page later. Crawling confirms a fetch; index selection is a separate decision. Google’s troubleshooting guidance says a page may be omitted if Google considers its value or user demand insufficient. Other problems can arise earlier or during processing:
- Google has not fetched the URL: The URL may be new or difficult to discover, the sitemap may not have been processed, or crawling may not yet be scheduled.
- The server or network failed: Timeouts, server errors, or unstable availability can interrupt fetching or discourage frequent requests.
- The page cannot be rendered as intended: Required scripts, stylesheets, or other resources may be blocked, unavailable, or dependent on a user interaction.
- A directive prevents the desired outcome: The page may have a noindex directive, or robots.txt may block Google from fetching it and seeing the page.
- Google selected another version: Duplicate or near-duplicate URLs can be grouped, and Google may choose a canonical version other than the URL you expected.
- The fetched page is not considered suitable for Search: Fetching alone does not establish that the page merits inclusion. Review its usefulness and distinct value rather than treating repeated submission as a guarantee.
Google advises that updates are checked and indexed in a reasonably timely manner, but for most sites this means three days or more; this is guidance, not a guaranteed indexing deadline. Same-day indexing should not be expected for ordinary pages. Google’s crawling troubleshooting guidance discusses the possible states and checks.
How to check whether Google can access a page
- Inspect one URL in Search Console. Use URL Inspection to check the URL’s reported indexing state and available details about Google’s last crawl. Compare the tested URL with the canonical URL you intend Google to use.
- Check the Page Indexing report for patterns. A site-level pattern can be more informative than repeatedly inspecting isolated URLs. Group affected pages by template, URL pattern, response behavior, and reported status.
- Review robots.txt and page directives. Confirm that the URL and its required resources are not disallowed unintentionally. For a noindex plan, make sure Google can fetch the page and that the response contains the intended meta tag or header.
- Check the sitemap and links. Confirm the sitemap lists the intended canonical URLs, is reachable, and has truthful modification dates. Check that important pages are linked from other crawlable pages.
- Check server logs and availability. Look for response codes, timeouts, repeated failures, and resource requests around Googlebot activity. Check server or network problems before changing crawl directives.
- Verify that a supposed Googlebot request is genuine. A user-agent string can be spoofed. Google recommends reverse-DNS verification or comparison with its published crawler IP ranges; do not treat a matching user-agent alone as proof. See Google’s crawling overview for verification guidance.
For a broader diagnosis, combine the URL Inspection result with Page Indexing and Crawl Stats reports, the actual HTTP response, and server logs. No single signal proves that every resource loaded correctly or that a URL will be selected for Search. Google’s troubleshooting guide provides additional checks for crawl errors.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Or skip the browser setup
If your immediate task is to capture a rendered page image or PDF for a review workflow, ScreenshotNeo is a screenshot API, not a Googlebot tester: a screenshot does not establish whether Google can crawl or index that URL. Its API can return a screenshot or PDF from one request. This cURL example saves a WebP image of the example page; replace the target URL with the page you want to capture. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python request:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js request:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server offers screenshot tools for AI agents, including Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
What to do first when a page is missing
Use the evidence to identify the stage that is failing. If Google has not discovered the URL, improve crawlable linking and keep the sitemap accurate. If it cannot fetch the page reliably, resolve server, network, or robots access issues. If it fetches but cannot see the intended content or directive, check rendering and resource access. If it has crawled the URL but not indexed it, review directives, canonical selection, duplication, and whether the page offers distinct value. Repeatedly submitting the same URL cannot guarantee a different result.
Frequently Asked Questions
Does Google crawl every page on the internet?
No. Google’s systems choose URLs based on factors including discovery, crawl demand, and site capacity. A URL being public or listed in a sitemap does not guarantee a fetch.
Recommended Free Tools
Can I make Google index a page immediately?
There is no guarantee of immediate indexing. Search Console inspection and a sitemap can help you check or signal a URL, but Google makes the indexing decision.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

