Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web crawling fails at scale when two finite systems collide: a crawler has limited bandwidth, time, and worker capacity, while every target host has finite serving capacity and an uneven supply of URLs. The practical fix is to identify which stage is failing—discovery, fetching, rendering, or indexing—then remove low-value URL work, keep important URLs easy to discover, protect origin capacity, and measure recovery with logs and crawler telemetry.

Do not treat “not indexed” as proof that a page was never crawled. Google separates crawling from indexing: a URL can be fetched and still be excluded because it lacks sufficient value or user demand.

What scale changes in a crawl

A small crawl can hide inefficient URL design and slow responses because there is spare capacity. At larger volumes, the same defects multiply into queue growth, host overload, missed recrawls, and stale data.

Constraint What becomes scarce Typical symptom
Crawler-side capacity Bandwidth, elapsed time, worker instances, queue slots Important URLs remain undiscovered or wait too long
Host-side capacity Origin CPU, database connections, CDN throughput, rendering resources Latency rises, 429/5xx responses increase, and crawl rate falls
Uneven crawl demand Attention for URLs that Google considers useful or fresh High-value pages are recrawled while duplicates consume requests

Crawl budget has two parts

Google defines crawl budget as “the number of URLs Googlebot can and wants to crawl.” Crawl rate is constrained by response health and server capacity. Crawl demand reflects how much Google wants to fetch particular URLs for its index. Adding servers can remove a rate bottleneck, but it cannot create demand for low-value pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Why host boundaries matter

Crawl demand is not distributed evenly across a site. A busy product host, an image host, an API subdomain, and a rarely changed archive may each need different limits and monitoring. Analyze capacity per hostname and URL pattern instead of averaging all requests into one site-wide number.

Find the failing stage before changing crawl settings

Stage Question Evidence to collect
Discovery Can a crawler find the URL through links or sitemaps? Internal-link graphs, sitemap contents, referring URLs, crawl queues
Fetching Can the host return the document reliably? Access logs, status codes, time to first byte, timeouts, CDN and origin health
Rendering Can required scripts and resources complete in a reasonable time? Resource waterfalls, JavaScript errors, blocked requests, render timing
Indexing Did the search engine decide the page belongs in results? URL Inspection, canonical signals, content quality, duplication, demand

Use Google Search Console Crawl Stats and URL Inspection, then verify requests in server logs. A report that says “not indexed” needs this separation; improving crawl access alone does not guarantee indexing.

The failure modes that consume crawl capacity

Unbounded URL spaces

Faceted navigation, date calendars, internal search parameters, sort orders, tracking parameters, proxy URLs, and infinite combinations can produce millions of addresses with little unique content. Shopping carts, session actions, and other state-changing URLs are not content inventories and should not be exposed as discovery paths.

  • Publish stable, canonical URLs for indexable content.
  • Link to important pages with ordinary crawlable links.
  • Keep parameters that do not change content out of internal links where possible.
  • Use durable crawl restrictions for known traps rather than repeatedly toggling rules.

Origin overload and poor availability

When a host is slow, unavailable, or returns many errors, Googlebot reduces crawling. Compare Crawl Stats host-availability graphs with origin and CDN logs, deployment events, database saturation, and the exact URL patterns that failed. If the serving limit is repeatedly reached while important URLs remain unprocessed, add capacity or reduce expensive work and observe whether successful requests recover.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Slow documents, heavy resources, and redirect chains

Google limits crawling by bandwidth, time, and available crawler instances. Faster responses can permit more fetching, but speed does not make a page valuable. Prioritize important templates: reduce server and render time, remove unnecessary redirect hops, bound resource sizes, and reuse stable resource URLs so caches can work. Conditional requests such as If-Modified-Since and If-None-Match can avoid regenerating unchanged content when the crawler supports them; they are not sent on every request.

Misunderstood robots and indexing controls

robots.txt controls whether a crawler may request a URL. It is not authentication, secrecy, or an indexing guarantee. Use a noindex directive when the crawler must fetch a page but you do not want it indexed, and use authentication for private material. The IETF Robots Exclusion Protocol (RFC 9309, September 2022) explicitly says robots rules are not authorization. Google also advises using robots.txt for stable controls, not as a temporary dial for reallocating crawl budget.

A diagnostic workflow that scales

  1. Define the missing set. List the URLs that matter—revenue pages, documentation, changed pages, or other critical templates. Classify each as undiscovered, not fetched, failed during fetch, failed during render, or crawled but not indexed.
  2. Measure crawler activity. Review Search Console Crawl Stats and URL Inspection. Confirm actual Googlebot requests in logs; user-agent strings can be spoofed, so validate Googlebot with reverse DNS and the published Google IP ranges.
  3. Group failures. Break data down by status code, latency band, response size, host, path pattern, query parameters, robots rule, and time. Clusters often reveal a facet explosion, redirect loop, slow template, or deployment-related outage.
  4. Check capacity at the serving layer. Correlate request volume with CPU, memory, database pools, queue depth, CDN cache status, and upstream timeouts. A high request count alone is not proof of a crawl problem; show that important URLs are losing successful service.
  5. Fix overload errors first. Resolve widespread 5xx and 429 responses, then watch whether successful Googlebot requests and crawl activity rise gradually. Do not leave emergency overload responses enabled indefinitely.
  6. Reduce low-value work. Bound parameter combinations, remove state-changing links, collapse duplicates, shorten redirect chains, and prevent noncritical resources from being required for basic page understanding.
  7. Improve discovery and freshness. Maintain a sitemap containing important and recently changed URLs, use accurate lastmod values, and ensure valuable pages have crawlable internal links. A sitemap is a hint, not a command or an immediate-crawl guarantee.
  8. Review the trend. Compare important-URL coverage, successful request rate, error rate, latency, bytes transferred, and render failures before and after every change.

Design an efficient URL inventory

Prefer finite, canonical paths

Every indexable concept should have one durable preferred URL. Generate links from that inventory rather than exposing every combination of filters, calendars, or sort orders. Canonical tags help consolidate signals but do not make an unlimited URL space cheap to crawl; prevent needless discovery at the source.

Use links and sitemaps for different jobs

Internal links provide context and a path for discovery. Sitemaps provide a curated list of important or recently modified URLs. Keep both accurate. Remove obsolete entries, repair broken links, and do not inflate lastmod on every deployment when content did not materially change.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Control traps permanently

Use robots.txt for stable, well-understood crawl restrictions. Do not rely on obscure query names, intermittent 403 responses, or changing directory rules as security or budget controls. If a URL must remain private, require authentication; if it may be fetched but should not enter search results, use an indexing directive.

Protect the host without hiding useful content

Interpret overload status codes correctly

Response Google-specific effect Operational response
429 Too Many Requests Signals overload and can slow crawling Apply briefly during an incident, reduce pressure, and monitor recovery
503 Service Unavailable Signals temporary server trouble and can slow crawling Use for short emergency protection, then stop when capacity recovers
5xx responses Indicate server errors; persistent failures can lead to URLs being dropped Fix the underlying dependency, timeout, or capacity problem
401 or 403 Do not provide the same crawl-rate reduction behavior as 429 Use only when access really requires authentication or authorization

Google’s emergency guidance is temporary: stop returning 429 or 503 when crawl pressure falls, and do not keep them active for more than a day or two. Its documentation warns that errors lasting several days can cause URLs to be removed from Search. This is Google-specific operational guidance, not a universal retry policy for every crawler.

Set politeness in your own crawler

For a general-purpose crawler, enforce per-host concurrency, request timeouts, retry budgets, and exponential backoff. Pause or lengthen backoff after 429 responses, and avoid synchronized retries that create a second traffic spike. There is no single delay or concurrency value that is safe for every website; derive limits from observed latency, error rate, and the host’s published policy.

Improve useful pages per unit of work

Make the first response count

Return essential content quickly, keep redirects short, and avoid making a crawler execute an unnecessary chain of scripts before it can identify the page. Separate optional analytics, chat, and advertising resources from the critical rendering path. Stable asset URLs improve cache reuse.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Reuse unchanged responses

Honor conditional requests where appropriate and set cache policies that match content volatility. A crawler that receives a validated unchanged response spends less bandwidth and origin CPU than one that regenerates the entire document every time.

Prioritize by value and freshness

Queue newly published or materially changed important URLs ahead of low-value duplicates. Keep a bounded retry queue so a failing host cannot monopolize workers. Record the reason each URL entered the queue, the last successful fetch, and the next permitted retry time.

Monitor the system with actionable metrics

  • Coverage: percentage of priority URLs discovered, fetched successfully, rendered, and inspected for indexing.
  • Efficiency: successful useful pages per request, bytes, worker-second, and render-second.
  • Host health: latency percentiles, timeout rate, 429/5xx rate, origin saturation, and CDN hit ratio by hostname.
  • Inventory quality: duplicate URL ratio, parameter combinations, redirect hops, orphan pages, and sitemap accuracy.
  • Recovery: time from an overload event to normal successful crawl activity.

Alert on changes in these measures by URL pattern and host, not only on site-wide averages. A small set of expensive paths can be responsible for most wasted work.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Visual verification for rendered pages

Logs tell you that a document returned; they do not show whether a consent banner, newsletter modal, chat widget, or client-side failure obscured the useful content. For a representative sample, open the URL in a browser at the target viewport, wait for the page’s normal data to appear, and capture the full page and key selectors. Compare captures after template, JavaScript, consent, or CDN changes. Record the URL, viewport, timestamp, response status, and build identifier so visual changes can be correlated with crawl telemetry.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For automated captures, ScreenshotNeo is a website screenshot API and MCP server. It can load lazy images for full-page shots, capture one CSS-selected element, wait for a selector, delay, or network idle, run custom JavaScript or CSS, click an element, hide selectors, block ads, trackers, requests, or resource types, and set headers, cookies, user agents, authorization, timezone, and geolocation. It also supports 12 device presets plus custom viewports, retina scale, dark mode, transparent backgrounds, image resizing, PDF paper and page options, HTML/CSS-to-image, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and parameter names used by other screenshot APIs.

Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Or skip the browser setup

One GET request returns PNG, JPEG, WebP, or PDF. ScreenshotNeo accepts the cookie or consent banner before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for parameters and response headers.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so an AI agent can inspect representative pages without a custom browser harness. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to start.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshoot common symptoms

Symptom Likely cause Fix and verification
Many URLs discovered, few valuable pages fetched Facets, calendars, duplicate parameters, or proxy URLs Constrain links and sitemap entries; compare the duplicate ratio and priority-URL coverage
Crawl activity drops during deployments Origin saturation, timeout changes, or a broken dependency Correlate deploy time with logs and Crawl Stats; roll back or restore capacity, then watch successful requests
High 429/5xx rate Serving limit, database pool exhaustion, or synchronized retries Protect the host briefly, fix the bottleneck, back off retries, and remove emergency responses after recovery
Pages are crawled but absent from results Indexing decision, canonical conflict, duplication, low value, or low demand Use URL Inspection and page-quality signals; do not assume more crawl capacity will solve it
Rendered content is missing Slow or blocked JavaScript, oversized resources, or an overlay Inspect resource waterfalls and visual captures; make essential content available within a bounded render path
Robots changes have unpredictable effects Temporary rules used as a budget switch, or an unreachable robots.txt Restore stable rules, verify fetch behavior, and use authentication for private data

What a durable fix looks like

A reliable large-scale crawl is not achieved by one larger crawler or one robots.txt edit. It is a feedback loop: maintain a finite URL inventory, expose valuable pages through links and accurate sitemaps, serve them quickly and consistently, apply host-aware politeness, and compare crawler logs with indexing diagnostics. Capacity work is justified when serving limits are demonstrably blocking successful requests; inventory and quality work is required when the crawler is spending its budget on URLs nobody needs.

Frequently Asked Questions

Does a larger server guarantee that Google will crawl more URLs?

No. Extra serving capacity can remove a crawl-rate bottleneck, but Google’s crawl demand still depends on perceived usefulness, freshness, and other indexing signals.

What does RFC 9309’s 500 KiB figure mean for robots.txt?

RFC 9309 requires implementations to support parsing at least 500 KiB. It is a protocol implementation limit, not a statistic about how often crawling fails.

Should every crawler retry a failed request immediately?

No. Immediate retries can amplify an outage. Use host-specific limits, exponential backoff, and a retry budget; pause or slow down after 429 responses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.