The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →A small site can expose Googlebot to a surprisingly large URL inventory when filter combinations, page numbers and language variants each create separate addresses. The fix is not automatically to “save crawl budget.” First decide which URLs deserve to be found in Search, then make those pages easy to discover and constrain the URL patterns that add no value. For most small, slowly changing sites, Google recommends an up-to-date sitemap and periodic checks of the Page Indexing report—not advanced crawl-budget tuning.
Why can a small site have so many URLs for Googlebot to crawl?
A URL is not necessarily a unique piece of useful content. A product listing with several filter parameters, multiple pages of results, and separate language paths can expose many URLs even if the site has relatively few underlying products or articles. Some of those URLs are useful landing pages; others are duplicate, empty, invalid or low-value variations.
Google defines crawl budget as “the set of URLs that Google can and wants to crawl,” combining crawl capacity—the resources Google can spend on a host—with crawl demand—its interest in revisiting known URLs. Capacity can be affected by response time, latency, server errors such as 5xx responses, and 429 rate limiting. Demand depends on factors including the URL inventory Google perceives, popularity, freshness and relevance. Crawling a URL does not guarantee that Google will index it.
Google’s guidance is aimed mainly at very large sites—roughly 1 million or more unique pages with moderate change—and medium-or-larger sites—roughly 10,000 or more unique pages with very rapid change. These are rough examples, not universal thresholds. Sites outside those groups generally need an accurate sitemap and periodic Page Indexing report checks. Google treats a hostname as a site for crawl-budget purposes, so separate hostnames may have separate budgets.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
Do faceted URLs waste crawl budget?
They can. Faceted navigation lets visitors narrow a listing by attributes such as size, color or price. If every combination creates a distinct parameter URL, the number of possible addresses can grow much faster than the number of useful landing pages. Googlebot may fetch many combinations before it can determine that they are not useful, while crawling indexable facets can also add server load and slow discovery of newer useful URLs.
When filtered pages should not appear in Google Search
If visitors need the filters but the resulting pages have no search value, prevent crawling of the unwanted patterns with a narrow robots.txt rule. Google recommends this as the direct way to keep crawlers from fetching faceted URLs; its examples show how to block selected filter parameters while allowing the desired unfiltered listing. URL fragments are generally not crawled or indexed by Google Search, so using fragments for filters also avoids creating crawlable parameter URLs.
Rank #2
Robots.txt is a crawl control, not a guarantee that a URL will never be known or shown: Google may know a blocked URL without fetching its content. A noindex directive serves a different purpose. Google must crawl the page to read it, so noindex is unsuitable when the goal is to stop crawl requests. A canonical can help consolidate duplicate URLs, but it is not a guaranteed crawl block; Google says canonicalization and nofollow are less effective long-term for controlling facet crawling than robots.txt or fragment-based filtering. For nofollow to affect a URL, every link to that URL must carry the attribute.
When filtered pages should be eligible for Search
Keep the URL space predictable: use conventional parameter separators such as &, or a consistent canonical ordering for path-encoded filters; avoid duplicate filters; and ensure each intended landing page has a stable URL. Return a real 404 under the requested URL for empty results, duplicate or nonsensical filter sets, and invalid pagination URLs. Do not send empty results to a shared error URL, because that hides the requested URL’s actual not-found status.
How should Googlebot discover paginated pages?
Give each meaningful page in a sequence its own stable URL, such as ?page=2, and a canonical pointing to itself. Do not canonicalize every page to page one: later pages may contain items that are not present on the first page. Link sequentially from one page to the next with ordinary <a href="..."> links; links back to the first page can also help visitors and crawlers navigate the sequence.
Do not use URL fragments such as #page=2 for pagination. Google generally ignores fragments for crawling, so it may not discover the next page as a distinct URL. Google no longer uses rel="next" and rel="prev" to identify pagination relationships, although other search engines may use them.
Rank #4
Load more and infinite scroll
A load-more button or infinite-scroll interface can work for visitors, but it should expose persistent paginated URLs and crawlable sequential links for content that needs to be discoverable. Google generally follows URLs in href attributes; it does not click buttons and generally does not trigger JavaScript interactions that require a user action. A sitemap can supplement links, and product catalogs can also use Merchant Center feeds as a discovery aid.
Sorting and filtering on long sequences can create duplicate variations. Use an appropriate noindex directive when Google may crawl a variation but should exclude it from results, or robots.txt when fetching itself should be prevented. Keep any blocking rule narrow enough not to hide useful pages in the pagination sequence.
Best Value
How do you make every language version reachable?
Give each language version its own URL rather than relying on cookies or browser settings to change one address’s content. Google recommends separate locale URL configurations annotated with hreflang. These annotations identify relationships among versions; they do not create missing pages or guarantee crawling, indexing or ranking. Each desired version still needs an accessible URL and a path Google can discover.
Locale-adaptive pages can be missed when they vary only in response to inferred country, preferred language, cookies or browser settings. Google’s default Googlebot requests do not set Accept-Language, and its default crawler IPs appear to be US-based, although Google also crawls from geographically distributed locations. Do not assume that one locale-dependent URL will expose every version consistently.
- Keep the visible content and navigation on each page in one language; avoid side-by-side translations that make the page’s language unclear.
- Provide ordinary links so visitors can switch language or region rather than forcing an automatic redirect based on guessed preferences.
- Use hreflang annotations or sitemaps to identify alternate versions, and apply robots directives consistently across locale URLs.
Which control should you use for unwanted URLs?
Choose the control by the outcome you want. Blocking fetches and excluding a crawled page from search results are different jobs; a canonical is a consolidation signal, not a substitute for either guarantee.
| Control | What it does | Use it when |
|---|---|---|
| robots.txt | Prevents Googlebot from fetching matching URLs; it does not ensure a blocked URL is unknown or absent from results. | The URL pattern should not consume crawl requests. |
noindex |
Excludes a crawled page from Google Search when Google can fetch and read the directive. | The page can be crawled, but should not be eligible for indexing. |
| Canonical | Signals a preferred URL among duplicate or similar pages; it does not guarantee that variants will stop being crawled. | There are duplicates to consolidate and a preferred version to identify. |
| 404 or 410 response | Reports that the requested URL is not available. | The URL is invalid, removed or represents an empty result that should not resolve as a page. |
How to diagnose crawl problems before changing URL rules
- Check whether crawl-budget tuning is proportionate. If the site is small and changes infrequently, start with an up-to-date sitemap and regular reviews of the Search Console Page Indexing report.
- Check serving health. Review Search Console’s Crawl Stats report for availability, response behavior and crawl patterns. Slow responses, 5xx errors or 429 rate limiting can restrict crawling independently of how many URLs the site exposes.
- Inspect what Googlebot actually requests. Use server logs to review URL-level crawl history. Search Console does not offer a crawl-history filter for arbitrary paths. Group logged URLs by patterns such as filter parameters, sort orders, session identifiers, page numbers, locale paths, and empty or error-result URLs.
- Classify the URL groups. Decide which patterns correspond to useful pages that should be discoverable and which are redundant or invalid. A large crawl count alone does not show that useful pages are being neglected.
- Make desired pages consistent. For useful pages, check that URLs are stable and distinct, internal links are crawlable, canonicals are appropriate, locale annotations are consistent, and sitemap entries are included where useful. A sitemap is a discovery hint, not a promise that Google will crawl or index every listed URL.
- Constrain waste and verify the result. Use a narrow robots.txt pattern or simplify how navigation generates URLs; return true 404 or 410 responses for removed or invalid URLs. Ensure rules are consistent across locales and do not block resources needed to understand important pages. Recheck logs and reports after changes, then assess indexing separately from crawling.
Do not assume that blocking a group of already crawled URLs automatically transfers that crawl activity to other pages. Google says that blocking or hiding fetched pages does not shift crawl budget elsewhere unless Google is already reaching the site’s serving limits. The practical goal is to remove needless URL paths and make valuable pages accessible, not to manage a transferable quota.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

