Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web crawlers are automated clients that discover and fetch web resources. They follow links, read sitemaps and other URL sources, request pages, and may execute JavaScript before passing what they found to an indexing or monitoring system. Search engines use crawlers, but so do site-auditing, archiving, price-monitoring and security tools.

A crawler fetching a page does not mean that page will be indexed or shown in search. Crawling, indexing and serving results are separate stages.

What is a web crawler?

RFC 9309 defines crawlers as automated clients. A crawler (also called a spider, robot or bot) starts with one or more URLs, retrieves resources over HTTP, parses the responses and queues additional URLs it discovers. Its output may be raw pages, extracted links, structured data, monitoring records or documents prepared for a search index.

Crawlers are not inherently good or bad. Googlebot and other search bots support discovery; aggressive or deceptive bots can waste bandwidth or collect data against an operator’s wishes. A responsible crawler identifies itself, limits request rates, honors published access rules and avoids bypassing authentication.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Web-Crawler
  • SUPERHERO AND VEHICLE FIGURE SET: Many adventures with this Spidey and His Amazing Friends set, which includes a figure, vehicle, and accessory
  • ARTICULATED FIGURE: This 4" figure features multiple points of articulation for lots of action
  • TEAM SPIDEY ADVENTURES: Kids can be part of Team Spidey and create their own epic adventures with this Spidey and His Amazing Friends Vehicle Set
  • INSPIRED BY MARVEL'S CHILDREN'S DRAWING: Little kids can imagine saving the day with their favorite superheroes with this Spidey and His Amazing Friends toy, inspired by the cute kids show
  • ENDLESS ADVENTURES WITH SPIDEY AND HIS AMAZING FRIENDS TOYS: Other Spidey and His Amazing Friends Toys Available (sold separately and subject to availability)

How a crawler finds and processes pages

  1. Discovery: URLs come from hyperlinks, XML sitemaps, feeds, submitted URL lists, APIs and previously known pages. A discovered URL is only a candidate; it is not a promise that the crawler will fetch it.
  2. Scheduling: The crawler prioritizes URLs and chooses concurrency, delays, retries and recrawl times. Google says its algorithm considers how often and how many pages to fetch, and slows when responses indicate that a server is overloaded.
  3. Robots check: Before an automatic crawl, a crawler normally retrieves /robots.txt, parses the applicable rules and decides whether the URL may be requested.
  4. HTTP retrieval: It sends a request with a declared user agent and receives HTML, images, stylesheets, scripts, feeds or other resources. Status codes, redirects, headers, caching and content type affect what happens next.
  5. Parsing and link extraction: The crawler reads the response, extracts links and other metadata, normalizes URLs and adds new candidates to its queue.
  6. Rendering: A rendering-capable crawler runs JavaScript and fetches referenced resources. A simple crawler may process only the original HTML.
  7. Indexing or storage: A search system analyzes text, images, video, titles, alt text, canonical relationships and duplicates, then stores an eligible representation. Other tools may save the response or produce an audit report.
  8. Serving: For a search engine, a separate retrieval and ranking system selects results for a user’s query. Crawling alone does not determine ranking or visibility.

How Googlebot crawls a website

Google describes Search as crawling, indexing and serving. Googlebot discovers URLs through links, sitemaps and known addresses, checks access rules, fetches pages and can render JavaScript with a Chrome-based rendering service. It then sends content and metadata through indexing systems that assess duplicates, canonical URLs and eligibility.

Google operates Smartphone and Desktop Googlebot variants. Both use the Googlebot product token in robots.txt, so that token cannot selectively allow one variant and block the other. Google says most Search crawling uses the mobile crawler. Crawling runs across many machines, and Google adjusts its rate when server errors or overload signals appear.

Not every discovered page reaches every stage. Access restrictions, server health, redirects, duplicate content, directives and quality systems can prevent fetching, indexing or serving.

Crawling versus indexing versus ranking

Stage What happens What it does not guarantee
Crawling A bot discovers and requests a URL and its resources. That the content is stored or appears in search.
Indexing A system analyzes content, metadata, links, duplicates and canonical signals and stores an eligible representation. That the page will rank or be shown for a query.
Serving Retrieval and ranking systems choose results for a particular search. Permanent placement, traffic or a guarantee that every indexed URL is displayed.

What robots.txt does—and does not do

Google describes robots.txt as a way to tell crawlers which URLs they may access and mainly to manage crawl traffic. RFC 9309 makes clear that these rules are not access authorization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • A disallow rule can stop a compliant crawler from fetching a path.
  • It is not a security boundary. Use authentication and authorization to protect private data.
  • Blocking a URL does not guarantee removal from Google. A blocked URL can still appear without a snippet if Google learns its address elsewhere.
  • To keep content out of Google, use authentication for private material or an appropriate noindex directive where the crawler can access the response.

Rules apply according to the crawler’s implementation. Check the documentation for the bot you are dealing with, and ensure the file is available at the site’s root over the expected scheme and host.

How crawlers handle JavaScript

There are two broad models. An HTML crawler parses the server response and sees links and content present in that response. A rendering crawler runs JavaScript, builds a page, fetches CSS and scripts, and then extracts the resulting content and links. Rendering costs more time and resources, so it may be delayed, limited or omitted.

For search visibility, put essential text and links in server-rendered HTML when practical. Use descriptive titles, useful alt attributes, stable canonical signals and ordinary crawlable links. Test the rendered result separately from the initial response: a browser can display content that a crawler cannot reach because a script failed, an API requires authentication or a link exists only after an interaction.

Designing or evaluating a crawler

Compare crawler implementations on the dimensions that affect results:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Discovery: links, sitemaps, feeds, APIs or submitted URL lists.
  • Fetch policy: concurrency, rate limits, retries, caching, redirect handling and response limits.
  • Rendering: static HTML only, or JavaScript execution with resource fetching.
  • Compliance: robots.txt handling, truthful user-agent identification, authentication boundaries and opt-out support.
  • Output: raw pages, links, structured data, index documents, screenshots or monitoring reports.
  • Freshness and scale: recrawl schedules, change detection, storage and distributed operation.

A minimal, polite Python crawler

This example stays on one host, honors robots.txt, limits requests and records links. It is a learning example, not a production search engine.

import time
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup

start = "https://example.com/"
host = urlparse(start).netloc
agent = "ExampleResearchBot/1.0 (+https://example.com/bot-info)"
robots = RobotFileParser(urljoin(start, "/robots.txt"))
robots.read()
q, seen = deque([start]), {start}
s = requests.Session()
s.headers["User-Agent"] = agent
while q and len(seen) <= 100:
    url = q.popleft()
    if not robots.can_fetch(agent, url):
        continue
    try:
        r = s.get(url, timeout=15)
        if "text/html" not in r.headers.get("content-type", ""):
            continue
        soup = BeautifulSoup(r.text, "html.parser")
        print(r.status_code, url, soup.title.get_text(strip=True) if soup.title else "")
        for a in soup.select("a[href]"):
            link = urldefrag(urljoin(url, a["href"]))[0]
            if urlparse(link).scheme in ("http", "https") and urlparse(link).netloc == host and link not in seen:
                seen.add(link); q.append(link)
    except requests.RequestException as e:
        print("fetch failed", url, e)
    time.sleep(1)

In production, add bounded response sizes, redirect and content-type policies, persistent queues, retries with backoff, duplicate detection, metrics, cancellation and stronger URL normalization. Never crawl authenticated areas without explicit authorization.

Or skip the browser setup

If your task is to obtain a reliable rendered view rather than build a discovery crawler, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF; it can render JavaScript, load lazy images, wait for network idle or a selector, and apply headers, cookies, user agents, geolocation and other controls.

It removes cookie-consent banners, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info and capture_pdf—let Claude, Cursor and other MCP clients request captures.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo documentation for the full option set, including CSS selectors, device presets, PDFs, custom JavaScript, blocking rules, caching, signed links, asynchronous jobs, bulk capture and usage data. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is included on every plan. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting crawler problems

The crawler gets 403 or 429 responses

Check authorization, rate limits and your user-agent identification. Reduce concurrency, add backoff and obtain permission; do not try to evade a site's controls.

Important links are missing

Inspect the raw HTML, not only the browser view. Confirm links are real anchors, not produced after a failed script, and publish an XML sitemap for large or difficult sites.

A page is discovered but not indexed

Discovery is not indexing. Check whether the server is healthy, the page is blocked, marked noindex, a duplicate with another canonical, or otherwise ineligible.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

robots.txt seems ignored

Verify the exact host, scheme, path and file syntax, then confirm the bot actually supports the rules. Remember that robots.txt controls compliant fetching, not security or guaranteed removal.

Rendered content differs from the browser

Look for JavaScript errors, delayed API calls, authentication requirements, resource blocking and timing limits. Compare a server-rendered fetch with a JavaScript-capable render.

Frequently Asked Questions

Do all crawlers obey robots.txt?

No. Robots.txt is a voluntary protocol for compliant automated clients, not an access-control mechanism.

Can crawling be prevented completely?

You can use authentication, network controls and operational defenses, but public pages may still be requested by clients that ignore robots.txt.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does a sitemap force Google to crawl every URL?

No. A sitemap supplies discovery hints; Google still schedules URLs based on its systems and site conditions.

Why can Google show a blocked URL?

Google may learn the URL from links or other sources even when robots.txt prevents fetching its content, resulting in limited or no snippet text.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.