Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A scalable Python crawler is a controlled pipeline, not merely an asynchronous loop. Put URLs in a durable frontier, fetch them with connection reuse and bounded per-host concurrency, parse and normalize links, deduplicate before scheduling, persist results and state, and measure queue, latency, errors and host-level request rates. Start with one well-scoped crawler; add processes or machines only after you can explain how politeness, shared state and duplicate suppression will still work.

The pipeline you should build first

Keep the first version deliberately explicit. Every URL should move through a small number of stages:

  1. Scope and seeds: define starting URLs, allowed hosts, depth or path rules, accepted content types and a maximum response size.
  2. Frontier: store URLs with status, discovery time, retry count, next-eligible time and (when useful) the referring URL. Normalize and deduplicate before enqueueing.
  3. Fetcher: reuse connections, set connect and read timeouts, validate schemes and redirects, cap bytes read and apply a host-specific scheduler.
  4. Robots and politeness: identify the crawler, fetch and parse each host’s robots.txt, enforce delay and concurrency limits, and back off on errors or blocking responses.
  5. Parser: extract the record you need and candidate links. Canonicalize cautiously; query parameters can change content.
  6. Storage and observability: persist records and crawl state, then expose counters and timings so a restart does not silently repeat work.

This separation lets you replace an in-memory queue with a database or distributed queue without rewriting parsing code.

A small, bounded asyncio crawler

For a narrow crawl, a custom client makes the control flow visible. The example below uses aiohttp and beautifulsoup4 and intentionally favors safe defaults over maximum concurrency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m pip install aiohttp beautifulsoup4
import asyncio
import time
from collections import defaultdict
from urllib.parse import urldefrag, urljoin, urlparse, urlunparse
from urllib.robotparser import RobotFileParser

import aiohttp
from bs4 import BeautifulSoup

SEEDS = ["https://example.com/"]
ALLOWED_HOSTS = {"example.com"}
MAX_PAGES = 500
MAX_BYTES = 2_000_000
PER_HOST_CONCURRENCY = 2
PER_HOST_DELAY = 1.0
USER_AGENT = "ExampleResearchCrawler/1.0 (+mailto:ops@example.com)"


def normalize(raw, base=None):
    value = urljoin(base, raw) if base else raw
    value, _ = urldefrag(value)
    p = urlparse(value)
    if p.scheme not in {"http", "https"} or not p.netloc:
        return None
    host = p.hostname.lower() if p.hostname else ""
    if host not in ALLOWED_HOSTS:
        return None
    return urlunparse((p.scheme, p.netloc, p.path or "/", "", p.query, ""))


class HostPolicy:
    def __init__(self):
        self.limits = defaultdict(lambda: asyncio.Semaphore(PER_HOST_CONCURRENCY))
        self.last_request = defaultdict(float)
        self.robots = {}
        self.robots_lock = asyncio.Lock()
        self.host_locks = defaultdict(asyncio.Lock)

    async def allowed(self, session, url):
        p = urlparse(url)
        origin = f"{p.scheme}://{p.netloc}"
        async with self.host_locks[p.netloc]:
            if p.netloc not in self.robots:
                robots_url = origin + "/robots.txt"
                try:
                    async with session.get(robots_url, allow_redirects=True) as r:
                        if 400 <= r.status < 500:
                            text = ""
                        elif r.status != 200:
                            self.robots[p.netloc] = False
                            return False
                        else:
                            text = await r.text(errors="replace")
                    parser = RobotFileParser()
                    parser.set_url(robots_url)
                    parser.parse(text.splitlines())
                    self.robots[p.netloc] = parser
                except (aiohttp.ClientError, asyncio.TimeoutError):
                    # Conservative policy for server/network failure.
                    self.robots[p.netloc] = False
            rule = self.robots[p.netloc]
            return rule is not False and rule.can_fetch(USER_AGENT, url)

    async def wait_turn(self, host):
        async with self.limits[host]:
            gap = PER_HOST_DELAY - (time.monotonic() - self.last_request[host])
            if gap > 0:
                await asyncio.sleep(gap)
            self.last_request[host] = time.monotonic()


async def crawl():
    timeout = aiohttp.ClientTimeout(total=30, connect=10)
    connector = aiohttp.TCPConnector(limit=20, limit_per_host=PER_HOST_CONCURRENCY)
    frontier = asyncio.Queue()
    seen = set()
    policy = HostPolicy()
    for seed in SEEDS:
        url = normalize(seed)
        if url:
            seen.add(url)
            await frontier.put(url)

    async with aiohttp.ClientSession(
        timeout=timeout, connector=connector,
        headers={"User-Agent": USER_AGENT}
    ) as session:
        pages = 0
        while pages < MAX_PAGES and not frontier.empty():
            url = await frontier.get()
            host = urlparse(url).netloc
            if not await policy.allowed(session, url):
                frontier.task_done()
                continue
            await policy.wait_turn(host)
            try:
                async with session.get(url, allow_redirects=True) as response:
                    if response.status != 200:
                        continue
                    if response.content_length and response.content_length > MAX_BYTES:
                        continue
                    body = await response.content.read(MAX_BYTES + 1)
                    if len(body) > MAX_BYTES:
                        continue
                    content_type = response.headers.get("content-type", "")
                    if "text/html" not in content_type:
                        continue
                    soup = BeautifulSoup(body, "html.parser")
                    title = soup.title.get_text(" ", strip=True) if soup.title else ""
                    print({"url": str(response.url), "title": title})
                    pages += 1
                    for tag in soup.select("a[href]"):
                        child = normalize(tag["href"], str(response.url))
                        if child and child not in seen:
                            seen.add(child)
                            await frontier.put(child)
            except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
                print("fetch error", url, repr(exc))
            finally:
                frontier.task_done()

if __name__ == "__main__":
    asyncio.run(crawl())

This is a teaching-sized fetcher, not a production frontier. Add durable URL and record tables, retry state, content hashing, structured logging and a shutdown path before relying on it. A robots.txt network failure is treated as complete disallow here; a 4xx response is treated as unavailable and therefore permits access, matching the conservative interpretation of RFC 9309.

Robots.txt, identity and host politeness

RFC 9309 places UTF-8 robots rules at the top-level /robots.txt. After a successful download, parseable rules must be followed. The protocol says crawlers should follow at least five consecutive redirects. A 4xx response means the file is unavailable and may allow access; a server or network failure requires assuming complete disallow. Do not use a cached file for more than 24 hours unless the file is unreachable.

Matching uses the most specific applicable path rule. If Allow and Disallow are equivalent, Allow wins. Robots exclusion is guidance, not authorization or authentication: RFC 9309 states, “The Robots Exclusion Protocol is not a substitute for valid content security measures.” Obtain permission for restricted data and never treat a robots file as proof that private information is protected.

  • Use a stable, descriptive User-Agent and include a contact address when possible.
  • Keep concurrency and delay limits per host, not just globally.
  • Increase delay after 429, 503, connection resets or repeated timeouts.
  • Honor redirects only within your scope and re-check the destination host’s policy.
  • Do not download unbounded bodies; reject unexpected content types and schemes.

When Scrapy is the better starting point

Scrapy supplies a project structure, scheduler, downloader middleware, item pipelines, retry machinery and operational settings. Its documentation describes AsyncCrawlerProcess and AsyncCrawlerRunner for running spiders from scripts or integrating with an existing event loop. Coroutine callbacks can await additional requests; asyncio libraries such as aiohttp require asyncio support to be enabled in that integration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A minimal spider looks like this:

import scrapy

class ArticleSpider(scrapy.Spider):
    name = "articles"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    custom_settings = {
        "USER_AGENT": "ExampleResearchCrawler/1.0 (+mailto:ops@example.com)",
        "ROBOTSTXT_OBEY": True,
        "CONCURRENT_REQUESTS": 16,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "DOWNLOAD_DELAY": 1.0,
        "AUTOTHROTTLE_ENABLED": True,
        "AUTOTHROTTLE_START_DELAY": 1.0,
        "AUTOTHROTTLE_MAX_DELAY": 30.0,
        "DOWNLOAD_TIMEOUT": 30,
    }

    def parse(self, response):
        yield {
            "url": response.url,
            "title": response.css("title::text").get(),
        }
        yield from response.follow_all(
            response.css("a::attr(href)").getall(), callback=self.parse
        )

Scrapy’s global and per-domain limits, download delay and AutoThrottle are per crawler. Running four crawler processes does not preserve the same aggregate rate: their limits can multiply. Treat internal throughput and permitted target-site request rate as separate budgets. Choose Scrapy when scheduling, extraction, retries and project conventions matter more than owning every event-loop detail. Choose a small asyncio client for a deliberately narrow or educational job where you are prepared to implement the missing machinery. Neither is universally faster; target behavior, parsing, storage and network conditions determine useful throughput.

What “scaling” changes

Scale within one process

Raise connection and concurrency limits only after watching latency, memory, error rate and each host’s request rate. Parsing large HTML, decompression, text extraction and database writes can become the bottleneck while sockets remain idle. Use bounded queues so a fast producer cannot exhaust memory.

Run independent spiders

If jobs are independent, schedule separate spider runs with separate scopes and budgets. Give each run its own output namespace and monitoring. Aggregate host-level rates across all runs before increasing limits.

Partition one large crawl across machines

Scrapy’s documentation is explicit: “Scrapy doesn’t provide any built-in facility for running crawls in a distributed (multi-server) manner.” Its documented approach is to partition URL inputs across separate runs and machines. You must then provide a shared or partitioned frontier, durable state, cross-worker deduplication, retry ownership, robots policy and result aggregation. A hash of normalized host or URL can provide deterministic ownership, but links discovered later still need a route to the correct worker.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Processes and machines add overhead: duplicate DNS and connection pools, coordination traffic, storage contention and the risk of multiplying load on a single site. Add workers only when queue depth and service-level goals justify that complexity.

Frontier, storage and observability design

A durable frontier commonly stores url, normalized URL hash, host, status, attempt count, next-attempt time, lease owner and timestamps. Lease a URL before fetching and make completion idempotent so a crashed worker can safely retry it. Store extracted records separately from crawl state; a response can be successfully fetched even when parsing or persistence fails.

Track these engineering signals rather than relying on a single “pages per second” number:

  • queued, leased, completed, retried and permanently failed URLs;
  • queue depth and age of the oldest pending URL;
  • DNS, connect, time-to-first-byte and total latency;
  • HTTP status distribution, bytes received and response-size rejects;
  • duplicate rate before and after normalization;
  • parser and storage failure counts;
  • memory, open connections and per-host request rate.

Set alerts on sustained queue growth, rising retries, host-specific 429/503 responses and stale leases. Keep a crawl manifest containing code version, scope, policy settings and start time so results are reproducible.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common failures and fixes

The crawler overwhelms a site

Check aggregate concurrency across every process, not only one configuration file. Lower per-host concurrency, add delay, enable adaptive throttling and pause hosts returning 429 or 503.

Robots handling is inconsistent

Log robots fetch status, redirect chain, cache age and the rule that matched. Distinguish a 4xx unavailable file from a timeout or server error; the latter requires complete disallow under RFC 9309.

The queue grows without bound

Enforce host and path scope before enqueueing, normalize fragments and default ports, cap depth or page count, and deduplicate atomically in the frontier store.

Results contain duplicates

Do not strip every query string. Remove only parameters proven irrelevant to your target. Canonical links, redirects and URL case rules should be recorded rather than silently rewriting evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Memory rises during a long run

Use bounded queues, stream or cap response bodies, release parsed documents after extraction, and persist records incrementally. Inspect whether retries or a never-ending discovery loop are retaining references.

JavaScript pages appear empty

HTTP clients receive the server response, not a browser-rendered DOM. Decide whether the data has an underlying JSON endpoint, whether a permitted browser renderer is necessary, and how that renderer’s extra load fits your host policy.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your crawl needs clean screenshots or PDFs of pages, ScreenshotNeo handles the browser capture call without your maintaining a rendering stack. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A one-call capture is:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It also supports full-page and element captures, device presets, custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, timezone and geolocation, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.

Further reading

Web Scraping with Python, 3rd Edition by Ryan Mitchell (O’Reilly, February 2024) covers crawler models, traversing sites, Scrapy, storage, parallel scraping and proxies. The publisher describes it as an intermediate-to-advanced, 352-page book; verify current availability and terms directly with the publisher before buying.

Frequently Asked Questions

Should I store the complete response body?

Only when your retention, privacy and replay requirements justify it. Otherwise persist the fields needed for extraction, a content hash, status metadata and a short-lived diagnostic sample.

How should I handle a URL that changes content by cookie or login?

Model the session as part of crawl scope. Keep credentials and cookies out of logs, obtain authorization, and deduplicate on the session-aware representation rather than assuming one public URL has one stable document.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a canonical link enough to merge two URLs?

No. Treat it as a useful signal, then confirm with redirects, content hashes and site-specific parameter rules. Canonical tags can be missing, stale or intentionally different from the URL you need to preserve.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.