Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To crawl a website with Python, build a bounded queue-and-parse loop: start with seed URLs, fetch each page, extract the fields and links you need, normalize and deduplicate URLs, enforce a same-site rule and page budget, then save structured records. Python’s standard-library URL tools and urllib.robotparser are enough for a small crawl; add Beautiful Soup for convenient HTML extraction, or use Scrapy when you need reusable spiders, pagination, pipelines, exports and crawl controls.

This walkthrough starts with a small, runnable crawler and then shows how to make it safer and more reliable. It uses a descriptive user agent, checks robots.txt, stays on one host and stops after 50 pages. The example is intentionally conservative: crawling rules, terms of service, privacy obligations, copyright and local law still apply.

What a Python crawler actually does

A crawler is not a single scraping command. It is a workflow with six recurring operations:

  1. Seed: put one or more starting URLs in a queue.
  2. Fetch: request a page with a timeout and identifying user agent.
  3. Parse: read the response and extract titles, text, metadata or other fields.
  4. Discover: collect links, resolve relative URLs and remove fragments.
  5. Control: deduplicate URLs, enforce host/path scope, obey robots rules and cap pages.
  6. Persist: write records incrementally so a failure does not lose earlier work.

The queue is normally breadth-first: take the oldest URL, process it, then append newly discovered links. A set of canonicalized URLs prevents loops caused by repeated links, fragments or alternate spellings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prerequisites and a responsible crawl plan

Install the small-script dependencies

Python’s urllib, collections and URL utilities are built in. Install Beautiful Soup for robust HTML parsing:

python -m pip install beautifulsoup4

Use a virtual environment for repeatable projects:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
python -m pip install beautifulsoup4

Set boundaries before sending requests

  • Choose an explicit host and, if necessary, a path prefix.
  • Set a page budget, response timeout and conservative delay.
  • Read the site’s robots.txt and apply the rules to the user agent you send. Robots.txt is a technical signal, not complete legal permission; a disallowed URL can still be discovered through links.
  • Review terms of service, privacy requirements, copyright and applicable law.
  • Do not crawl login, checkout, private or clearly restricted areas.
  • Collect only fields needed for the stated purpose and protect personal data.

A complete small crawler with urllib and Beautiful Soup

Save this as crawl.py. It follows same-host links, removes URL fragments, honors robots rules, limits the crawl to 50 pages, validates content type, handles common HTTP failures and writes JSON Lines records as it goes.

from collections import deque
from json import dumps
from time import sleep
from urllib.error import HTTPError, URLError
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.request import Request, urlopen
from urllib.robotparser import RobotFileParser

from bs4 import BeautifulSoup

START_URL = "https://example.com/"
USER_AGENT = "ExampleResearchBot/1.0 (+https://example.com/bot-info)"
MAX_PAGES = 50
TIMEOUT_SECONDS = 20
DELAY_SECONDS = 1.0
ALLOWED_HOST = urlparse(START_URL).netloc


def canonicalize(url):
    url, _ = urldefrag(url)
    parsed = urlparse(url)
    if parsed.scheme not in {"http", "https"}:
        return None
    return url


robots = RobotFileParser(urljoin(START_URL, "/robots.txt"))
try:
    robots.read()
except Exception as exc:
    print(f"Could not read robots.txt: {exc}")

queue = deque([canonicalize(START_URL)])
queued = {canonicalize(START_URL)}
seen = set()

with open("pages.jsonl", "w", encoding="utf-8") as output:
    while queue and len(seen) < MAX_PAGES:
        url = queue.popleft()
        if not url or url in seen:
            continue
        parsed = urlparse(url)
        if parsed.netloc != ALLOWED_HOST:
            continue
        if not robots.can_fetch(USER_AGENT, url):
            print(f"Blocked by robots.txt: {url}")
            continue

        request = Request(url, headers={"User-Agent": USER_AGENT})
        try:
            with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
                content_type = response.headers.get_content_type()
                if content_type not in {"text/html", "application/xhtml+xml"}:
                    print(f"Skipping {content_type}: {url}")
                    seen.add(url)
                    continue
                html = response.read()
        except HTTPError as exc:
            print(f"HTTP {exc.code} for {url}")
            continue
        except (URLError, TimeoutError) as exc:
            print(f"Request failed for {url}: {exc}")
            continue

        seen.add(url)
        soup = BeautifulSoup(html, "html.parser")
        title = soup.title.get_text(" ", strip=True) if soup.title else ""
        record = {"url": url, "title": title}
        output.write(dumps(record, ensure_ascii=False) + "n")
        output.flush()
        print(record)

        for link in soup.select("a[href]"):
            next_url = canonicalize(urljoin(url, link["href"]))
            if not next_url:
                continue
            if urlparse(next_url).netloc != ALLOWED_HOST:
                continue
            if next_url not in seen and next_url not in queued:
                queue.append(next_url)
                queued.add(next_url)

        sleep(DELAY_SECONDS)

Replace START_URL and the example bot-information URL with values that identify your project. The output file contains one JSON object per line, making it easy to resume, stream into a data warehouse or inspect with command-line tools.

What to customize

  • Fields: add selectors such as soup.select_one("article h1"), JSON-LD extraction or metadata parsing.
  • Scope: add a path-prefix check, for example parsed.path.startswith("/docs/").
  • Budget: lower MAX_PAGES for a trial; never remove the cap without an operational reason.
  • Rate: increase DELAY_SECONDS for a busy server. Retry only transient failures and stop after repeated server errors.
  • Storage: persist status, HTTP code and fetch time if you need auditing or incremental recrawls.

URL normalization and duplicate control

urljoin converts relative links such as /about into absolute URLs. urldefrag removes fragments because /guide#install and /guide#api normally identify the same HTTP document. The example also tracks URLs when queued, not only when fetched, so a page linking to the same destination repeatedly does not inflate the queue.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For larger projects, define a canonicalization policy explicitly. Decide whether trailing slashes, default ports, case differences in hosts, query parameters and tracking parameters represent the same resource. Do not blindly remove query parameters: they can select meaningful content. Record the original URL if you transform it so results remain explainable.

When Beautiful Soup is enough

Beautiful Soup is a practical choice for a small, focused extraction job: it parses HTML/XML and offers readable CSS selectors. It does not provide a complete crawl scheduler, feed-export system or concurrency policy, so those remain your responsibility. It also does not execute page JavaScript; if the data is absent from the initial HTML, a browser-rendering integration or an API may be required.

When to choose Scrapy instead

Scrapy is an application framework for crawling websites and extracting structured data. Choose it when the project needs reusable spiders, recursive following, pagination, CSS/XPath selectors, crawl-depth restrictions, caching, middleware, feed exports or pipelines. Its documented feature set covers those concerns, while a hand-written queue requires you to implement each one.

The Scrapy project currently labels 2.19.0 as its latest release in September 2026; release details change, so verify the version before pinning it. Scrapy’s tutorial demonstrates a quotes spider, extraction, exports and recursive following. JavaScript-heavy sites still need a browser-rendering integration; Scrapy alone does not make client-rendered data appear in the original response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need urllib + Beautiful Soup Scrapy
One site or small page budget Good fit with minimal setup Works, but adds framework setup
Recursive crawling and pagination Manual queue and rules Spider/request patterns are built in
CSS/XPath selectors Beautiful Soup CSS selectors Selectors plus XPath
Feed exports and pipelines Implement them yourself Built-in support
Depth, caching and middleware Implement and maintain them Documented components
JavaScript-rendered pages Usually insufficient alone Add browser-rendering integration

Production safeguards

Reliability and performance

  • Use connect/read timeouts and cap response sizes before parsing.
  • Validate content type; skip binaries unless they are part of the stated objective.
  • Write records incrementally and retain a failure log.
  • Cache responses where appropriate to avoid repeated load during development.
  • Use bounded concurrency only after measuring server tolerance; concurrency is not a substitute for a crawl budget.
  • Retry a small number of times for transient network errors, with exponential backoff. Do not repeatedly retry 4xx responses or robots-denied URLs.

Privacy and access control

Keep credentials out of URLs and source control. Never guess or bypass authentication, paywalls, CAPTCHAs or access controls. Minimize personal-data collection, define retention, and secure the resulting files.

Troubleshooting common failures

HTTP 403 or 429

The server may block the user agent or rate. Confirm the site’s rules and terms, identify your crawler clearly, slow down, reduce scope and stop rather than rotate identities to evade controls.

Timeouts and connection resets

Lower concurrency, increase the timeout modestly, add backoff and check whether the host is overloaded. Persist completed records so you can resume without refetching everything.

Empty titles or missing content

Inspect the raw response. The desired data may be generated by JavaScript, loaded behind an API, or selected with the wrong CSS selector. Do not assume a browser view equals the HTML returned to urlopen.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots parser errors

A missing or malformed robots file is not permission to crawl without limits. Log the condition, apply your conservative policy, and obtain clarification from the site owner when the project matters.

Queue grows without bound

Enforce a page budget, host/path allowlist and URL policy. Query parameters, calendars and faceted navigation can generate effectively infinite URL spaces.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than HTML data extraction, ScreenshotNeo provides a single-call website screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Read the full parameter list in the ScreenshotNeo API documentation. This cURL example captures Stripe as WebP:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every plan includes its options, including full-page and element capture, device presets, custom CSS/JavaScript, waits, blocking rules, headers and cookies, PDF controls, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call and a usage API. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can I crawl a site without Beautiful Soup?

Yes. The standard library can fetch pages and find links, but Beautiful Soup makes HTML parsing and selector-based extraction considerably safer and clearer for small crawlers.

Does robots.txt make scraping legal?

No. It communicates crawl preferences and traffic controls. You must also review terms, privacy, copyright and applicable law.

Why does my crawler see less than Chrome?

Many pages render content in JavaScript after the initial response. A basic urllib crawler receives the original HTML; use an appropriate API or browser-rendering integration when the data is client-rendered.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should I resume an interrupted crawl?

Persist each successful record and a failure/status log as you go, then initialize the queue from unprocessed URLs rather than starting over.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.