Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a single page, Python’s urllib.request can fetch a URL. To visit pages systematically, follow links, extract structured data, and export results, use Scrapy: define a spider that starts with URLs, parses responses, and schedules relevant links. Before crawling, identify your crawler and check the site’s robots.txt, terms, and applicable legal requirements.

Fetch one page or crawl a site?

Fetching retrieves a URL; crawling repeatedly discovers and visits pages. Use the standard library for a one-off fetch or small script. Choose Scrapy when you need link traversal, request scheduling, structured items, exports, or controls for crawl behavior. Scrapy’s framework includes a scheduler, downloader, spiders, items, pipelines, and feed exports (Scrapy overview).

Fetch a single URL with urllib

from urllib.request import urlopen

url = "https://example.com/"
with urlopen(url, timeout=20) as response:
    html = response.read().decode("utf-8", errors="replace")

print(html[:500])

Replace the example URL with a page you are permitted to fetch. This retrieves the response body; it does not discover links, manage a crawl queue, or create structured records for you. A page’s character encoding may differ from UTF-8, so decoding with replacement is a simple display-oriented choice, not a guarantee of lossless text extraction. Python’s urllib HOWTO documents the basic urlopen() pattern.

Build a basic Scrapy crawler

The example below creates a small project, visits pages on a site you control or have permission to crawl, extracts titles and page URLs, and follows links only within the starting host. The extraction selectors are examples; adjust them to match the target page’s HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

1. Install Scrapy and create a project

python -m pip install scrapy
scrapy startproject sitecrawl
cd sitecrawl
scrapy genspider pages example.com

Use the installation instructions for your Python environment and consult the Scrapy installation documentation if installation fails. The documentation surfaced for this guide is Scrapy 2.19.0; check the documentation matching your installed version because commands and defaults can change.

2. Set an identifiable user agent

In the project’s sitecrawl/settings.py, set a descriptive user agent with a contact method you actually monitor. Do not copy a fictitious contact address.

USER_AGENT = "SiteCrawl/1.0 (+https://your-domain.example/contact)"
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 4
DOWNLOAD_DELAY = 1.0

Replace the sample contact URL with a real one, or use a monitored email address in the user-agent text. The concurrency and delay values here are conservative starting settings, not universal requirements: choose a pace appropriate to the site and its instructions. Scrapy supports concurrency and crawl controls; high speed should not be the goal.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

3. Write a spider that extracts and follows links

Replace sitecrawl/spiders/pages.py with the following. The spider starts at the site root, yields a record for each response, and schedules links whose host matches the starting host.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import urlparse

import scrapy


class PagesSpider(scrapy.Spider):
    name = "pages"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    def parse(self, response):
        title = response.css("title::text").get()
        yield {
            "url": response.url,
            "status": response.status,
            "title": title.strip() if title else None,
        }

        for href in response.css("a::attr(href)").getall():
            target = response.urljoin(href)
            if urlparse(target).hostname == "example.com":
                yield response.follow(target, callback=self.parse)

Change allowed_domains and start_urls together to your target host. The hostname check keeps the example from following external links; for subdomains or a multi-host site, define scope deliberately rather than removing the check without replacement. The spider’s callback can yield both extracted items and more requests, which is Scrapy’s central crawl loop (Scrapy spiders).

4. Run the crawl and export records

scrapy crawl pages -O pages.jsonl

-O writes an output feed and overwrites an existing file with the same name. Use -o when you want to append to an existing feed instead. Scrapy’s tutorial covers project setup, spider execution, and item export (Scrapy tutorial). For larger workflows, item pipelines can validate, clean, and store records, while feed exports can write to supported destinations.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Choose how the crawler discovers pages

Approach Best fit Trade-off
urllib.request One-off retrieval or a small fetch script Simple, but link traversal, scheduling, and structured output must be built separately.
Scrapy Spider Custom traversal and parsing Flexible callback-driven control; you write and maintain the crawl logic.
CrawlSpider Regular sites whose links fit rule-based following Convenient rules, but not suited to every site; custom callbacks need care.
SitemapSpider A target with useful sitemap URLs Discovers URLs from sitemap structure rather than relying only on page links.

Scrapy documents CrawlSpider and SitemapSpider. Pick based on the target’s structure, available sitemap, extraction needs, and how much custom logic is required. A plain spider is often clearest when traversal is unusual.

Set scope, pace, and permissions before running

Check robots.txt and site requirements

The Robots Exclusion Protocol places a robots file at the site’s top-level /robots.txt path. For example, check https://example.com/robots.txt before crawling that host. RFC 9309 defines the protocol and specifies UTF-8 encoding for the file (RFC 9309). Scrapy supports robots.txt handling, and ROBOTSTXT_OBEY = True enables the setting in the example.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots rules are crawler instructions, not proof that a crawl is permitted and not a substitute for reviewing site terms or applicable law. Requirements can depend on the site, data, purpose, and jurisdiction. The technical sources here do not determine legal permission for a particular crawl.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Constrain what gets crawled

  • Start from the narrowest useful set of URLs and keep allowed hosts explicit.
  • Filter links by path or other criteria when only a section of a site is in scope.
  • Choose concurrency and delay in light of the site’s published guidance and observed behavior.
  • Stop or reduce requests if the site returns errors, indicates overload, or asks you to stop.

Extract fields that match the page

The basic spider extracts the document title and response status. For a real task, inspect representative HTML and choose selectors for the fields you need. CSS selectors such as response.css("h1::text").get() retrieve a first match; .getall() retrieves all matches. Use response.urljoin(href) for relative links so they become absolute URLs before scheduling.

Do not assume every page uses the same markup or has a title. Missing values should be handled explicitly, as the example does for the title. If extraction requires pages rendered by client-side JavaScript, a basic HTTP spider may not see the rendered content; the sources cited here do not establish a universal rendering solution. Scrapy’s overview points to browser-rendering extensions, but choose and configure one only if the target’s behavior requires it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and output

Scrapy schedules requests and supports concurrent downloads, which is useful for crawling more than one page. Concurrency also increases load on the target, so tune it to the site rather than maximizing it. Delays, scope limits, and robots compliance should be decided before a large crawl, not after excessive requests have been sent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Expect incomplete coverage when links are inaccessible, the site changes its structure, responses fail, or pages require a rendering approach outside the spider’s setup. No crawler can be assumed to reach every page. For reliable downstream use, validate extracted fields in an item pipeline and record enough context—such as URL and status—to identify missing or failed content. Scrapy’s feed export is sufficient for small jobs; larger workflows can use pipelines and configured export destinations.

Troubleshooting common problems

  • The spider yields no items: Confirm the spider name, project directory, and start URL; run scrapy list to inspect discovered spiders. Check that the callback is actually yielding a dictionary or item.
  • No links are followed: Inspect the response HTML and verify that the page contains ordinary anchor elements matching a::attr(href). Check the hostname filter and ensure the target host matches allowed_domains.
  • Pages are outside the intended scope: Tighten the host and path checks before rerunning. A hostname-only filter allows every path on that host.
  • Some fields are empty: Inspect the page source and update selectors to match its actual markup. Check for missing fields rather than assuming every response has identical structure.
  • Requests are disallowed or blocked: Review the site’s robots instructions and terms, reduce request pressure where appropriate, and stop if asked. Changing the user agent to disguise the crawler is not an appropriate fix.
  • The export file seems incomplete: Review the crawl output for response errors and confirm whether you used -O (overwrite) or -o (append). Validate the feed after the crawl before relying on it.
  • Scrapy is not found after installation: Ensure the environment where you ran python -m pip install scrapy is the same environment from which you invoke scrapy; activate the intended virtual environment and retry.

Or skip the browser setup

If your task is to capture page images or PDFs rather than extract records and follow links, ScreenshotNeo is a website screenshot API and MCP server. Its one-request API can return an image or PDF, but it is not a replacement for a crawler that discovers and parses a site’s pages. See the ScreenshotNeo site and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server provides screenshot, page-info, and PDF-capture tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for free screenshots.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.