For a single page, Python’s urllib.request can fetch a URL. To visit pages systematically, follow links, extract structured data, and export results, use Scrapy: define a spider that starts with URLs, parses responses, and schedules relevant links. Before crawling, identify your crawler and check the site’s robots.txt, terms, and applicable legal requirements.
Fetch one page or crawl a site?
Fetching retrieves a URL; crawling repeatedly discovers and visits pages. Use the standard library for a one-off fetch or small script. Choose Scrapy when you need link traversal, request scheduling, structured items, exports, or controls for crawl behavior. Scrapy’s framework includes a scheduler, downloader, spiders, items, pipelines, and feed exports (Scrapy overview).
Fetch a single URL with urllib
from urllib.request import urlopen
url = "https://example.com/"
with urlopen(url, timeout=20) as response:
html = response.read().decode("utf-8", errors="replace")
print(html[:500])
Replace the example URL with a page you are permitted to fetch. This retrieves the response body; it does not discover links, manage a crawl queue, or create structured records for you. A page’s character encoding may differ from UTF-8, so decoding with replacement is a simple display-oriented choice, not a guarantee of lossless text extraction. Python’s urllib HOWTO documents the basic urlopen() pattern.
Build a basic Scrapy crawler
The example below creates a small project, visits pages on a site you control or have permission to crawl, extracts titles and page URLs, and follows links only within the starting host. The extraction selectors are examples; adjust them to match the target page’s HTML.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
1. Install Scrapy and create a project
python -m pip install scrapy
scrapy startproject sitecrawl
cd sitecrawl
scrapy genspider pages example.com
Use the installation instructions for your Python environment and consult the Scrapy installation documentation if installation fails. The documentation surfaced for this guide is Scrapy 2.19.0; check the documentation matching your installed version because commands and defaults can change.
2. Set an identifiable user agent
In the project’s sitecrawl/settings.py, set a descriptive user agent with a contact method you actually monitor. Do not copy a fictitious contact address.
USER_AGENT = "SiteCrawl/1.0 (+https://your-domain.example/contact)"
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 4
DOWNLOAD_DELAY = 1.0
Replace the sample contact URL with a real one, or use a monitored email address in the user-agent text. The concurrency and delay values here are conservative starting settings, not universal requirements: choose a pace appropriate to the site and its instructions. Scrapy supports concurrency and crawl controls; high speed should not be the goal.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
3. Write a spider that extracts and follows links
Replace sitecrawl/spiders/pages.py with the following. The spider starts at the site root, yields a record for each response, and schedules links whose host matches the starting host.
Free tools Windows power users keep installed
One-click scans. No signup required.
from urllib.parse import urlparse
import scrapy
class PagesSpider(scrapy.Spider):
name = "pages"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/"]
def parse(self, response):
title = response.css("title::text").get()
yield {
"url": response.url,
"status": response.status,
"title": title.strip() if title else None,
}
for href in response.css("a::attr(href)").getall():
target = response.urljoin(href)
if urlparse(target).hostname == "example.com":
yield response.follow(target, callback=self.parse)
Change allowed_domains and start_urls together to your target host. The hostname check keeps the example from following external links; for subdomains or a multi-host site, define scope deliberately rather than removing the check without replacement. The spider’s callback can yield both extracted items and more requests, which is Scrapy’s central crawl loop (Scrapy spiders).
4. Run the crawl and export records
scrapy crawl pages -O pages.jsonl
-O writes an output feed and overwrites an existing file with the same name. Use -o when you want to append to an existing feed instead. Scrapy’s tutorial covers project setup, spider execution, and item export (Scrapy tutorial). For larger workflows, item pipelines can validate, clean, and store records, while feed exports can write to supported destinations.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Choose how the crawler discovers pages
| Approach | Best fit | Trade-off |
|---|---|---|
urllib.request |
One-off retrieval or a small fetch script | Simple, but link traversal, scheduling, and structured output must be built separately. |
| Scrapy Spider | Custom traversal and parsing | Flexible callback-driven control; you write and maintain the crawl logic. |
CrawlSpider |
Regular sites whose links fit rule-based following | Convenient rules, but not suited to every site; custom callbacks need care. |
SitemapSpider |
A target with useful sitemap URLs | Discovers URLs from sitemap structure rather than relying only on page links. |
Scrapy documents CrawlSpider and SitemapSpider. Pick based on the target’s structure, available sitemap, extraction needs, and how much custom logic is required. A plain spider is often clearest when traversal is unusual.
Set scope, pace, and permissions before running
Check robots.txt and site requirements
The Robots Exclusion Protocol places a robots file at the site’s top-level /robots.txt path. For example, check https://example.com/robots.txt before crawling that host. RFC 9309 defines the protocol and specifies UTF-8 encoding for the file (RFC 9309). Scrapy supports robots.txt handling, and ROBOTSTXT_OBEY = True enables the setting in the example.
Robots rules are crawler instructions, not proof that a crawl is permitted and not a substitute for reviewing site terms or applicable law. Requirements can depend on the site, data, purpose, and jurisdiction. The technical sources here do not determine legal permission for a particular crawl.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Constrain what gets crawled
- Start from the narrowest useful set of URLs and keep allowed hosts explicit.
- Filter links by path or other criteria when only a section of a site is in scope.
- Choose concurrency and delay in light of the site’s published guidance and observed behavior.
- Stop or reduce requests if the site returns errors, indicates overload, or asks you to stop.
Extract fields that match the page
The basic spider extracts the document title and response status. For a real task, inspect representative HTML and choose selectors for the fields you need. CSS selectors such as response.css("h1::text").get() retrieve a first match; .getall() retrieves all matches. Use response.urljoin(href) for relative links so they become absolute URLs before scheduling.
Do not assume every page uses the same markup or has a title. Missing values should be handled explicitly, as the example does for the title. If extraction requires pages rendered by client-side JavaScript, a basic HTTP spider may not see the rendered content; the sources cited here do not establish a universal rendering solution. Scrapy’s overview points to browser-rendering extensions, but choose and configure one only if the target’s behavior requires it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, reliability, and output
Scrapy schedules requests and supports concurrent downloads, which is useful for crawling more than one page. Concurrency also increases load on the target, so tune it to the site rather than maximizing it. Delays, scope limits, and robots compliance should be decided before a large crawl, not after excessive requests have been sent.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Expect incomplete coverage when links are inaccessible, the site changes its structure, responses fail, or pages require a rendering approach outside the spider’s setup. No crawler can be assumed to reach every page. For reliable downstream use, validate extracted fields in an item pipeline and record enough context—such as URL and status—to identify missing or failed content. Scrapy’s feed export is sufficient for small jobs; larger workflows can use pipelines and configured export destinations.
Troubleshooting common problems
- The spider yields no items: Confirm the spider name, project directory, and start URL; run
scrapy listto inspect discovered spiders. Check that the callback is actually yielding a dictionary or item. - No links are followed: Inspect the response HTML and verify that the page contains ordinary anchor elements matching
a::attr(href). Check the hostname filter and ensure the target host matchesallowed_domains. - Pages are outside the intended scope: Tighten the host and path checks before rerunning. A hostname-only filter allows every path on that host.
- Some fields are empty: Inspect the page source and update selectors to match its actual markup. Check for missing fields rather than assuming every response has identical structure.
- Requests are disallowed or blocked: Review the site’s robots instructions and terms, reduce request pressure where appropriate, and stop if asked. Changing the user agent to disguise the crawler is not an appropriate fix.
- The export file seems incomplete: Review the crawl output for response errors and confirm whether you used
-O(overwrite) or-o(append). Validate the feed after the crawl before relying on it. - Scrapy is not found after installation: Ensure the environment where you ran
python -m pip install scrapyis the same environment from which you invokescrapy; activate the intended virtual environment and retry.
Or skip the browser setup
If your task is to capture page images or PDFs rather than extract records and follow links, ScreenshotNeo is a website screenshot API and MCP server. Its one-request API can return an image or PDF, but it is not a replacement for a crawler that discovers and parses a site’s pages. See the ScreenshotNeo site and API documentation.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server provides screenshot, page-info, and PDF-capture tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for free screenshots.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

