Scrapy is a Python framework for asynchronous web crawling and structured data extraction. It coordinates requests, link discovery, selectors, retries, throttling, validation, and exports so you can run repeatable crawls instead of maintaining a one-off parser. The current official documentation is labeled Scrapy 2.17.0 (checked August 18, 2026) and requires Python 3.10 or newer. Scrapy is excellent for mostly HTTP-accessible pages, but it is not a browser: JavaScript rendering, interactive logins, canvas content, and sophisticated anti-bot systems may require an API or browser integration.
This guide builds a complete crawler against the safe training site quotes.toscrape.com, then covers selectors, pagination, pipelines, exports, reliability, JavaScript diagnosis, testing, deployment, and alternatives.
What Scrapy does
Crawling means discovering and requesting pages. Scraping means selecting fields from those responses. Data extraction turns those fields into stable records, while automation adds scheduling, retries, throttling, persistence, and monitoring. Scrapy provides all four layers through spiders, requests, responses, selectors, a scheduler, downloader middleware, item pipelines, and feed exporters. Its official documentation also lists data mining, monitoring, and automated testing as uses. See the Scrapy documentation.
When Scrapy is a good fit
- Multi-page or multi-domain crawls with link following and pagination.
- Recurring jobs that need deduplication, retries, caching, throttling, and statistics.
- Structured exports to JSON, JSON Lines, CSV, XML, databases, or queues.
- Projects that need a maintainable separation between extraction, cleaning, and storage.
When a smaller tool is better
For one static page collected once, requests plus Beautiful Soup or lxml usually involves less setup. Scrapy becomes worthwhile when request orchestration and repeatability matter.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Scrapy architecture in one flow
Spider
↓ yields Requests
Engine
├── Scheduler
└── Downloader
↓
Response
↓
Spider callback
├── new Requests
└── Items
↓
Item Pipeline
↓
Feed exporter / database
The engine coordinates components. The scheduler queues requests, the downloader performs HTTP work, spiders interpret responses and yield new requests or items, and pipelines validate and transform items before storage. This design supports controlled concurrency rather than a manually managed loop.
Install Scrapy correctly
The official installation guide requires Python 3.10 or newer, supports CPython and PyPy, and recommends a dedicated virtual environment.
- Create and activate an environment:
python -m venv .venvmacOS/Linux:
source .venv/bin/activateWindows Command Prompt:
.venvScriptsactivate.batWindows PowerShell:
.venvScriptsActivate.ps1 - Install Scrapy:
python -m pip install Scrapy - Verify the executable and detailed version:
scrapy version scrapy version -v scrapy bench
Conda users can install from conda-forge:
conda install -c conda-forge scrapy
The current documentation is labeled 2.17.0, while an official Zyte tutorial still shows pip install scrapy==2.14.2. Treat that as a tutorial pin, not proof of the latest release. Either install the current package with python -m pip install Scrapy or deliberately pin and test a project version:
python -m pip install "Scrapy==2.17.0"
See the Zyte tutorial for its version-specific example.
Recommended Free Tools
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Create a project and first spider
- Create the project and enter it:
scrapy startproject quotes_project cd quotes_project - The generated layout includes
scrapy.cfg,items.py,middlewares.py,pipelines.py,settings.py, and aspiders/package. - Generate a spider skeleton:
scrapy genspider quotes quotes.toscrape.com - Replace the generated spider with:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
allowed_domains = ["quotes.toscrape.com"]
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
"url": response.url,
}
next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
Run it and export records:
scrapy crawl quotes -O quotes.json
This follows each “Next” link until none remains. The workflow mirrors the official tutorial.
Selectors: CSS, XPath, and the Scrapy shell
Selectors operate on the response body. Use .get() for the first match, .getall() for every match, and .re() or .re_first() when a regular expression is useful.
CSS examples
response.css("h1::text").get()
response.css(".price_color::text").get()
response.css("article.product_pod").getall()
response.css("a::attr(href)").getall()
XPath examples
response.xpath("//h1/text()").get()
response.xpath("//a[contains(., 'Next')]/@href").get()
response.xpath("//article[contains(@class, 'product_pod')]").getall()
XPath is useful when selection depends on text or document relationships, such as locating a link labeled “Next Page.”
Test before running a full crawl
scrapy shell "https://quotes.toscrape.com/"
response.css("div.quote span.text::text").getall()
response.css("small.author::text").getall()
response.xpath("//li[@class='next']/a/@href").get()
The shell separates a bad selector from missing content, a JavaScript-only page, a redirect, or a block, shortening the edit-run-debug cycle.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Pagination and detail pages
response.follow() resolves relative URLs against the current response:
next_href = response.css("li.next a::attr(href)").get()
if next_href:
yield response.follow(next_href, callback=self.parse)
For many links, use follow_all():
yield from response.follow_all(
response.css("article a::attr(href)"),
callback=self.parse_detail,
)
Keep list-page and detail-page callbacks separate when the fields differ. Stop pagination when the link is absent, and let Scrapy’s request fingerprinting prevent duplicate requests. Cursor-based APIs, POST pagination, and infinite scroll require requests that reproduce the site’s actual network behavior rather than blindly constructing page numbers.
Items, cleaning, and validation
Yielding dictionaries is adequate for a small crawl:
yield {
"name": name,
"price": price,
"url": response.url,
}
For a larger project, define a stable schema:
import scrapy
class ProductItem(scrapy.Item):
name = scrapy.Field()
price = scrapy.Field()
currency = scrapy.Field()
url = scrapy.Field()
Use an item pipeline to normalize and validate data:
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
from decimal import Decimal
class CleanPricePipeline:
def process_item(self, item, spider):
raw_price = item.get("price")
if raw_price:
item["price"] = Decimal(
raw_price.replace("$", "").replace(",", "").strip()
)
return item
Enable it in settings.py:
ITEM_PIPELINES = {
"quotes_project.pipelines.CleanPricePipeline": 300,
}
Pipelines are the right place for whitespace cleanup, numeric and date conversion, required-field checks, dropping incomplete records, deduplication, and database or queue writes. A crawler that exits successfully can still produce zero or incorrect records, so validate counts and field types explicitly.
Export results safely
scrapy crawl quotes -O quotes.json
scrapy crawl quotes -o quotes.jsonl
scrapy crawl quotes -O quotes.csv
scrapy crawl quotes -O quotes.xml
-Ooverwrites the destination.-oappends.- Appending repeatedly to a normal JSON array can create invalid JSON; JSON Lines is safer for incremental output because each record occupies one line.
Set encoding when needed:
FEED_EXPORT_ENCODING = "utf-8"
For production, send exports to durable object storage, a database, or a downstream queue instead of treating a local file as the system of record.
Control request rate and crawl behavior
Start conservatively in settings.py:
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 1
CONCURRENT_REQUESTS_PER_DOMAIN = 2
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_MAX_DELAY = 10
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
Concurrency improves throughput; delays and lower concurrency reduce load and may reduce blocking. AutoThrottle adjusts pacing from observed latency. Tune values for each site. The Zyte training tutorial’s CONCURRENT_REQUESTS_PER_DOMAIN = 8 and DOWNLOAD_DELAY = 0.01 are specific to its safe example, not universal production defaults.
ROBOTSTXT_OBEY is an operational signal, not a legal permission. Terms of service, copyright, privacy, authentication requirements, and applicable law are separate considerations.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBest Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Handle statuses, retries, and failures
| Result | What to check |
|---|---|
| 200 | The response arrived, but selectors may still be wrong. |
| 301/302 | Inspect the final URL and redirected content. |
| 403 | Authentication, bot detection, permissions, or missing headers may be involved. |
| 404 | The link may be stale or the item removed. |
| 429 | Reduce request pressure and respect rate limits. |
| 500–599 | Retry selectively and inspect server or gateway behavior. |
| Empty selector result | The markup changed, content is dynamic, or the response is a block page. |
Log enough context to diagnose silent failures:
self.logger.info(
"status=%s url=%s title=%r",
response.status,
response.url,
response.css("title::text").get(),
)
Use an errback for transport-level failures:
def parse(self, response):
yield scrapy.Request(
"https://example.com/detail",
callback=self.parse_detail,
errback=self.handle_error,
)
def handle_error(self, failure):
self.logger.error("Request failed: %r", failure)
Retries do not solve a 403 caused by a missing browser session or authentication; diagnose the cause first.
When JavaScript changes the approach
- Compare the browser’s live DOM with “View Source.”
- Inspect network requests in developer tools.
- Look for JSON or GraphQL endpoints carrying the data.
- Test whether an authorized direct request can retrieve that endpoint.
- Add browser rendering only when the underlying request cannot be reproduced reliably.
Use this escalation path:
| Target condition | Approach |
|---|---|
| Data in initial HTML | Scrapy requests and selectors. |
| Public JSON endpoint | Scrapy requests with JSON parsing. |
| JavaScript-only rendering or interaction | Scrapy plus Playwright/Selenium integration. |
| Anti-bot, geolocation, or difficult infrastructure | Managed browser, proxy, or extraction API. |
| Authenticated or restricted data | An authorized API or approved access method. |
Scrapy does not execute JavaScript automatically. Its documentation includes guidance on dynamic content and developer-tools inspection; see Scrapy documentation.
Testing and production maintainability
- Keep selectors centralized where practical and maintain fixtures with representative HTML.
- Test required fields, field types, pagination termination, and duplicate handling.
- Run with small limits before broad crawling.
- Log response status, final URL, crawl statistics, and item counts.
- Alert on sudden drops, spikes, or all-null fields instead of trusting process exit status.
- Pin dependencies and protect API keys and credentials with secrets management.
- Store results durably and define crawl, time, and spending limits.
- Review site rules and applicable obligations before each deployment.
Choosing an execution model
Local or self-hosted Scrapy
This offers the lowest software cost and maximum control, but you must operate scheduling, workers, observability, proxies, storage, and upgrades.
Scrapy Cloud
Scrapy Cloud is hosted execution and scheduling for Scrapy spiders. Zyte lists plans from $9 per Scrapy Unit per month, with one unit described as 1 GB RAM and one concurrent crawl; pricing and included retention should be confirmed at Scrapy Cloud pricing. It suits teams that already have spiders and mainly need hosted runs, not teams requiring complete infrastructure control or advanced browser automation. Signup is at Zyte signup.
Zyte API
Zyte API provides managed HTTP fetching, browser rendering, proxy and anti-blocking capabilities, with pricing based on target difficulty and request type. The pricing page displayed $0.13–$1.27 per 1,000 HTTP requests and $1.01–$16.08 per 1,000 browser-rendered requests when checked August 16, 2026; verify current rates at Zyte API pricing. Standard signup included $5 of first-month credit according to Zyte API pricing details. It is most appropriate when rendering, geolocation, or anti-bot infrastructure—not ordinary static HTML—is the main problem.
Managed datasets
Zyte Data is aimed at organizations that need recurring structured data without maintaining parsers and infrastructure. Its signup page showed plans from $450 per month when checked August 16, 2026; confirm current terms at Zyte signup. Buying data is a poor fit for learning, one-off research, or projects requiring complete schema and crawl control.
Quick Recap
Scrapy compared with alternatives
| Option | Best fit | Main trade-off |
|---|---|---|
requests + Beautiful Soup/lxml |
Small, one-off static extraction | Less orchestration and persistence. |
| Scrapy | Repeatable HTTP crawls with pipelines and exports | More project structure than a single script; no built-in browser. |
| Playwright | JavaScript-heavy pages and interactive workflows | Heavier browser resource use. |
| Selenium | Mature browser automation | Usually less efficient than Scrapy for high-volume static crawling. |
| Managed scraping API | Proxy, rendering, geolocation, or anti-bot operations | Usage cost, vendor dependence, and less infrastructure control. |
Production checklist
- Python 3.10+ and a pinned, verified Scrapy version.
- Selectors tested against saved fixtures.
- Pagination, deduplication, and required-field tests.
- Conservative concurrency, delay, AutoThrottle, and crawl limits.
- Durable output storage and monitoring for item-count anomalies.
- Secrets management for credentials and API keys.
- Documented response to 403, 429, redirects, schema changes, and dynamic content.
- Review of robots directives, terms, privacy, copyright, authentication, and applicable law.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

