Scrapy is a Python framework for crawling websites and turning responses into structured records. The current official documentation covers Scrapy 2.19.0, whose documented minimum is Python 3.10. A dependable first project is: create an isolated environment, generate a Scrapy project, write a spider, test selectors against the downloaded response, yield items, follow links, and export a feed. This guide walks through that workflow and explains what to do when a browser shows content that Scrapy cannot see.
What Scrapy does
Scrapy describes itself as “a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages.” A spider creates requests and parses responses. The scheduler queues requests, the downloader fetches them, and the engine coordinates the flow. Spiders yield either extracted items or more requests.
Items are key-value records. Feed exports serialize those records to JSON, CSV, XML and other supported formats. Item pipelines receive each item for cleanup, validation, duplicate filtering or custom persistence. Downloader middleware handles request and response behavior such as headers, authentication, retries, redirects and proxies; spider middleware processes objects entering callbacks or leaving them. Settings configure all of these components, and a spider can override project settings with custom_settings. Extensions are suited to cross-cutting work such as statistics and crawl-progress logging.
Scrapy 2.19.0 is the version represented by the current documentation index and was listed by the project as the latest release in September 2026. Recheck the documentation and release notes before pinning a version. The release includes a RemoteControl extension used by the Scrapy MCP server and an experimental aiohttp download handler; neither is required for a beginner crawl.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Install Scrapy safely
Use a virtual environment
- Install Python 3.10 or newer.
- Create and activate an environment in a new directory:
python -m venv .venv # macOS/Linux source .venv/bin/activate # Windows PowerShell .venvScriptsActivate.ps1 - Install Scrapy with pip:
python -m pip install --upgrade pip python -m pip install Scrapy - Confirm the command is available:
scrapy version
The official installation guide also documents conda-forge. On Windows, pip may require Microsoft C++ Build Tools because of compiled dependencies; conda-forge can avoid many of those setup problems. Optional extras add integrations such as HTTPX, S3, Google Cloud Storage, image pipelines and shell interfaces, but a basic crawl does not need them. If installation fails, use the platform-specific notes in the installation guide rather than mixing system and virtual-environment packages.
Create your first spider
The official tutorial uses the deliberately simple training site quotes.toscrape.com. Use it as a contained exercise, then independently check that any real target and your intended use are appropriate under its terms, access rules and applicable law.
Generate a project
scrapy startproject quotesdemo
cd quotesdemo
scrapy genspider quotes quotes.toscrape.com
This creates a project with a spiders directory. Replace the generated spider with:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
allowed_domains = ["quotes.toscrape.com"]
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
The callback extracts each quote and follows the next-page link until no link remains. response.follow() resolves a relative URL safely and schedules another request.
Free tools Windows power users keep installed
One-click scans. No signup required.
Inspect and test selectors before coding
Selectors are integrated through response.css() and response.xpath(). They wrap Parsel, which uses lxml. CSS is often quick for classes and attributes; XPath is useful for relationships, conditions and text nodes. Neither is universally better: choose the expression that matches the response structure and that your team can maintain.
Open a shell for a downloaded page:
scrapy shell "https://quotes.toscrape.com/"
Then try selectors interactively:
response.css("div.quote span.text::text").getall()
response.xpath("//div[contains(@class, 'quote')]//span[@class='text']/text()").getall()
response.css("li.next a::attr(href)").get()
Use .get() for one value, .getall() for a list, and strip or normalize values only after confirming the response. A selector is tied to the current HTML; a site redesign can make a previously valid expression return an empty list.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Run the crawl and export records
From the project directory, write JSON directly with a feed export:
scrapy crawl quotes -O quotes.json
For CSV, use:
scrapy crawl quotes -O quotes.csv
The -O option overwrites the file. Use -o when you want to append to an existing feed. Feed exports are the simplest route when records only need serialization; no custom pipeline is necessary just to write JSON or CSV.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Items, pipelines and settings
When an item class helps
For a small spider, yielding dictionaries is clear. Larger projects can declare an item so fields and downstream behavior are explicit:
import scrapy
class QuoteItem(scrapy.Item):
text = scrapy.Field()
author = scrapy.Field()
tags = scrapy.Field()
Use a pipeline for item-level rules
Pipelines are the right place for cleanup, validation, duplicate filtering and persistence. This example rejects records without text and trims whitespace:
from itemadapter import ItemAdapter
class CleanQuotePipeline:
def process_item(self, item, spider):
adapter = ItemAdapter(item)
text = adapter.get("text")
if not text:
raise ValueError("quote has no text")
adapter["text"] = text.strip()
adapter["author"] = (adapter.get("author") or "").strip()
return item
Enable it in settings.py:
ITEM_PIPELINES = {
"quotesdemo.pipelines.CleanQuotePipeline": 300,
}
Integer priorities run from lower to higher values. Put normalization before a persistence pipeline when the latter should receive cleaned data. The documentation’s book example similarly extracts title and price, validates them in a pipeline, then exports the item; its selectors are an example, not a promise about another site’s markup.
Follow links without losing control
Link following turns a one-page parser into a crawler. Keep a narrow allowed_domains, follow only links relevant to the task, and stop when a pagination link is absent. For a site with many pages, tune concurrency and delay rather than assuming maximum speed is appropriate.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Scrapy provides download delay, per-domain concurrency limits and AutoThrottle. AutoThrottle attempts to adapt settings to server load. A conservative project-level starting point is:
DOWNLOAD_DELAY = 1
CONCURRENT_REQUESTS_PER_DOMAIN = 4
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_MAX_DELAY = 10
These are configuration choices, not universal performance settings. Increase or decrease them according to response times, workload and the target’s access rules. Review terms, robots directives and relevant law; a robots.txt file alone does not establish permission for every use.
Why browser content is missing in Scrapy
Scrapy downloads an HTTP response; it does not automatically execute the same JavaScript and browser interactions as a modern browser. If text appears in a browser but not in response.text, diagnose the source in this order.
Find the underlying request
- Open browser developer tools and select the Network panel.
- Reload the page and filter for Fetch/XHR requests.
- Inspect response bodies, query parameters, request headers and pagination values.
- Reproduce the data request directly in a Scrapy callback when the data is available from that endpoint.
The dynamic-content guide also notes that data may be embedded in JavaScript or loaded from an external resource. Parse that source when practical instead of rendering an entire browser.
Escalate to a headless browser only when needed
If the desired content exists only after scripts run and cannot be reached through an underlying request, a headless browser is an option. It adds browser binaries, memory, startup time and more failure modes, so treat it as an escalation rather than the default for every page. Keep the extraction logic separate from browser-control code so that a direct endpoint remains easy to use if the site changes.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when your workflow needs a rendered visual or PDF instead of building browser automation yourself. One GET request returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.
Use the API directly (see the ScreenshotNeo documentation):
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
Troubleshooting common failures
“scrapy: command not found”
The virtual environment is not active, or Scrapy was installed into another interpreter. Activate .venv and run python -m pip show Scrapy; reinstall with that same Python executable.
Dependency build errors on Windows
Compiled dependencies may need Microsoft C++ Build Tools. Follow the installation guide’s Windows notes or use the documented conda-forge route.
Selectors return an empty list
Inspect the actual response in scrapy shell. Check that the selector matches the downloaded HTML, not an element created later by JavaScript, and verify class names, nesting and URL redirects.
Only some pages are crawled
Print or log the next URL, confirm it is inside allowed_domains, and check whether pagination uses a different link or an API cursor. Avoid generating duplicate requests.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Items are missing fields or never reach the output file
Confirm the callback yields the item, run with increased logging, and check that pipeline validation is not rejecting it. Feed export writes what survives the pipeline.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
The page returns a challenge, timeout or blank body
Record status codes and response headers, reduce concurrency, add a suitable delay, and determine whether the content requires authentication, a different request, or browser rendering. Do not assume a proxy or headless browser is required until the response path is understood.
Operational checklist
- Pin and periodically recheck your Scrapy and Python versions.
- Keep secrets out of source code; use settings or environment variables for credentials.
- Test selectors against saved responses so a site change is detectable.
- Record item counts, HTTP errors and pipeline drops.
- Use feed exports for straightforward files and pipelines for transformation or storage.
- Set delay, concurrency and AutoThrottle deliberately for each crawl.
- Review the target’s terms, access controls and applicable legal requirements before collecting data.
Frequently Asked Questions
Can Scrapy scrape a site without JavaScript?
Yes. If the required data is present in the downloaded HTML or an underlying API response, a normal Scrapy spider can extract it without a browser.
Should I choose CSS or XPath selectors?
Use whichever expression best matches the response structure and your team’s maintenance skills; Scrapy supports both through response.css() and response.xpath().
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Do I need an item pipeline to create a JSON file?
No. Feed exports can serialize yielded dictionaries or items directly. Add a pipeline when you need validation, cleanup, filtering or custom persistence.
What Python version does Scrapy 2.19 document?
The current installation documentation lists Python 3.10 or newer.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

