Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesTo crawl a page with Scrapy, create a Python project, write a spider that requests a starting URL, extract fields from each response, and yield more requests for links you want to follow. Then run the spider and export its results to a file. This walkthrough uses the official Scrapy tutorial site, quotes.toscrape.com, so you can try the complete workflow with predictable example markup.
The commands and spider syntax below follow Scrapy’s official documentation version 2.19.0, presented on September 30, 2026. Scrapy’s installation guidance requires Python 3.10 or newer; check the current documentation if you are using a different version or platform.
What Scrapy does—and what a spider is
Scrapy is a Python framework for crawling websites and extracting structured data. A spider is the part of your project that defines which pages to request and how to parse their responses. As the official documentation puts it, “Spiders are classes that you define and that Scrapy uses to scrape information from a website (or a group of websites).”
This is different from fetching one page manually with an HTTP library: Scrapy provides a request-and-response workflow, link following, item processing, and feed exports. You describe the pages and data you want; Scrapy schedules requests and passes downloaded responses to your callbacks.
#1 Best Overall
A crawler should identify itself and follow the target site’s rules. Permission, terms, applicable law, and acceptable request rates depend on the site and use case; a tutorial example does not grant permission to crawl another website.
Install Scrapy and create a project
Use a virtual environment so the project’s dependencies do not interfere with other Python projects or system packages. Scrapy’s current installation documentation lists Python 3.10 or newer and supports installation with pip or conda-forge. This walkthrough uses pip.
-
Check that Python is available. On many systems, use
python3 --version; on Windows,py --versionmay be the appropriate command. Confirm the version is 3.10 or newer. -
Create and activate a virtual environment from the directory where you want to work:
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.python -m venv .venv # macOS/Linux source .venv/bin/activate # Windows PowerShell .venvScriptsActivate.ps1 # Windows Command Prompt .venvScriptsactivate.batUse only the activation command for your shell. If your system uses
python3instead ofpython, substitute it when creating the environment. -
Install Scrapy and create the tutorial project:
python -m pip install Scrapy scrapy startproject tutorial cd tutorial -
Check that Scrapy is installed and the project command is available:
scrapy version
The generated project includes settings, item and pipeline modules, and a spiders directory. A spider can be run from this project directory, where Scrapy can discover it.
Installation issues to watch for
Scrapy relies on packages including lxml, parsel, w3lib, Twisted, cryptography, and pyOpenSSL. Some dependencies can require platform-specific setup. If installation fails while building or installing a dependency, first verify the Python version, activate the intended environment, and consult the current Scrapy installation instructions for your operating system rather than mixing packages into the system Python.
Free tools Windows power users keep installed
One-click scans. No signup required.
Create a spider that extracts data and follows pages
Create a file named quotes_spider.py in tutorial/spiders/. The spider below extracts quote text, author, and tags from each page, then follows the “next” link until there are no more pages.
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
async def start(self):
yield scrapy.Request("https://quotes.toscrape.com/")
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
"source_url": response.url,
}
next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
The unique name identifies the spider within the project; Scrapy uses it when you run the spider from the command line. In Scrapy 2.19’s tutorial, start() is an asynchronous generator that yields request objects. You may encounter older tutorials using a different starting-request interface, so match code to the Scrapy version you have installed.
parse() receives a downloaded response. Each dictionary yielded from it is an item for Scrapy to process or export. The example’s selectors match the tutorial site’s markup, not a universal website structure. Inspect the target page’s actual HTML and adapt the selectors before relying on them.
Set an identifying user agent
Before crawling, set a descriptive USER_AGENT in tutorial/settings.py. The setting lets site owners identify and contact the crawler operator. For example, replace the value with an identifier and contact address you control:
USER_AGENT = "ExampleResearchBot (+https://example.com/contact)"
Do not copy the example identity as if it were yours. Use an accurate description of your crawler.
Choose CSS or XPath selectors from the page markup
Scrapy responses provide both response.css() and response.xpath(). CSS selectors are often concise when the target is easy to identify by element, class, or attribute. XPath can be convenient when selection depends on document structure or text content. Scrapy converts CSS selectors to XPath internally; neither method is universally better.
| Selector type | Useful when | Example | Trade-off |
|---|---|---|---|
| CSS | You can identify an element by its tag, class, or attribute. | response.css("li.next a::attr(href)").get() |
Concise for common HTML patterns; complex content-based conditions may be less direct. |
| XPath | The selection depends on relationships in the document or text content. | response.xpath("//li[@class='next']/a/@href").get() |
Can express structural and content conditions, but may be harder to read if you are unfamiliar with XPath. |
Use Scrapy’s shell to inspect a response and test selectors instead of guessing at the markup. From the project directory, run:
scrapy shell https://quotes.toscrape.com/
At the shell prompt, try expressions such as response.css("div.quote span.text::text").getall() and response.xpath("//div[@class='quote']//span[@class='text']/text()").getall(). Compare the returned values with the page you intend to collect. If a selector returns an empty list, the live markup may differ, the content may not be present in the downloaded response, or the selector may be wrong.
Rank #3
Run the spider and export its items
From the project directory, run the spider and write its yielded dictionaries to JSON Lines:
scrapy crawl quotes -O quotes.jsonl
The capital -O overwrites the output file if it already exists. If you want to append to an existing feed rather than overwrite it, Scrapy’s feed export command also supports lowercase -o; choose deliberately so repeated runs do not produce an unexpected file.
For a regular JSON array instead of one JSON object per line, use a JSON feed path:
scrapy crawl quotes -O quotes.json
Feed formats are inferred from the file extension. Scrapy can export yielded items to supported feed formats, which is usually the simplest way to save an introductory crawl. Inspect the resulting file to confirm it contains the expected fields and number of records before using it downstream.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11How the crawl proceeds, and how to control its scope
The spider starts with one request, parses its response, yields extracted records, and requests the next page when the selector finds a link. Scrapy resolves relative links in response.follow() against the response URL, so pagination links such as /page/2/ work without manually joining URLs.
For a real target, make the crawl scope explicit. Only follow the links needed for the task; a broad link-following rule can collect unrelated pages or create a much larger crawl than intended. Check that pagination terminates, and do not assume the example site’s markup or linking pattern applies elsewhere.
Pass a spider argument for a variable starting URL
Scrapy spiders can accept command-line arguments. For example, add an optional start_url argument to the spider and use it in start():
class QuotesSpider(scrapy.Spider):
name = "quotes"
def __init__(self, start_url=None, **kwargs):
super().__init__(**kwargs)
self.start_url = start_url or "https://quotes.toscrape.com/"
async def start(self):
yield scrapy.Request(self.start_url)
def parse(self, response):
# Keep the extraction and pagination logic here.
...
Run it with a supplied URL:
scrapy crawl quotes -a start_url=https://quotes.toscrape.com/ -O quotes.jsonl
This changes the starting URL; it does not make the extraction selectors fit arbitrary sites. A different site may need different parsing logic and a different link-following strategy.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →When to add an item pipeline
For the first crawl, feed export is enough. Add a pipeline when items need validation, cleaning, deduplication, or storage beyond a simple exported feed. A pipeline receives yielded items and can process them before they reach the output feed or another storage destination.
To activate a pipeline, add its dotted Python class path to ITEM_PIPELINES in tutorial/settings.py. The setting maps pipeline paths to numeric priorities; lower numbers run before higher ones. For example:
ITEM_PIPELINES = {
"tutorial.pipelines.ValidateQuotePipeline": 300,
}
The class path must match the module and class you actually create. Avoid adding a pipeline merely to save a basic crawl: start with feed export, then introduce processing when you have a concrete validation or storage requirement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common Scrapy crawl problems
-
scrapyis not recognized or not found. The virtual environment may not be active, or Scrapy may have been installed under a different Python interpreter. Activate.venvand runpython -m pip show Scrapy; install it into that same environment if needed.Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
Installation fails on a dependency. Confirm the active Python meets the current minimum requirement and follow the platform-specific installation guidance for the dependency error. Scrapy’s dependency stack includes compiled or platform-sensitive packages, so a failed install is not necessarily an error in your spider.
-
The spider is not found. Ensure the file is under the project’s
spidersdirectory, the class subclassesscrapy.Spider, it has a uniquename, and you runscrapy crawl quotesfrom the project directory. -
The crawl runs but exports no items. Check the log for download or callback errors, open the response in
scrapy shell, and test each selector against the actual HTML. The page may have changed, or the content may not exist in the downloaded response. -
Only the first page is collected. Inspect the next-link selector with the shell. Verify that it matches an
hrefon the first page and that subsequent pages use the same pattern. The callback must yield the result ofresponse.follow()for Scrapy to schedule it.What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.Best Value
-
A relative next-page URL fails. Use
response.follow(next_page, callback=self.parse)with the response that contains the link. It resolves relative links against that response’s URL; avoid treating the relative path as a complete URL yourself. -
The output file looks different from expected. Check the file extension and whether you used
-O(overwrite) or-o(append). JSON Lines and a JSON array are different feed formats; open a small output sample before building another step around it.
Performance, reliability, and responsible use
Scrapy is designed to manage crawling requests and parsing in a framework, but this walkthrough makes no benchmark claim about crawl speed. Actual performance depends on the target, network, response sizes, selectors, and crawl configuration. Begin with a small, narrow crawl, monitor its output and logs, and expand only when the target’s rules and your data requirements permit.
Reliability starts with selectors that are checked against real responses and output that is validated after export. Page markup can change, and a successful HTTP response does not guarantee that a selector extracted the intended data. For longer-running jobs, consider whether an item pipeline is needed to validate or deduplicate records before storing them.
Recommended Free Tools
Review the target website’s terms and applicable requirements before crawling. Scrapy documentation describes how to build the workflow; it does not determine whether a particular collection is permitted.
Or skip the browser setup
If your goal is a screenshot rather than structured records, Scrapy is not the right tool: it crawls and extracts data, while a screenshot API returns an image or PDF. ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return PNG, JPEG, WebP, or PDF.
For example, this cURL request saves a WebP screenshot of the tutorial site:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://quotes.toscrape.com/ -o shot.webp
See the ScreenshotNeo API documentation for the request options and response details. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Sign up for ScreenshotNeo free to try 1,000 screenshots a month without a card.
Frequently Asked Questions
Can Scrapy crawl a site whose content is rendered by JavaScript?
This walkthrough’s documented workflow extracts from Scrapy response content. Whether that response contains the content you need depends on the site; inspect it in Scrapy shell before choosing an approach.
Can I use Scrapy to take screenshots?
Scrapy is for crawling and extracting structured data, not producing screenshots. For a screenshot or PDF, use a screenshot tool such as ScreenshotNeo.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →

