Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To crawl a page with Scrapy, create a Python project, write a spider that requests a starting URL, extract fields from each response, and yield more requests for links you want to follow. Then run the spider and export its results to a file. This walkthrough uses the official Scrapy tutorial site, quotes.toscrape.com, so you can try the complete workflow with predictable example markup.

The commands and spider syntax below follow Scrapy’s official documentation version 2.19.0, presented on September 30, 2026. Scrapy’s installation guidance requires Python 3.10 or newer; check the current documentation if you are using a different version or platform.

What Scrapy does—and what a spider is

Scrapy is a Python framework for crawling websites and extracting structured data. A spider is the part of your project that defines which pages to request and how to parse their responses. As the official documentation puts it, “Spiders are classes that you define and that Scrapy uses to scrape information from a website (or a group of websites).”

This is different from fetching one page manually with an HTTP library: Scrapy provides a request-and-response workflow, link following, item processing, and feed exports. You describe the pages and data you want; Scrapy schedules requests and passes downloaded responses to your callbacks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A crawler should identify itself and follow the target site’s rules. Permission, terms, applicable law, and acceptable request rates depend on the site and use case; a tutorial example does not grant permission to crawl another website.

Install Scrapy and create a project

Use a virtual environment so the project’s dependencies do not interfere with other Python projects or system packages. Scrapy’s current installation documentation lists Python 3.10 or newer and supports installation with pip or conda-forge. This walkthrough uses pip.

  1. Check that Python is available. On many systems, use python3 --version; on Windows, py --version may be the appropriate command. Confirm the version is 3.10 or newer.

  2. Create and activate a virtual environment from the directory where you want to work:

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
    python -m venv .venv
    # macOS/Linux
    source .venv/bin/activate
    # Windows PowerShell
    .venvScriptsActivate.ps1
    # Windows Command Prompt
    .venvScriptsactivate.bat

    Use only the activation command for your shell. If your system uses python3 instead of python, substitute it when creating the environment.

  3. Install Scrapy and create the tutorial project:

    python -m pip install Scrapy
    scrapy startproject tutorial
    cd tutorial
  4. Check that Scrapy is installed and the project command is available:

    scrapy version

The generated project includes settings, item and pipeline modules, and a spiders directory. A spider can be run from this project directory, where Scrapy can discover it.

Installation issues to watch for

Scrapy relies on packages including lxml, parsel, w3lib, Twisted, cryptography, and pyOpenSSL. Some dependencies can require platform-specific setup. If installation fails while building or installing a dependency, first verify the Python version, activate the intended environment, and consult the current Scrapy installation instructions for your operating system rather than mixing packages into the system Python.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Create a spider that extracts data and follows pages

Create a file named quotes_spider.py in tutorial/spiders/. The spider below extracts quote text, author, and tags from each page, then follows the “next” link until there are no more pages.

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        yield scrapy.Request("https://quotes.toscrape.com/")

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
                "tags": quote.css("div.tags a.tag::text").getall(),
                "source_url": response.url,
            }

        next_page = response.css("li.next a::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

The unique name identifies the spider within the project; Scrapy uses it when you run the spider from the command line. In Scrapy 2.19’s tutorial, start() is an asynchronous generator that yields request objects. You may encounter older tutorials using a different starting-request interface, so match code to the Scrapy version you have installed.

parse() receives a downloaded response. Each dictionary yielded from it is an item for Scrapy to process or export. The example’s selectors match the tutorial site’s markup, not a universal website structure. Inspect the target page’s actual HTML and adapt the selectors before relying on them.

Set an identifying user agent

Before crawling, set a descriptive USER_AGENT in tutorial/settings.py. The setting lets site owners identify and contact the crawler operator. For example, replace the value with an identifier and contact address you control:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
USER_AGENT = "ExampleResearchBot (+https://example.com/contact)"

Do not copy the example identity as if it were yours. Use an accurate description of your crawler.

Choose CSS or XPath selectors from the page markup

Scrapy responses provide both response.css() and response.xpath(). CSS selectors are often concise when the target is easy to identify by element, class, or attribute. XPath can be convenient when selection depends on document structure or text content. Scrapy converts CSS selectors to XPath internally; neither method is universally better.

Selector type Useful when Example Trade-off
CSS You can identify an element by its tag, class, or attribute. response.css("li.next a::attr(href)").get() Concise for common HTML patterns; complex content-based conditions may be less direct.
XPath The selection depends on relationships in the document or text content. response.xpath("//li[@class='next']/a/@href").get() Can express structural and content conditions, but may be harder to read if you are unfamiliar with XPath.

Use Scrapy’s shell to inspect a response and test selectors instead of guessing at the markup. From the project directory, run:

scrapy shell https://quotes.toscrape.com/

At the shell prompt, try expressions such as response.css("div.quote span.text::text").getall() and response.xpath("//div[@class='quote']//span[@class='text']/text()").getall(). Compare the returned values with the page you intend to collect. If a selector returns an empty list, the live markup may differ, the content may not be present in the downloaded response, or the selector may be wrong.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run the spider and export its items

From the project directory, run the spider and write its yielded dictionaries to JSON Lines:

scrapy crawl quotes -O quotes.jsonl

The capital -O overwrites the output file if it already exists. If you want to append to an existing feed rather than overwrite it, Scrapy’s feed export command also supports lowercase -o; choose deliberately so repeated runs do not produce an unexpected file.

For a regular JSON array instead of one JSON object per line, use a JSON feed path:

scrapy crawl quotes -O quotes.json

Feed formats are inferred from the file extension. Scrapy can export yielded items to supported feed formats, which is usually the simplest way to save an introductory crawl. Inspect the resulting file to confirm it contains the expected fields and number of records before using it downstream.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the crawl proceeds, and how to control its scope

The spider starts with one request, parses its response, yields extracted records, and requests the next page when the selector finds a link. Scrapy resolves relative links in response.follow() against the response URL, so pagination links such as /page/2/ work without manually joining URLs.

For a real target, make the crawl scope explicit. Only follow the links needed for the task; a broad link-following rule can collect unrelated pages or create a much larger crawl than intended. Check that pagination terminates, and do not assume the example site’s markup or linking pattern applies elsewhere.

Pass a spider argument for a variable starting URL

Scrapy spiders can accept command-line arguments. For example, add an optional start_url argument to the spider and use it in start():

class QuotesSpider(scrapy.Spider):
    name = "quotes"

    def __init__(self, start_url=None, **kwargs):
        super().__init__(**kwargs)
        self.start_url = start_url or "https://quotes.toscrape.com/"

    async def start(self):
        yield scrapy.Request(self.start_url)

    def parse(self, response):
        # Keep the extraction and pagination logic here.
        ...

Run it with a supplied URL:

scrapy crawl quotes -a start_url=https://quotes.toscrape.com/ -O quotes.jsonl

This changes the starting URL; it does not make the extraction selectors fit arbitrary sites. A different site may need different parsing logic and a different link-following strategy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to add an item pipeline

For the first crawl, feed export is enough. Add a pipeline when items need validation, cleaning, deduplication, or storage beyond a simple exported feed. A pipeline receives yielded items and can process them before they reach the output feed or another storage destination.

To activate a pipeline, add its dotted Python class path to ITEM_PIPELINES in tutorial/settings.py. The setting maps pipeline paths to numeric priorities; lower numbers run before higher ones. For example:

ITEM_PIPELINES = {
    "tutorial.pipelines.ValidateQuotePipeline": 300,
}

The class path must match the module and class you actually create. Avoid adding a pipeline merely to save a basic crawl: start with feed export, then introduce processing when you have a concrete validation or storage requirement.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common Scrapy crawl problems

Performance, reliability, and responsible use

Scrapy is designed to manage crawling requests and parsing in a framework, but this walkthrough makes no benchmark claim about crawl speed. Actual performance depends on the target, network, response sizes, selectors, and crawl configuration. Begin with a small, narrow crawl, monitor its output and logs, and expand only when the target’s rules and your data requirements permit.

Reliability starts with selectors that are checked against real responses and output that is validated after export. Page markup can change, and a successful HTTP response does not guarantee that a selector extracted the intended data. For longer-running jobs, consider whether an item pipeline is needed to validate or deduplicate records before storing them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Review the target website’s terms and applicable requirements before crawling. Scrapy documentation describes how to build the workflow; it does not determine whether a particular collection is permitted.

Or skip the browser setup

If your goal is a screenshot rather than structured records, Scrapy is not the right tool: it crawls and extracts data, while a screenshot API returns an image or PDF. ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return PNG, JPEG, WebP, or PDF.

For example, this cURL request saves a WebP screenshot of the tutorial site:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://quotes.toscrape.com/ -o shot.webp

See the ScreenshotNeo API documentation for the request options and response details. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Sign up for ScreenshotNeo free to try 1,000 screenshots a month without a card.

Frequently Asked Questions

Can Scrapy crawl a site whose content is rendered by JavaScript?

This walkthrough’s documented workflow extracts from Scrapy response content. Whether that response contains the content you need depends on the site; inspect it in Scrapy shell before choosing an approach.

Can I use Scrapy to take screenshots?

Scrapy is for crawling and extracting structured data, not producing screenshots. For a screenshot or PDF, use a screenshot tool such as ScreenshotNeo.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.