Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy is a Python framework for crawling websites and extracting structured data. The beginner workflow is: install Scrapy in a Python 3.10-or-newer virtual environment, create a project and spider, parse each response with CSS or XPath selectors, yield dictionaries or item objects, and export the results as JSON, CSV, XML, or another supported feed. Add an item pipeline when you need cleaning, validation, duplicate removal, or custom storage.

This guide builds that workflow from an empty folder, then covers selectors, pagination, feeds, pipelines, settings, responsible crawling, debugging, and a browser-free screenshot option.

What Scrapy does

Scrapy manages the repetitive parts of a crawler: scheduling requests, downloading responses, invoking callbacks, extracting fields, following more links, processing items, and exporting data. A spider contains the site-specific logic; selectors read HTML; item pipelines process records; feed exports serialize them; and settings connect and tune those components.

The framework is designed for crawling and structured extraction, with documented uses including data mining, monitoring, and automated testing. Unlike a one-off request-and-parse script, a Scrapy project gives each concern a defined place and can handle many requests through its scheduler and concurrency controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The request-to-item loop

  1. The spider creates an initial request, commonly through start_urls or a start_requests() method.
  2. Scrapy downloads the response and calls the spider’s callback, usually parse().
  3. The callback uses response.css() or response.xpath() to select values.
  4. It yields dictionaries or item objects, and can yield follow-up requests for detail pages or pagination.
  5. Scrapy sends each yielded item through enabled pipelines and then to any feed export requested at the command line or in settings.

Install Scrapy in an isolated project

Current Scrapy 2.19 documentation requires Python 3.10 or newer. A project-specific virtual environment keeps Scrapy and its dependencies separate from system packages and from unrelated projects.

  1. Verify Python:
    python --version

    Use the Python executable that reports 3.10 or a later version.

  2. Create and enter a working directory:
    mkdir scrapy-learning
    cd scrapy-learning
  3. Create a virtual environment:
    python -m venv .venv
  4. Activate it using the command appropriate to your shell, then install Scrapy from PyPI:
    python -m pip install --upgrade pip
    python -m pip install Scrapy
  5. Check the installation:
    scrapy version
  6. Create a project:
    scrapy startproject tutorial
    cd tutorial

The same installation guide also documents conda-forge installation. Whichever installer you use, keep the environment active whenever you run the project.

Write a first spider

Generate a spider inside the project with:

scrapy genspider example example.com

Replace the generated file with this complete example. It reads the page title and headings, follows a conventional a.next link when one exists, and safely handles missing fields.

import scrapy


class ExampleSpider(scrapy.Spider):
    name = "example"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    def parse(self, response):
        title = response.css("title::text").get()
        headings = response.css("h1::text").getall()

        yield {
            "url": response.url,
            "title": title.strip() if title else None,
            "headings": [value.strip() for value in headings],
        }

        next_href = response.css("a.next::attr(href)").get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

Save the file under tutorial/spiders/example.py. The allowed_domains value prevents accidental off-site crawling when links are followed. The example uses response.follow(), which resolves a relative URL against the current response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run the spider and export a feed

From the directory containing scrapy.cfg, run:

scrapy crawl example -O items.json

This starts the spider and writes yielded items as JSON. Feed exports also support JSON Lines, CSV, and XML, along with other documented storage and serialization choices. For example:

scrapy crawl example -O items.jsonl
scrapy crawl example -O items.csv
scrapy crawl example -O items.xml

Choose the format that matches the next system in your workflow. JSON is convenient for nested values, CSV is convenient for spreadsheet-like rows, and JSON Lines is useful when processing records one at a time.

Extract values with CSS and XPath

Scrapy selectors support both CSS and XPath. Choose based on the actual structure of the page and whichever expression your team can maintain; the documentation does not establish that one is universally more robust.

CSS selectors

# One element or None
name = response.css("article h1::text").get()

# Every matching value
prices = response.css(".price::text").getall()

# An attribute
href = response.css("a.product::attr(href)").get()

.get() returns the first match or None when there is no match. .getall() returns every match, or an empty list when none exists. Strip whitespace and test for None before calling string methods.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XPath selectors

name = response.xpath("//article//h1/text()").get()
prices = response.xpath("//span[contains(@class, 'price')]/text()").getall()
href = response.xpath("//a[contains(@class, 'product')]/@href").get()

Use XPath when you need relationships such as “the text in the heading next to this label.” CSS is often easier to read for class and element matches. Inspect the downloaded HTML, not only what a browser displays after JavaScript changes the page.

Normalize and validate extracted data

Selectors return strings as they appear in the response. Normalize whitespace, convert numeric text deliberately, and reject incomplete records in a pipeline rather than silently treating a missing selector as valid data.

raw_price = response.css(".price::text").get()
price = raw_price.strip() if raw_price else None

if price is not None:
    price = price.replace("$", "").replace(",", "")

Follow pagination and detail links

A callback can yield both an item and new requests. For a list page, extract each card, follow its detail URL, and schedule the next list page:

def parse(self, response):
    for card in response.css(".card"):
        detail_href = card.css("a::attr(href)").get()
        if detail_href:
            yield response.follow(detail_href, callback=self.parse_detail)

    next_href = response.css("a.next::attr(href)").get()
    if next_href:
        yield response.follow(next_href, callback=self.parse)

def parse_detail(self, response):
    yield {
        "url": response.url,
        "name": response.css("h1::text").get(),
        "description": response.css(".description::text").get(),
    }

Guard every optional link. A missing “next” link should end the crawl, not create a request for None. Also decide how to identify duplicates when a site exposes the same content through multiple URLs.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose between feed exports and item pipelines

Feed exports are the simpler route when you only need supported serialization and storage. Pipelines are the right place for item-level work such as cleaning, validation, duplicate removal, database writes, or a custom destination.

Need Prefer Reason
Write extracted records to JSON, JSON Lines, CSV, or XML Feed export Minimal configuration and no custom item-processing code.
Normalize fields or reject invalid records Pipeline Processing runs once per yielded item.
Remove duplicates during a run Pipeline A component can keep a set or apply a project-specific key.
Save to a custom database or service Pipeline Storage logic stays separate from page parsing.

A small cleaning pipeline

class CleanExamplePipeline:
    def process_item(self, item, spider):
        for key, value in item.items():
            if isinstance(value, str):
                item[key] = " ".join(value.split())
        return item

Enable it in the project’s settings.py:

ITEM_PIPELINES = {
    "tutorial.pipelines.CleanExamplePipeline": 300,
}

Pipeline components run in numeric priority order from lower values to higher values. Put validation before storage when invalid items must never reach the destination, and make duplicate handling explicit so that a restart does not produce surprising results.

Settings that affect a crawl

Settings configure components and runtime behavior. Keep site-specific values in the project rather than scattering them through spider code. Scrapy exposes concurrency and crawl-rate controls, but there is no universally safe request rate: the appropriate pace depends on the target site, its current instructions, your authorization, and applicable requirements.

  • Concurrency: control how many requests can be in flight so your crawl does not overwhelm a server or your own machine.
  • Delay and throttling: use a deliberate delay or adaptive crawl-rate controls when the target’s capacity and rules call for them.
  • Robots and terms: check the particular site’s current robots instructions, terms, privacy expectations, copyright constraints, and jurisdiction before collecting data. Scrapy settings cannot grant permission.
  • Request identity: use honest, stable headers and identify your crawler where appropriate; do not use settings to evade access controls.

Start with a small, observable crawl. Log the request count, response status, item count, and error count, then increase scope only when the output and behavior are correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Debug a Scrapy project

“scrapy” or “python” is not found

The virtual environment is probably inactive, or Scrapy was installed into a different Python interpreter. Activate the environment and run python -m pip show Scrapy. If necessary, install with python -m pip install Scrapy while that environment is active.

The spider is not listed

Run scrapy list from the directory containing scrapy.cfg. Confirm the file is under the project’s spiders package, the class inherits from scrapy.Spider, and its name is unique.

Selectors return None or an empty list

Print or save the response body and compare it with the selector. Check for a changed class name, a wrong response type, whitespace-only text nodes, or content that is not present in the downloaded HTML. Use .get() only when one value is expected and .getall() for repeated values.

The server returns 403 or another error

Stop and verify that your crawl is authorized and follows the site’s current instructions. A different status can indicate a blocked or incomplete request, but it is not a reason to evade a bot check or access control. Inspect logs, reduce scope, and correct request configuration only within the site’s rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Items are yielded but the output file is empty

Check that the callback actually reaches a yield, that the spider exits without an exception, and that you ran the command from the project directory. Run with a feed format explicitly, such as scrapy crawl example -O debug.json, and inspect the crawl log for item counts.

The pipeline does not run

Confirm the fully qualified class path in ITEM_PIPELINES, verify indentation and commas in settings.py, and remember that a pipeline receives only items that the spider yields. Its priority number controls order, not whether it is enabled.

JavaScript-generated content is missing

Scrapy parses the response it receives. If the desired data is absent from that HTML, inspect whether the page obtains it through a separate request and whether that endpoint can be accessed lawfully. Scrapy’s documentation also covers dynamic content and related advanced topics; do not assume that adding a CSS selector can extract content that was never in the response.

Operational and cost considerations

Scrapy itself is a framework you install in your Python environment; the retrieved documentation does not establish a universal hosting provider, deployment price, speed benchmark, or safe crawl-rate percentage. Plan capacity around your page count, response sizes, concurrency, storage, retries, and the target’s limits. Test with a narrow URL set before a broad crawl, keep raw evidence when you need to audit extraction, and make a restart strategy so a transient failure does not duplicate or lose records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For production work, separate parsing code from settings and pipelines, pin the environment used for a run, record the spider version with exported data, and monitor logs for status-code and selector changes. Scrapy’s documentation index includes debugging, optimization, deployment, security, contracts, and dynamic-content subjects for the next level of project-specific work.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your immediate need is a clean visual capture rather than a structured crawl, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture, then removes more than 60 known consent platforms along with newsletter popups and chat widgets. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the shot was billed.

Use the documented options when you need full-page or element captures, lazy images, a device preset or custom viewport, dark mode, retina scale, PDF paper settings, custom CSS or JavaScript, clicks, waits, blocked resources, headers, cookies, authorization, timezone, geolocation, transparency, resizing, a chosen cache TTL, signed image links, asynchronous webhooks, bulk capture, usage data, or an OpenAPI specification.

One-call cURL example

See the ScreenshotNeo documentation for the complete parameter reference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to get started.

FAQ

Can one spider yield both items and new requests?

Yes. A callback can yield dictionaries or item objects and follow-up requests in the same method; Scrapy schedules each yielded request and sends each item through the configured processing path.

Should I define Scrapy Item classes before learning selectors?

No. Dictionaries are sufficient for a first spider. Introduce item classes when field definitions, validation, or shared structure makes them useful.

Where should site-specific business rules live?

Keep page navigation and extraction in the spider, item-level cleaning and validation in pipelines, and reusable runtime behavior in project settings. This separation makes a selector change less likely to affect storage code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can one spider yield both items and new requests?

Yes. A callback can yield dictionaries or item objects and follow-up requests in the same method; Scrapy schedules each yielded request and sends each item through the configured processing path.

Should I define Scrapy Item classes before learning selectors?

No. Dictionaries are sufficient for a first spider. Introduce item classes when field definitions, validation, or shared structure makes them useful.

Where should site-specific business rules live?

Keep page navigation and extraction in the spider, item-level cleaning and validation in pipelines, and reusable runtime behavior in project settings. This separation makes a selector change less likely to affect storage code.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.