Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a Scrapy item pipeline when scraped items need application-specific processing—such as cleaning, validation, deduplication, or database writes. Use feed exports when you simply need Scrapy to serialize items and deliver them to a file or supported storage service. You can combine the two: pipelines process each item in sequence, and items that remain in the chain can also be exported.

Choose between a pipeline and a feed export

The choice depends on what must happen to an item after a spider yields it. An item pipeline gives your project code control over each item; a feed export is a built-in route for serializing items to a destination with little or no custom code. These are not mutually exclusive: a successful pipeline can return an item so that later stages and configured feed exports can receive it.

Need Better fit Why
Clean values, validate required fields, transform data, or filter duplicates Item pipeline Pipeline components inspect and modify each item, and can discard it with DropItem. Scrapy item pipelines
Write items to a database with application-specific behavior Item pipeline A component can use a database client to persist each item and apply project-specific write logic. Scrapy item pipelines
Produce JSON, JSON Lines, CSV, or XML with minimal custom logic Feed export Scrapy serializes items and writes them to a configured feed destination. Scrapy feed exports
Deliver files to local storage, FTP/FTPS, S3, GCS, or standard output Feed export The feed URI selects a documented storage backend; some cloud destinations may need optional extras. Scrapy feed exports
Run indexed queries or controlled updates as soon as a crawl completes Database pipeline A database is more suitable than a flat export when downstream applications need queryable records or controlled persistence.
Retain a crawl artifact for inspection or downstream batch processing Feed export A file or object-storage feed can be inspected or handed to another process; retention and overwrite behavior should be configured deliberately. Scrapy feed exports

How an item pipeline processes data

A spider yields an item, then Scrapy passes it through enabled pipeline components in sequence. Each component implements process_item(self, item, spider). It must return the item to continue processing or raise DropItem to stop that item from proceeding. Components only run when enabled through ITEM_PIPELINES, usually in the project’s settings.py. Lower numeric priorities run earlier in the chain. Scrapy item pipelines

Keep each component focused

A useful chain separates jobs: normalize fields first, validate next, filter duplicates if needed, then persist. Because the order matters, choose priorities that put prerequisite transformations before checks or writes that depend on them. If a pipeline raises DropItem, later pipeline components do not receive that item.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Implement a database writer

Here is a compact MongoDB example following Scrapy’s documented pipeline pattern. Install and configure the MongoDB driver required by your project, then adapt the collection and database settings to your environment. Scrapy’s MongoDB pipeline example

import pymongo


class MongoPipeline:
    @classmethod
    def from_crawler(cls, crawler):
        return cls(
            mongo_uri=crawler.settings.get("MONGO_URI"),
            mongo_db=crawler.settings.get("MONGO_DATABASE", "scrapy"),
        )

    def __init__(self, mongo_uri, mongo_db):
        self.mongo_uri = mongo_uri
        self.mongo_db = mongo_db

    def open_spider(self, spider):
        self.client = pymongo.MongoClient(self.mongo_uri)
        self.db = self.client[self.mongo_db]
        self.collection = self.db["items"]

    def close_spider(self, spider):
        self.client.close()

    def process_item(self, item, spider):
        self.collection.insert_one(dict(item))
        return item

The return at the end allows later pipeline stages to run and leaves the item available for other configured item consumers, including feed exports. Raise DropItem only when the item should be discarded rather than passed onward.

Enable the component and set connection values

Register the pipeline class in the project’s settings. The numeric priority controls where it runs relative to other enabled components.

# settings.py
ITEM_PIPELINES = {
    "myproject.pipelines.MongoPipeline": 300,
}

MONGO_URI = "mongodb://localhost:27017"
MONGO_DATABASE = "scrapy"

Keep credentials out of source control; load them from environment-specific configuration in a real deployment. Decide how the chosen database client should handle retries, indexes, connection lifecycle, and idempotency. Those details depend on the database and driver rather than on Scrapy’s pipeline interface.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clean, validate, and filter before persistence

Normalize values

Use a component to make representations consistent before later checks or writes. For example, trim a title and normalize whitespace so that duplicate comparisons do not treat cosmetic spacing as a different value.

class NormalizePipeline:
    def process_item(self, item, spider):
        title = item.get("title")
        if title:
            item["title"] = " ".join(title.split())
        return item

Reject invalid records

Use DropItem when a record cannot be used downstream. Include a useful reason in the exception so crawl logs can distinguish missing data from other failures.

from scrapy.exceptions import DropItem


class RequiredFieldsPipeline:
    def process_item(self, item, spider):
        if not item.get("url") or not item.get("title"):
            raise DropItem("missing required url or title")
        return item

Plan duplicate handling around your identity key

A duplicate filter needs a stable definition of “same record,” such as a canonical URL or a source-specific identifier. A process-local set can filter repeats during one crawl, but it does not by itself provide durable deduplication across runs or coordinate multiple workers. For durable behavior, use a database uniqueness constraint or an upsert strategy appropriate to your database, and decide what should happen when a record changes.

Export JSON, CSV, or another feed format

Configure feed exports with the FEEDS setting. Each entry maps a destination URI to options such as format, encoding, selected fields, overwrite behavior, empty-feed behavior, batching, and post-processing. Scrapy’s built-in feed formats include JSON, JSON Lines, CSV, and XML; custom formats can be added through FEED_EXPORTERS. Scrapy feed exports

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write a local JSON Lines file

JSON Lines stores one JSON object per line, which is useful when records should be consumed incrementally or appended by downstream tools. For a straightforward local feed:

# settings.py
FEEDS = {
    "output/items.jsonl": {
        "format": "jsonlines",
        "encoding": "utf8",
        "overwrite": True,
    },
}

Scrapy’s format name for JSON Lines is jsonlines. Set overwrite deliberately: replacing a prior crawl file may be correct for a repeatable snapshot, but wrong if the file is intended as an archive.

Write CSV with selected fields

CSV is convenient for spreadsheet review and simple tabular interchange. Specify the columns when stable ordering matters; ensure the fields represent flat values that make sense in a row-oriented format.

# settings.py
FEEDS = {
    "output/items.csv": {
        "format": "csv",
        "encoding": "utf8",
        "fields": ["url", "title", "price"],
        "overwrite": True,
    },
}

JSON and XML can represent nested structures more naturally than CSV. Check how your chosen format handles missing or nested fields before treating an export as a long-term interchange contract.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use dynamic paths and multiple feeds

Feed URI placeholders such as %(time)s and %(name)s can make a path specific to the crawl time or spider. You can configure more than one feed when the same crawl needs, for example, both a local review file and a cloud delivery destination. Confirm the expanded paths and retention policy before scheduling repeated jobs. Scrapy feed exports

Choose where feed files go

Scrapy documents local filesystem, FTP, FTPS, Amazon S3, Google Cloud Storage, and standard output as feed storage backends. S3 and GCS may require optional extras, so install the dependencies appropriate to the Scrapy version and deployment you use. Scrapy feed exports

Destination Good fit Watch for
Local filesystem Development, debugging, or a job whose output is consumed on the same host Disk capacity, path permissions, and whether each run overwrites prior output
FTP or FTPS Delivery to an existing file-transfer workflow Credentials, network access, and destination-side retention rules
Amazon S3 or Google Cloud Storage Durable feed delivery and downstream batch or data-lake workflows Optional dependencies, access configuration, object naming, and lifecycle/retention policy
Standard output Piping a feed into another process or capturing it with a job runner Keep logs separate from feed output so they do not corrupt the stream

Storage destinations are delivery targets, not substitutes for database semantics. If downstream consumers need indexed queries, updates, or record-level constraints, store records in a database or have a later ingestion step load the feed into one.

Combine database writes and exports safely

A pipeline can validate and persist an item, then return it so it continues through the chain. This lets a crawl both update a database and retain a feed artifact. Decide what an export should contain: items rejected by a pipeline will not proceed, while returned items can continue to later consumers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Put normalization and validation before persistence when the database should receive only cleaned, acceptable values.
  • Return successfully processed items when they should continue to later components or be included in a feed.
  • Use DropItem for intentional exclusion, not as a substitute for handling database or network errors.
  • Define a recovery approach for partial failures, such as a database write succeeding while a later export fails; Scrapy’s pipeline interface does not make writes across separate destinations one atomic transaction.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational choices: reliability, performance, and cost

Database writes

A per-item synchronous write is easy to understand but can limit crawl throughput when the database responds slowly. For larger workloads, consider batching or asynchronous driver patterns supported by the client and Scrapy integration you choose. Bound retries and make writes idempotent where possible, so a retry does not create unintended duplicate records. Monitor connection failures and ensure resources close when a spider ends.

Feed exports

Exports reduce custom persistence code, but large feeds still consume storage and transfer capacity. Choose file naming and overwrite behavior to match whether each run is a replaceable snapshot or a retained historical artifact. For cloud storage, use the documented backend configuration and validate permissions before relying on a scheduled crawl. Scrapy documents the available backends and feed options, but does not prescribe your organization’s retention, storage cost, or recovery policy. Scrapy feed exports

Troubleshooting Scrapy data handling

Symptom Likely cause What to check
Pipeline code never runs The component is not enabled or the class path is wrong Confirm its dotted import path and priority under ITEM_PIPELINES in the active project settings.
Items disappear from later stages or the feed A component raises DropItem, or returns something other than the item Inspect each component’s return path and log the reason for intentional drops.
Records reach the database before normalization Pipeline priorities put the writer before the normalization component Lower numeric priorities run earlier; order components accordingly.
No feed file appears The feed setting is absent, malformed, or points to an unavailable destination Check the FEEDS URI and options, destination permissions, and whether the run produced items.
Cloud feed setup fails Optional storage dependencies or credentials are missing Install the appropriate optional extras and validate the configured access for the target backend.
A previous export was replaced The feed overwrite policy or destination behavior allowed replacement Review overwrite, URI naming, and backend-specific behavior; use unique run paths when keeping history.
CSV columns are missing or inconsistent Exported items have differing fields or the intended field selection was not configured Set the fields order explicitly and verify each yielded item contains the expected data.

Or skip the browser setup

If the data source needs browser rendering before Scrapy can process its pages, ScreenshotNeo is an alternative website screenshot API and MCP server. A GET request with a URL returns an image or PDF; its documented options include custom CSS and JavaScript, cookies, headers, waits, and full-page capture. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; those cleanup steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server exposes screenshot and page-information tools to AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo or sign up for the free plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.