To keep a long web-scraping job within a memory budget, find which part is growing before changing limits. In Scrapy, compare scheduler queues, active responses and process memory at several points in the crawl. Then constrain the specific source—queued requests, oversized responses or slow processing—while checking that the change does not drop valid pages or overload the target site.
Why long scraping jobs run out of memory
Memory use is not just the response body currently being downloaded. A crawler may hold scheduled requests, active responses, parsed selector trees, items awaiting pipelines, and objects retained by callbacks or custom components. In browser-based jobs, pages and request history add their own state.
The key distinction is whether memory growth tracks a bounded workload or continues without an apparent queue explanation. Scrapy’s optimization guide describes an ever-growing scheduler memory queue as a direct cause of exhaustion; if memory rises without that queue growing, retained objects or a leak deserve attention. Scrapy 2.19.0 optimization documentation (live master documentation; defaults and implementation details may change).
Measure the crawl before changing settings
Take readings at multiple stages rather than diagnosing from a single peak. Scrapy’s engine status exposes useful signals:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
len(engine.downloader.active): active downloader requests.len(engine.scheduler.mqs): requests in in-memory scheduler queues.engine.scraper.slot.active_size: response data currently active in the scraper.engine.scraper.slot.needs_backout(): whether the scraper signals that incoming work should back off.
Use these as observations during a running crawl, not as values to optimize in isolation. Record process memory alongside them and note how each changes as URLs are discovered, downloaded, parsed and passed through pipelines. Engine internals can vary by Scrapy version, so confirm the relevant status interface for the version you run.
Read the pattern, not one number
- Scheduler queue climbs continuously: the spider may be discovering or scheduling requests faster than the downloader consumes them. Request objects and their metadata are part of the memory cost.
- Active response size approaches its limit: parsing callbacks or item pipelines may be falling behind the incoming response flow. Reducing backlog or response volume is usually a better first move than increasing concurrency.
- Memory rises while both queues remain steady: inspect references retained by callbacks, middleware, pipelines, extensions and request metadata. Custom components can keep objects alive longer than intended.
- Memory is stable but disk or CPU is saturated: this is a different bottleneck. Compare response volume with available bandwidth and inspect persistent job state, caches and media pipelines.
Control queued requests without losing crawl coverage
Producing requests early helps keep downloaders busy, but requests wait somewhere until they can be processed. Scrapy summarizes the tradeoff: “Each of these trades memory for speed: a request produced before the downloader can take it waits in the scheduler, or on disk if you set JOBDIR.” Scrapy optimization guide.
Reduce the in-memory backlog
- Limit how far callbacks or start-request generation run ahead of available downloader and parsing capacity.
- Review request priorities if low-value work is accumulating ahead of useful requests.
- Avoid eagerly materializing very large collections of start URLs when they can be yielded progressively.
These measures may reduce downloader utilization if request production becomes too conservative. Watch both the queue and crawl throughput; the goal is a manageable queue, not an empty one at all times.
Use disk-backed job state when it fits
Scrapy’s JOBDIR can store scheduled requests on disk, reducing the need to keep all queued state in memory. This exchanges RAM pressure for disk use and I/O, and persistent job state adds operational considerations such as disk capacity and job-directory management. It is useful when scheduled work is the source of memory growth, not as a remedy for a leak in callbacks or a large active response backlog.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Bound response and parsing memory
A response body can understate its eventual memory cost. Scrapy selectors build an in-memory tree for the full response, which can consume several times the body size. Multiple in-flight responses and parsed objects can compound that cost.
Set a response-size cap deliberately
Scrapy 2.19.0’s live security documentation describes DOWNLOAD_MAXSIZE as defaulting to 1 GiB per response. That is a framework default, not a recommended cap for every workload, and may change in later versions. Set the value based on observed legitimate response sizes and the memory budget. A smaller cap protects against unexpectedly huge bodies, but responses above it can be dropped, including valid pages. See Scrapy’s security documentation.
Before applying a cap broadly, inspect the largest responses your crawl actually needs: long reports, product catalogs, embedded data or unusually large pages may be legitimate. Treat oversize failures as coverage decisions and monitor them.
Constrain active scraper data
SCRAPER_SLOT_MAX_ACTIVE_SIZE is a soft limit for response data under processing. Lowering it may constrain active processing and memory pressure, but can reduce throughput. If active size is persistently near the limit, investigate callback and item-pipeline work as well; throttling incoming work does not make slow processing faster.
Rank #3
For media pipelines, Scrapy’s optimization guide also identifies MEDIA_CACHE_SIZE as a relevant control. Media processing and whole-response handling can create separate memory and disk demands, so identify which resource is growing before tuning it.
Tune concurrency for processing capacity and site tolerance
Global concurrency, per-domain concurrency and download delay shape how much work is in flight and how quickly new requests are issued. Higher concurrency can improve utilization when resources are available, but it also means more work being processed or waiting. It is not a free speed multiplier.
Increase or decrease concurrency gradually while observing memory, response latency and target responses. Rising HTTP 429 or 503 responses, more retries, or worsening latency are signs that the target may be throttling the crawl or that the chosen rate is counterproductive. Respect the site’s documented access rules and keep the crawl rate within its tolerance. There is no universal concurrency number or memory cap that fits every job.
Choose the right scaling approach
| Approach | What it can bound or improve | Tradeoff to check |
|---|---|---|
| Limit request production or adjust priorities | In-memory scheduled work | Too little work ahead can leave download capacity idle. |
JOBDIR disk-backed scheduling |
RAM used by scheduled requests | Uses disk and I/O; manage persistent job state and available space. |
DOWNLOAD_MAXSIZE |
Largest accepted response body | Can exclude legitimate responses above the chosen cap. |
| Lower active scraper-data limit | Response data being processed concurrently | May constrain throughput; address processing backlog too. |
| Lower concurrency or add delay | In-flight work and pressure on the target | Can extend crawl duration; excessive rates can trigger throttling or errors. |
| Separate worker processes | CPU use across more than one core | Adds operational complexity and does not fix unbounded growth within each job. |
Scrapy’s current optimization guidance describes the framework as a single process. CPU-bound Python code competes for the GIL even if moved to a thread; threads can keep CPU-heavy work from blocking an event loop, but they do not add CPU capacity for CPU-bound Python. Separate processes are the documented way to use more than one CPU core. First establish that CPU is the bottleneck: adding processes does not fix a scheduler backlog or memory leak, and each process has its own resource needs. See Scrapy’s optimization guidance.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesPlaywright: account for browser-held state
Browser automation has state beyond downloaded page bodies. In the current Playwright Python Page API, page.requests() provides up to the 100 most recent requests; older request objects may be collected to avoid unbounded memory growth. The API was introduced in Playwright v1.56, and the live documentation may change by release. If request data matters to your job, retrieve it promptly rather than assuming the full history remains available. Playwright also documents page.request_gc(). Manage page and browser-context lifecycles deliberately, and inspect application references that retain pages or request objects. Playwright Python Page API documentation.
Troubleshoot common memory-growth symptoms
Memory rises as the scheduler queue grows
Likely cause: request discovery is outpacing downloads. Try: pace request production, review priorities, progressively yield large start URL sets, or use JOBDIR if disk-backed scheduling suits the job. Verify that the queue trend changes and that coverage remains complete.
Memory rises with a large active-response backlog
Likely cause: callbacks or pipelines cannot process responses as quickly as they arrive. Try: reduce response volume or incoming concurrency, inspect expensive parsing and item handling, and consider the soft active-size limit. Do not raise concurrency until processing can keep up.
Large pages fail after adding a size limit
Likely cause: the cap is below a legitimate response size. Try: inspect failed URLs and response sizes, then choose a cap that reflects required content and available memory. A cap is a completeness tradeoff, not a neutral optimization.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- Used Book in Good Condition
Memory grows but queues do not
Likely cause: retained objects or a leak in custom code. Try: audit references held by callbacks, middleware, extensions, pipelines and request metadata; check whether parsed items, responses or browser objects remain attached to long-lived structures.
Retries, 429s, 503s or latency increase after speeding up
Likely cause: the target or the crawler’s processing path cannot sustain the rate. Try: lower concurrency or increase delay, then watch response behavior and memory again. Faster request issuance can produce a slower or less complete crawl when it triggers throttling.
Adding threads does not improve CPU-bound parsing
Likely cause: CPU-bound Python work is constrained by the GIL. Try: profile first; use separate processes if more CPU cores are needed, and separately diagnose each process’s memory behavior.
Or skip the browser setup
If the job is simply to capture pages as images or PDFs, a screenshot API can avoid maintaining a browser capture stack. ScreenshotNeo is a website screenshot API and MCP server for developers: it removes cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages and failed loads are not billed; and its MCP server lets AI agents use screenshot tools. One GET request returns a PNG, JPEG, WebP or PDF. See the ScreenshotNeo site and API documentation.
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
One thousand screenshots a month are free with no card; paid plans start at $5 for 3,000 shots. Start at ScreenshotNeo’s free sign-up.
Frequently Asked Questions
Should I add RAM before optimizing a scraper?
Only after measurements show the workload is bounded and genuinely needs more capacity. More RAM can postpone exhaustion from an unbounded queue or retained objects without fixing the cause.
How often should I sample memory during a crawl?
Sample at multiple meaningful stages—early crawl, sustained throughput and periods of growth—so you can compare process memory with scheduler and active-response trends. The sources do not prescribe one universal sampling interval.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

