Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If a crawl starts slowing, stalling, or returning worse data around request 10,000, that number is usually a symptom—not a universal limit. The failure is most often caused by target-site throttling, a Scrapy concurrency or delay setting, request-generation logic, callback and pipeline backpressure, CPU or memory pressure, or retries consuming your available capacity. Measure those signals before raising concurrency.

What “fails after 10,000 requests” actually means

Different failures require different fixes. Record the exact symptom: throughput falling, the process exiting, memory exhaustion, an empty result set, incomplete pagination, HTTP errors, a ban page, stale records, or malformed items. Official Scrapy guidance does not establish 10,000 requests as a general breakpoint; its recommendations are operational signals rather than population statistics. Treat your observed count as the point where this workload exposed a constraint.

For every response, log at least the URL or domain, status code, response latency, retry number, response size, and whether an item was produced. At the crawler level, sample active downloader requests, scheduler queue sizes, callback and pipeline duration, CPU, and memory. A time series makes it possible to distinguish a remote limit from a local bottleneck.

Failure mode 1: the target site is throttling or blocking you

Rising 429 Too Many Requests or 503 Service Unavailable responses, ban-page bodies, more retries, and worsening download latency as concurrency increases are strong signs that the target is receiving more traffic than it currently tolerates. Scrapy’s optimization documentation recommends treating these signals as a reason to slow down and re-check the site’s rules, not as proof that you need still more parallelism (Scrapy Optimization).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to check before changing settings

  • Read robots.txt, the site’s terms, and any published API or rate-limit documentation.
  • Compare status and latency by domain; one problematic host should not dictate settings for every host.
  • Inspect response bodies for a block page that uses HTTP 200 instead of an error status.
  • Check whether an authorized API, bulk export, or documented search endpoint supplies the same data with fewer page requests.

Do not make IP or proxy rotation your default remedy. It can increase load, violate terms, and hide the underlying pacing problem. Increase request rate only gradually while the site’s responses remain acceptable.

Failure mode 2: concurrency and delay settings are the ceiling

Scrapy applies several independent controls. CONCURRENT_REQUESTS limits simultaneous downloads globally; CONCURRENT_REQUESTS_PER_DOMAIN limits requests to one domain; and DOWNLOAD_DELAY imposes a minimum interval between requests to a domain. A large scheduler queue with underused downloader slots often means the per-domain cap, delay, or AutoThrottle is limiting dispatch rather than the global setting.

How to interpret the queue

  • Queued requests, low downloader activity: inspect per-domain concurrency, delay, and AutoThrottle.
  • Busy downloader, rising latency and errors: the target may be throttling you; reduce pressure.
  • Both scheduler and downloader nearly empty: request production, not downloading, is limiting throughput.

AutoThrottle adjusts per-site delays from observed latency toward a configured average concurrency. That target is a goal, not a hard cap; ordinary concurrency and delay settings still apply. Its design avoids reducing delay merely because fast non-200 responses arrive, since such responses can indicate an excessive request rate (Scrapy AutoThrottle). Enable it when adaptive pacing fits your crawl, then watch the resulting latency and status distribution rather than assuming it is a universal fix.

Translate robots directives into settings

Scrapy does not automatically apply Crawl-delay or Request-rate directives from robots.txt to your settings. When those directives are present and applicable, reflect them explicitly in download delay and concurrency values, and document the decision in your deployment configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Failure mode 3: the spider cannot produce requests fast enough

A crawl that follows pagination strictly one page at a time cannot use more concurrency than its discovery pattern allows. If the downloader and scheduler are mostly idle, profile the callback that creates requests. Look for serial browser-like workflows, unnecessary parsing, database calls, or a next-page link that is not being extracted after a template change.

Make discovery safely more independent

  • Seed independent category, sitemap, or detail URLs where the site permits it.
  • Separate discovery from detail fetching so known detail URLs can be scheduled together.
  • Preserve ordering only where the application requires it; do not serialize an entire crawl for a presentation-only order.
  • Keep the target’s documented rate and robots directives in force while increasing independence.

Measure requests generated per callback and callbacks per second. Raising downloader concurrency cannot help if no new requests are being yielded.

Failure mode 4: callbacks, pipelines, CPU, or memory are saturated

Responses can arrive faster than selectors, transformations, deduplication, or item pipelines can process them. Scrapy then applies backpressure. A scheduler queue that grows without settling means discovery is outpacing downloading; a growing in-memory workload can eventually exhaust the process.

Separate local pressure from remote pressure

  • CPU near one core while network is underused: profile selectors, parsing, compression, serialization, and synchronous library calls.
  • Memory climbs throughout the crawl: look for retained responses, unbounded lists, duplicate caches, or a pipeline leak.
  • Callbacks or pipelines take longer as volume grows: inspect database indexes, batch sizes, and external service latency.
  • Network errors rise only when concurrency rises: compare with target-side status and latency before changing local code.

Scrapy runs in one process; apart from DNS and work explicitly moved to a thread, most work runs in one thread. One CPU core can therefore become the ceiling (Scrapy Optimization). Profile first. More downloader concurrency can worsen memory pressure by delivering even more responses to a CPU-bound callback.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Failure mode 5: retries amplify the outage

Retries are useful for transient failures, but repeated timeout retries against a slow or failing site keep crawler capacity tied up. Scrapy’s broad-crawl documentation notes that this can substantially slow a broad crawl and prevent capacity from being reused for other domains (Scrapy Broad Crawls 2.7.1).

Set retry behavior by failure and crawl shape

  • Retry only statuses and exceptions that are plausibly transient.
  • Use a bounded retry count and record exhausted URLs for later review.
  • Do not treat ban pages or deterministic 4xx responses as network glitches.
  • For broad crawls, prevent one failing domain from consuming all global slots.

Track original attempts and retry attempts separately. A nominal 10,000-URL crawl can create many more requests when each timeout is retried.

A practical diagnostic sequence

  1. Define the failure. Write down whether the problem is speed, process exit, memory, empty output, partial pagination, HTTP errors, ban content, or data quality.
  2. Establish a baseline. Capture per-status counts, latency percentiles, retry counts, response sizes, items produced, and active requests for a representative interval.
  3. Compare queues and downloader activity. Queued work with unused slots points to a cap or delay; busy slots with rising errors points to target pressure; empty queues point to request production.
  4. Inspect processing. Measure callback and pipeline time, CPU utilization, memory trend, database wait, and scheduler growth.
  5. Change one control. Adjust only one variable—such as per-domain concurrency or delay—in a small step, then observe status, latency, and throughput.
  6. Back off on deterioration. If 429/503 counts, ban pages, latency, or memory rise, restore the previous value and investigate instead of stacking changes.
  7. Check documented access paths. An official API, bulk export, or search endpoint may be faster for your crawler and cheaper for the target than page crawling. Confirm availability, terms, authentication, and rate limits with the site.

Choosing the remedy by signal

Observed signal Likely location First response
429/503, ban pages, latency rises with concurrency Target-site limit Reduce pressure, honor published rules, and look for an authorized access method.
Queue grows while downloader is below global capacity Per-domain cap, delay, or AutoThrottle Inspect those settings and domain-specific statistics.
Downloader and scheduler nearly empty Request-generation logic Profile callbacks and safely widen independent discovery.
CPU saturated, network idle Parsing or pipeline work Profile selectors and synchronous operations; optimize processing.
Memory grows with queue or response volume Backpressure or leak Bound buffers, inspect retained objects, and reduce in-flight work.
Many timeout retries Retry amplification Bound retries, classify failures, and isolate unhealthy domains.

Performance, reliability, and cost trade-offs

Throughput is useful only when the resulting records are complete and valid. A faster crawl that triggers blocks, drops pages, or fills memory is less reliable than a slower crawl with stable status rates. Keep raw response metadata, retry outcomes, and failed URLs so you can audit omissions. Apply backpressure deliberately: bounded queues and batch writes protect the process, while unbounded buffering merely postpones failure.

For multi-domain crawls, fairness matters. A failing host should not monopolize global concurrency through long timeouts and retries. Conversely, do not lower every domain’s rate because one site is slow. Partition metrics and settings by domain when your architecture allows it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup: ScreenshotNeo for page images

If the job is collecting rendered page images rather than extracting structured records, a screenshot endpoint can remove an entire browser-management layer. ScreenshotNeo is a website screenshot API and MCP server. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status.

One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets or custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

cURL

See the ScreenshotNeo documentation for the complete option list. A basic capture is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo is also an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common errors and fixes

“Concurrency is high, but throughput fell”

Check 429/503 rates, ban content, and latency. Reduce concurrency or add delay, then verify recovery. Do not assume the host can absorb more traffic.

“The queue grows until the process dies”

Measure callback time, CPU, and memory. Bound in-flight work, remove retained objects, optimize parsing, and reduce concurrency if responses are arriving faster than they can be processed.

“No requests are active”

Inspect the callback that yields pagination and detail requests, link extraction after template changes, and any database or API call blocking the reactor.

“Retries dominate the request count”

Classify timeout and status failures, cap retries, and quarantine unhealthy domains. Review failed URLs separately rather than endlessly retrying during the main crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The crawl returns HTML but the data is empty”

Check for a block page with status 200, JavaScript-rendered content, consent overlays, and selector changes. Save representative responses for inspection before altering concurrency.

FAQ

Is 10,000 requests a hard limit?

No. The count is an observation from a particular workload; official Scrapy documentation does not define it as a universal failure threshold.

Should I simply increase CONCURRENT_REQUESTS?

Only after metrics show downloader capacity is the bottleneck and the target’s errors and latency remain acceptable. Otherwise, increasing it can worsen blocking or memory pressure.

Does AutoThrottle replace robots.txt compliance?

No. It adapts delay from latency; you must still interpret applicable robots directives and the site’s published terms yourself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When is an API better than crawling pages?

When the site offers an authorized API, export, or documented search endpoint containing the needed data at a stated rate. Check its availability, terms, and limits first.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.