Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable way to improve production scraping is to maximize valid, schema-checked records per attempted request—not raw request volume. Measure that yield per target, route and time window; find the concurrency and delay each site tolerates; classify failures before retrying; and separate transport health from extraction quality. There is no universal production success-rate benchmark, so your own validated baseline is the number to improve.

Define “success” before changing concurrency

A request that returns HTTP 200 is not necessarily a successful scrape. Define a denominator and numerator that match the business outcome. For example:

success rate = records passing parsing and schema validation ÷ attempted records

Attach the measurement to a named target, route and time window, such as “product-detail URLs on example.com during the last hour.” Track at least these separate measures:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Transport completion: responses received, connection failures and timeouts.
  • HTTP outcome: counts for 2xx, 3xx, 4xx and 5xx responses.
  • Retry behavior: attempts, retry reason and final outcome.
  • Parse validity: responses from which the expected fields were extracted.
  • Schema validity: records passing type, required-field and range checks.
  • Freshness: whether the record was collected within its business deadline.

Report these by host, route, status class, deployment version and time window. A rising completion rate with falling schema validity is a regression, not an improvement.

Start with the target’s permitted access path

Before tuning a crawler, check robots.txt, terms and published developer guidance. If the site offers an API, bulk export or search endpoint, prefer it: a documented interface is often faster for your system and less expensive for the site than page crawling.

When robots filtering is appropriate, enable Scrapy’s RobotsTxtMiddleware. Scrapy does not automatically enforce every Crawl-delay or Request-rate directive, so translate those instructions into explicit delays and concurrency limits. Document the permission decision and the owner responsible for reviewing changes.

Establish a per-host baseline

Run a representative, low-rate sample before increasing load. Keep target responses distinct from crawler-side exceptions: a 429 is a target response, while a DNS failure or socket timeout is a client-side failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimum baseline fields

  • Attempt timestamp, host, route and request identifier.
  • Status code and response size.
  • DNS, connect, TLS and total download latency when available.
  • Timeout or connection exception type.
  • Retry number and reason.
  • Parser result, missing-field details and schema result.
  • Target-side markers such as a known block page or CAPTCHA.

Use percentiles, not only averages. A p95 or p99 latency increase often appears before visible failure spikes. Preserve enough request metadata to replay a failed case safely without duplicating side effects.

Find the concurrency the site tolerates

There is no safe universal concurrency value. The limiting value is the one the target website tolerates. Increase per-host concurrency gradually while watching 429 responses, 503 responses, known ban pages, retries and latency. If those signals rise together, reduce load rather than adding workers.

A controlled ramp

  1. Begin with one or a few concurrent requests per host and a conservative delay.
  2. Hold each level long enough to observe a stable sample, rather than reacting to one response.
  3. Increase one variable at a time: concurrency, delay, request rate or rendering load.
  4. Stop the ramp when rate-limit responses, overload responses, ban pages or tail latency materially increase.
  5. Back off to the last stable level and record it as a provisional limit for that host and route.

Concurrency is not the same as request rate. Slow pages can create many in-flight requests at a modest request rate; JavaScript rendering can also consume substantially more target and local resources than a simple HTML fetch. Set separate limits for hosts and, where needed, expensive routes.

Use adaptive throttling, not a fixed “fast enough” delay

Scrapy AutoThrottle estimates delay from response latency and a target concurrency, averages the estimate with the previous delay, and respects configured minimum and maximum bounds. Non-200 response latency is not allowed to decrease the delay, which prevents an error response from making the crawler more aggressive. Treat target concurrency as an average goal, not an instantaneous cap.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Practical guardrails

  • Set a nonzero minimum delay even when pages are fast.
  • Set a maximum delay so a transient problem cannot stall the entire queue indefinitely.
  • Use a lower target concurrency for fragile hosts, authenticated sessions and browser-rendered routes.
  • Alert on sustained 429/503 rates and on latency percentiles, not just total throughput.
  • Keep an operator override for emergency backoff.

Adaptive throttling should respond to each target independently. A healthy host must not cause a struggling host to speed up.

Retry only failures that can recover

Retries are a bounded recovery mechanism, not a substitute for diagnosing access limits. Scrapy’s documented default is two retries after the initial download, and its retry middleware includes 429 and selected 5xx responses. Defaults are version-specific; verify the setting in the version you deploy.

Classify before retrying

  • Usually transient: connection resets, temporary DNS failures, gateway timeouts and selected 5xx responses.
  • Rate limiting: 429 responses should trigger backoff that honors any server guidance, not immediate parallel retries.
  • Likely permanent: malformed URLs, missing resources, rejected authentication and deterministic validation failures.
  • Access control: CAPTCHA, bot-check and ban pages require a policy decision and often a lower rate or documented access path, not more retries.

Use exponential backoff with jitter, a maximum attempt count and a queue-level retry budget. Record the final reason. Prevent retry amplification: if a site is overloaded, thousands of queued retries can turn a small incident into a sustained denial of service.

Investigate 4xx and 5xx responses differently

4xx responses

Check URL construction, stale links, authentication, authorization and the target’s access policy. A 404 may be a genuinely removed page; a 403 may be an intentional policy decision. Do not retry deterministic 4xx responses indefinitely. Sample response bodies carefully and redact credentials or personal data before storing them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5xx responses

A 5xx can originate at an intermediary such as a CDN or at the origin server. Compare the timing and distribution across routes, regions and request types, and check the target’s published status information where available. If only one route fails, inspect that route; if all routes fail simultaneously, treat it as a target or network incident. Keep backoff active while investigating.

Make extraction quality a first-class control loop

Validate records immediately after parsing. Required identifiers, URL shape, data types, date ranges and enumerated values should have explicit checks. Quarantine invalid records with a reason code instead of silently dropping them. Alert separately on:

  • Parser exceptions and selector miss rates.
  • Schema-invalid records by field.
  • Unexpected template or content-length changes.
  • Freshness lag and queue age.
  • Duplicate or conflicting records.

When a site redesign changes markup, transport metrics may remain healthy while valid yield collapses. A canary set of URLs with known expected fields can detect that failure before the full crawl completes.

Remove local bottlenecks and repeated work

During development, use HTTP caching to avoid downloading the same pages repeatedly. In production, distinguish target limits from scheduler, CPU, memory, storage, DNS and parser saturation. Narrow and reuse selectors; avoid expensive whole-document operations when a targeted selector is sufficient. Monitor queue depth and worker utilization alongside target responses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For idempotent jobs, deduplicate URLs before dispatch. Persist checkpoints so a worker restart resumes without replaying the entire crawl. Use bounded queues and circuit breakers per host so one failing target cannot consume all workers.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose an architecture by workload, not fashion

Option Best fit Questions to answer
Official API or export Documented, stable data access Quota, fields, freshness, pagination, authentication and change policy
Self-managed crawler Specialized routes, custom parsing or strict control Rendering, sessions, retries, observability, maintenance and target permissions
Managed extraction service Teams that need hosted browser/proxy operations Target coverage, JavaScript support, 429/5xx handling, replay, schema quality, latency and cost per valid record

Compare cost per valid record, not cost per request. Include engineering time, monitoring, incident response and data-quality remediation. A managed service is not automatically more reliable; evaluate it against a representative workload and an agreed success definition.

Or skip the browser setup

If your pipeline needs clean screenshots or rendered evidence rather than parsed HTML, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are free, with X-Page-Verdict and X-Billed headers explaining the result.

One GET request returns PNG, JPEG, WebP or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the full parameter set: full-page and selector capture, device presets, custom viewport and retina scale, dark mode, PDF controls, custom CSS/JavaScript, clicks, waits, blocking rules, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, 100-URL bulk capture, usage API and OpenAPI support. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. Every feature is included on every plan; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Production troubleshooting checklist

Symptom Likely cause Action
429 spike after a ramp Concurrency or rate exceeds the target’s tolerance Back off, honor server guidance, lower per-host limits and resume gradually
503 with rising latency Target or intermediary overload Pause aggressive retries, compare routes and check origin health
Many 404s Stale links or URL construction bug Inspect URL generation and stop retrying deterministic misses
200 responses but low valid yield Markup change, block page or parser regression Inspect samples, canary pages and schema-failure reasons
Timeouts only on rendered pages Heavy scripts or insufficient wait/resource limits Measure browser-stage timing, reduce concurrency and set explicit waits
Queue grows while targets look healthy CPU, memory, storage or parser bottleneck Profile workers and scale or optimize the local stage

FAQ

How much concurrency can my crawler safely use?

Only a measured, target-specific limit is defensible. Ramp gradually and hold the last level that does not increase rate limits, overload responses, ban pages or tail latency.

Is a 200 response a successful scrape?

No. Count a success only after the expected record parses and passes your schema and freshness rules.

Should I retry every 5xx response?

No. Retry selected transient failures within a bounded budget and back off when errors persist. Investigate whether the intermediary or origin is failing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.