Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliable web scraping is less about sending more requests and more about collecting the right pages, at a rate a site can handle, with rules and failures you can explain later. Start with an official API or feed when one provides the required fields and freshness. If page scraping is necessary, identify your crawler, read the exact site’s robots.txt, use conservative pacing, stop when access signals say to stop, and validate the resulting records.

The practices below combine the standards in IETF RFC 9309 and RFC 9110 with operational guidance from Amazon Web Services. They are implementation guidance, not legal advice; permission, terms and privacy obligations must be reviewed separately.

1. Check for an API or feed before scraping pages

Compare a documented API, RSS/Atom feed or data export with HTML scraping before writing a crawler. Record the trade-offs:

  • Permission and terms: what the interface explicitly allows versus what page access permits.
  • Fields and completeness: whether the API exposes every field you need or omits content rendered in a page.
  • Freshness: update timing, pagination and historical coverage.
  • Quotas and server impact: request limits, authentication and predictable load.
  • Operational complexity: authentication and schema changes versus HTML selectors and rendering.
  • Validation: how easily responses can be checked against a documented schema.

Choose the method that supplies the required data with the least unnecessary load. Neither method is universally superior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Read robots.txt for the exact origin and crawler

Fetch https://example.com/robots.txt (matching the protocol, host and port you will access). RFC 9309 defines groups by crawler identity, so give your software a stable product token and apply the matching group’s parseable rules. Google’s explanation of the specification is also useful when interpreting groups and patterns: robots.txt specification.

Cache the file for the duration of a crawl, record when it was retrieved, and recheck it before a later run. Treat syntax you cannot parse conservatively rather than guessing that a path is allowed.

3. Treat robots.txt as guidance, not authorization

RFC 9309 says the rules “are not a form of access authorization.” A permitted path is not a grant of credentials or a waiver of restrictions. Review the site’s terms, authentication requirements, technical controls and applicable privacy obligations independently. Also remember that robots.txt itself is public and can reveal sensitive-looking paths; it is not a security boundary.

4. Identify your crawler with a clear User-Agent

RFC 9110 §10.1.5 states that “A user agent SHOULD send a User-Agent header field in each request unless specifically configured not to do so.” Use a product name, version and contact URL or email, for example:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User-Agent: CatalogResearchBot/1.2 (+https://example.org/bot-info)

Do not impersonate a browser or another crawler. RFC 9110 cautions against needless detail, which can increase fingerprinting and latency; identify enough for an operator to understand the traffic.

5. Start with a conservative per-host rate

Rate limits are site-specific. Begin slowly, measure responses and reduce load when latency or errors rise. AWS gives illustrative examples—not universal safe thresholds—of one request every 10–15 seconds for small or medium sites, and one to two requests per second for larger sites or sites that explicitly permit crawling. Respect any published limit that is lower.

Use a per-origin scheduler, not a single global sleep. Add jitter so a batch does not create a synchronized burst, and cap concurrency. Keep a request log containing URL, timestamp, status, elapsed time and response size.

6. Use HTTP errors as feedback

Handle 429 without escalating

A 429 means “Too Many Requests.” Pause the affected host, honor a server-provided Retry-After value when present, and resume only at a lower rate. Do not multiply retries across workers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stop on persistent 403 responses

A continuing 403 (“Forbidden”) is an access restriction, not an invitation to rotate identities or increase traffic. AWS recommends considering a stop when 403 responses continue. Save the failed URLs and seek permission or an approved interface.

Bound every retry

Retry only transient failures (for example, selected 5xx responses and network timeouts), with a small maximum and exponential backoff. Make failures visible instead of looping indefinitely.

7. Focus discovery with sitemaps

Use the site’s sitemap or sitemap index to find canonical URLs and avoid crawling navigation permutations, search results and duplicate parameters. AWS recommends sitemaps for identifying important pages and reducing unnecessary discovery. Parse each sitemap’s last-modified value when available, but treat it as a hint and verify that the page still returns the data you need.

8. Crawl in small, restartable batches

Partition a URL set by host, sitemap section or stable ranges. AWS recommends smaller batches to distribute load and reduce timeout and resource problems. Persist a queue and checkpoint after each successful record so a process restart does not repeat the entire crawl.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Write discovered URLs to durable storage with a status of pending.
  2. Claim a bounded batch and mark each URL in_progress.
  3. Store response metadata and parsed output separately from the queue.
  4. Mark successes complete; schedule only bounded retries for failures.
  5. Stop the batch when rate limits, repeated 403s or an error budget is reached.

9. Make extraction deterministic and observable

Prefer stable semantic selectors and explicit pagination rules over brittle positional selectors. Record the selector or parser version with every dataset. Save the retrieval timestamp, source URL, HTTP status, content type and a hash of the relevant response where storage and terms permit. Keep raw responses only when your retention and privacy policies allow it.

For JavaScript-rendered pages, wait for a specific selector or a documented readiness condition rather than an arbitrary long sleep. A browser is not a workaround for access controls; it still must follow the site’s rules and rate limits.

10. Validate data quality before publishing or modeling

A successful HTTP response does not prove a successful extraction. Run checks appropriate to your schema:

  • Required fields are present and have the expected type.
  • Primary keys are unique, and duplicate URLs are explained.
  • Pagination reaches its intended end without silently truncating.
  • Numeric ranges, encodings and date formats are plausible.
  • Collected timestamps are consistent with the source and the run time.
  • Record counts are compared with the expected sitemap or batch size.
  • Parser failures and empty pages are quarantined for inspection.

Do not invent a universal pass threshold. Set thresholds from the source and your use case, alert on deviations, and retain a sample for manual review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

11. Recheck assumptions when a site changes

Pages, robots rules, content structures and delivery behavior change. Version selectors and crawl configuration, monitor empty-field rates and status-code distributions, and record each collection date. Before relying on a new run, compare a sample with the prior schema and re-read robots.txt. If a layout change breaks parsing, pause the affected job instead of publishing partial data that looks complete.

A minimal respectful Python crawler

This example demonstrates identity, robots checking, pacing, bounded retries and visible errors. It is intentionally conservative; adapt parsing and permission review to the target site.

import time, random
from urllib.parse import urlparse
import requests
from urllib.robotparser import RobotFileParser

UA = "ExampleResearchBot/1.0 (+https://example.org/bot-info)"

def allowed(url):
    p = urlparse(url)
    robots = RobotFileParser(f"{p.scheme}://{p.netloc}/robots.txt")
    robots.read()
    return robots.can_fetch(UA, url)

def fetch(url, attempts=3):
    if not allowed(url):
        raise RuntimeError(f"Disallowed by robots.txt: {url}")
    for attempt in range(attempts):
        try:
            r = requests.get(url, headers={"User-Agent": UA}, timeout=30)
            print(url, r.status_code, r.elapsed.total_seconds())
            if r.status_code == 429:
                wait = int(r.headers.get("Retry-After", "60"))
                time.sleep(wait)
                continue
            if r.status_code == 403:
                raise RuntimeError("403 received; stop and review permission")
            if 500 <= r.status_code < 600:
                time.sleep((2 ** attempt) + random.random())
                continue
            r.raise_for_status()
            return r.text
        except requests.RequestException:
            if attempt == attempts - 1:
                raise
            time.sleep((2 ** attempt) + random.random())
    raise RuntimeError("retry limit reached")

html = fetch("https://example.com/page")
# Parse only the fields your documented schema requires.

Or skip the browser setup

For pages where you need a rendered screenshot rather than parsed fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookies/consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.

One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page and element captures, lazy-image loading, device and viewport settings, dark mode, retina scale, custom CSS or JavaScript, waits, request blocking, headers and cookies, geolocation, timezone, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks and bulk capture up to 100 URLs per call. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo documentation for parameters and authentication. Example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is included on every plan. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting checklist

Robots rules appear to conflict

Confirm protocol, host, port and User-Agent group. Re-fetch the file, parse only rules you can understand, and ask the site owner when intent is unclear.

Responses are mostly 429

Stop the queue, honor Retry-After, lower per-host concurrency and increase the interval. Check that multiple workers are not sharing an uncoordinated limiter.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Responses are mostly 403

Stop retries. Review permission and terms, identify your crawler, and use an official API or request authorization.

Records are empty but status is 200

The page may require JavaScript, show a consent wall or have changed markup. Inspect the saved response, verify a readiness selector when rendering, and update a versioned parser. Do not treat an empty parse as success.

The crawl times out

Split the batch, use explicit connect and read timeouts, checkpoint progress and retry only transient failures. A smaller, restartable job is safer than a large unbounded process.

Cost, performance and reliability decisions

More concurrency can shorten elapsed time while increasing server impact, throttling and retry work. Optimize total useful records per request, not requests per second. Sitemaps, caching of unchanged inputs, deduplication and small batches usually improve both cost and reliability. Cloud functions can suit short-lived event-driven tasks, but ordinary scraping does not require cloud infrastructure; choose execution based on run duration, scheduling, storage and observability needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does a disallow rule make scraping illegal?

No. RFC 9309 describes robots.txt as crawler guidance, not access authorization. Permission, terms, technical controls and privacy obligations require separate review.

What should I do when a site has no robots.txt?

Do not interpret absence as unlimited permission. Identify your crawler, review terms and access controls, choose a conservative rate, and contact the operator when permission is uncertain.

How many retries should a scraper use?

There is no universal number. Use a small bounded limit, back off for transient failures, pause on 429 and stop on persistent 403 responses.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.