Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Reliable web scraping is less about sending more requests and more about collecting the right pages, at a rate a site can handle, with rules and failures you can explain later. Start with an official API or feed when one provides the required fields and freshness. If page scraping is necessary, identify your crawler, read the exact site’s robots.txt, use conservative pacing, stop when access signals say to stop, and validate the resulting records.
The practices below combine the standards in IETF RFC 9309 and RFC 9110 with operational guidance from Amazon Web Services. They are implementation guidance, not legal advice; permission, terms and privacy obligations must be reviewed separately.
1. Check for an API or feed before scraping pages
Compare a documented API, RSS/Atom feed or data export with HTML scraping before writing a crawler. Record the trade-offs:
- Permission and terms: what the interface explicitly allows versus what page access permits.
- Fields and completeness: whether the API exposes every field you need or omits content rendered in a page.
- Freshness: update timing, pagination and historical coverage.
- Quotas and server impact: request limits, authentication and predictable load.
- Operational complexity: authentication and schema changes versus HTML selectors and rendering.
- Validation: how easily responses can be checked against a documented schema.
Choose the method that supplies the required data with the least unnecessary load. Neither method is universally superior.
#1 Best Overall
2. Read robots.txt for the exact origin and crawler
Fetch https://example.com/robots.txt (matching the protocol, host and port you will access). RFC 9309 defines groups by crawler identity, so give your software a stable product token and apply the matching group’s parseable rules. Google’s explanation of the specification is also useful when interpreting groups and patterns: robots.txt specification.
Cache the file for the duration of a crawl, record when it was retrieved, and recheck it before a later run. Treat syntax you cannot parse conservatively rather than guessing that a path is allowed.
3. Treat robots.txt as guidance, not authorization
RFC 9309 says the rules “are not a form of access authorization.” A permitted path is not a grant of credentials or a waiver of restrictions. Review the site’s terms, authentication requirements, technical controls and applicable privacy obligations independently. Also remember that robots.txt itself is public and can reveal sensitive-looking paths; it is not a security boundary.
4. Identify your crawler with a clear User-Agent
RFC 9110 §10.1.5 states that “A user agent SHOULD send a User-Agent header field in each request unless specifically configured not to do so.” Use a product name, version and contact URL or email, for example:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →User-Agent: CatalogResearchBot/1.2 (+https://example.org/bot-info)
Do not impersonate a browser or another crawler. RFC 9110 cautions against needless detail, which can increase fingerprinting and latency; identify enough for an operator to understand the traffic.
5. Start with a conservative per-host rate
Rate limits are site-specific. Begin slowly, measure responses and reduce load when latency or errors rise. AWS gives illustrative examples—not universal safe thresholds—of one request every 10–15 seconds for small or medium sites, and one to two requests per second for larger sites or sites that explicitly permit crawling. Respect any published limit that is lower.
Use a per-origin scheduler, not a single global sleep. Add jitter so a batch does not create a synchronized burst, and cap concurrency. Keep a request log containing URL, timestamp, status, elapsed time and response size.
6. Use HTTP errors as feedback
Handle 429 without escalating
A 429 means “Too Many Requests.” Pause the affected host, honor a server-provided Retry-After value when present, and resume only at a lower rate. Do not multiply retries across workers.
Recommended Free Tools
Stop on persistent 403 responses
A continuing 403 (“Forbidden”) is an access restriction, not an invitation to rotate identities or increase traffic. AWS recommends considering a stop when 403 responses continue. Save the failed URLs and seek permission or an approved interface.
Bound every retry
Retry only transient failures (for example, selected 5xx responses and network timeouts), with a small maximum and exponential backoff. Make failures visible instead of looping indefinitely.
7. Focus discovery with sitemaps
Use the site’s sitemap or sitemap index to find canonical URLs and avoid crawling navigation permutations, search results and duplicate parameters. AWS recommends sitemaps for identifying important pages and reducing unnecessary discovery. Parse each sitemap’s last-modified value when available, but treat it as a hint and verify that the page still returns the data you need.
8. Crawl in small, restartable batches
Partition a URL set by host, sitemap section or stable ranges. AWS recommends smaller batches to distribute load and reduce timeout and resource problems. Persist a queue and checkpoint after each successful record so a process restart does not repeat the entire crawl.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
- Write discovered URLs to durable storage with a status of
pending. - Claim a bounded batch and mark each URL
in_progress. - Store response metadata and parsed output separately from the queue.
- Mark successes
complete; schedule only bounded retries for failures. - Stop the batch when rate limits, repeated 403s or an error budget is reached.
9. Make extraction deterministic and observable
Prefer stable semantic selectors and explicit pagination rules over brittle positional selectors. Record the selector or parser version with every dataset. Save the retrieval timestamp, source URL, HTTP status, content type and a hash of the relevant response where storage and terms permit. Keep raw responses only when your retention and privacy policies allow it.
For JavaScript-rendered pages, wait for a specific selector or a documented readiness condition rather than an arbitrary long sleep. A browser is not a workaround for access controls; it still must follow the site’s rules and rate limits.
10. Validate data quality before publishing or modeling
A successful HTTP response does not prove a successful extraction. Run checks appropriate to your schema:
- Required fields are present and have the expected type.
- Primary keys are unique, and duplicate URLs are explained.
- Pagination reaches its intended end without silently truncating.
- Numeric ranges, encodings and date formats are plausible.
- Collected timestamps are consistent with the source and the run time.
- Record counts are compared with the expected sitemap or batch size.
- Parser failures and empty pages are quarantined for inspection.
Do not invent a universal pass threshold. Set thresholds from the source and your use case, alert on deviations, and retain a sample for manual review.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors11. Recheck assumptions when a site changes
Pages, robots rules, content structures and delivery behavior change. Version selectors and crawl configuration, monitor empty-field rates and status-code distributions, and record each collection date. Before relying on a new run, compare a sample with the prior schema and re-read robots.txt. If a layout change breaks parsing, pause the affected job instead of publishing partial data that looks complete.
A minimal respectful Python crawler
This example demonstrates identity, robots checking, pacing, bounded retries and visible errors. It is intentionally conservative; adapt parsing and permission review to the target site.
import time, random
from urllib.parse import urlparse
import requests
from urllib.robotparser import RobotFileParser
UA = "ExampleResearchBot/1.0 (+https://example.org/bot-info)"
def allowed(url):
p = urlparse(url)
robots = RobotFileParser(f"{p.scheme}://{p.netloc}/robots.txt")
robots.read()
return robots.can_fetch(UA, url)
def fetch(url, attempts=3):
if not allowed(url):
raise RuntimeError(f"Disallowed by robots.txt: {url}")
for attempt in range(attempts):
try:
r = requests.get(url, headers={"User-Agent": UA}, timeout=30)
print(url, r.status_code, r.elapsed.total_seconds())
if r.status_code == 429:
wait = int(r.headers.get("Retry-After", "60"))
time.sleep(wait)
continue
if r.status_code == 403:
raise RuntimeError("403 received; stop and review permission")
if 500 <= r.status_code < 600:
time.sleep((2 ** attempt) + random.random())
continue
r.raise_for_status()
return r.text
except requests.RequestException:
if attempt == attempts - 1:
raise
time.sleep((2 ** attempt) + random.random())
raise RuntimeError("retry limit reached")
html = fetch("https://example.com/page")
# Parse only the fields your documented schema requires.
Or skip the browser setup
For pages where you need a rendered screenshot rather than parsed fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookies/consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page and element captures, lazy-image loading, device and viewport settings, dark mode, retina scale, custom CSS or JavaScript, waits, request blocking, headers and cookies, geolocation, timezone, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks and bulk capture up to 100 URLs per call. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.
See the ScreenshotNeo documentation for parameters and authentication. Example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is included on every plan. Create a free ScreenshotNeo account.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting checklist
Robots rules appear to conflict
Confirm protocol, host, port and User-Agent group. Re-fetch the file, parse only rules you can understand, and ask the site owner when intent is unclear.
Responses are mostly 429
Stop the queue, honor Retry-After, lower per-host concurrency and increase the interval. Check that multiple workers are not sharing an uncoordinated limiter.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Responses are mostly 403
Stop retries. Review permission and terms, identify your crawler, and use an official API or request authorization.
Best Value
Records are empty but status is 200
The page may require JavaScript, show a consent wall or have changed markup. Inspect the saved response, verify a readiness selector when rendering, and update a versioned parser. Do not treat an empty parse as success.
The crawl times out
Split the batch, use explicit connect and read timeouts, checkpoint progress and retry only transient failures. A smaller, restartable job is safer than a large unbounded process.
Cost, performance and reliability decisions
More concurrency can shorten elapsed time while increasing server impact, throttling and retry work. Optimize total useful records per request, not requests per second. Sitemaps, caching of unchanged inputs, deduplication and small batches usually improve both cost and reliability. Cloud functions can suit short-lived event-driven tasks, but ordinary scraping does not require cloud infrastructure; choose execution based on run duration, scheduling, storage and observability needs.
Frequently Asked Questions
Does a disallow rule make scraping illegal?
No. RFC 9309 describes robots.txt as crawler guidance, not access authorization. Permission, terms, technical controls and privacy obligations require separate review.
What should I do when a site has no robots.txt?
Do not interpret absence as unlimited permission. Identify your crawler, review terms and access controls, choose a conservative rate, and contact the operator when permission is uncertain.
How many retries should a scraper use?
There is no universal number. Use a small bounded limit, back off for transient failures, pause on 429 and stop on persistent 403 responses.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

