Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhen a scraper fails, first identify whether the problem is the request, the page’s rendering, an access restriction, or your extraction code. Inspect the browser’s network activity, prefer an approved data endpoint when one exists, and use browser automation only when the data depends on rendered page behavior. Crawl politely, validate what you collect, and stop rather than trying to defeat CAPTCHAs or other access controls.
Diagnose the failure before changing tools
A scraper can return a successful HTTP response and still fail to collect useful data. The page might deliver its content through JavaScript, the site might return a challenge instead of the expected page, or a selector might no longer match the HTML. Start by recording what actually happened on each request.
Log enough to distinguish the causes
- Record the requested URL, timestamp, response status, final URL after redirects, content type, and elapsed time.
- Save a small, access-controlled sample of the response body or a redacted excerpt. Check whether it is the expected page, an error, a consent screen, or a challenge.
- Record extraction outcomes as well as request outcomes: expected fields found, fields missing, duplicate records, and parser or selector errors.
- Do not log passwords, session cookies, authorization headers, or personal data unnecessarily. Restrict access to logs that could contain them.
These signals separate transport and access problems from parsing problems. If the response contains the expected content but your output is empty, inspect your parser. If the response is a challenge page or an unexpected redirect, changing a CSS selector will not fix it.
Why does a scraper get 403 errors, CAPTCHAs, or challenge pages?
A 403 means the server refused the request; it does not, by itself, identify why. Sites can apply layered controls, including web application firewall rules, IP restrictions, JavaScript checks, CAPTCHAs, authentication requirements, and geographic rules. Repeated requests or requests outside an allowed route may also trigger restrictions.
#1 Best Overall
Respond without trying to bypass the restriction
- Check that the URL is intended for public or otherwise authorized access, and review the site’s terms and any API documentation.
- Reduce your request rate and concurrency, stop retrying a persistent denial, and verify that your crawler is not repeatedly requesting the same pages.
- Look for an approved API, export, feed, or permissioned data route. Ask the site owner for access if the intended content requires it.
- If a CAPTCHA, WAF challenge, authentication wall, or other access-control signal remains, stop that collection path. Do not try to defeat the control with stealth techniques or credential workarounds.
A retry is appropriate for a temporary network failure or server error when access is authorized; it is not a way to turn a deliberate denial into permission. Treat the returned content and status as an operational signal, not as a challenge to work around.
How do you scrape JavaScript-rendered pages?
First determine whether the browser is fetching the data from a separate request. Scrapy’s guidance notes that data visible in a browser may be absent from the HTML downloaded by a normal request. Open the browser’s network panel, reload the page, and inspect requests that return JSON or other structured data. If the site provides an intended API or endpoint and your use is permitted, requesting that data directly is usually simpler than rendering the whole page.
When direct HTTP is enough
Use a direct request when the required content is in the response body or in an authorized data endpoint. It generally avoids the overhead of a browser, is easier to run at scale, and makes it simpler to inspect status codes and response data. Keep to the route’s documented rules, authentication requirements, and rate limits.
When a browser is needed
Use a browser automation framework such as Playwright when the required data appears only after browser-side behavior, such as client-side rendering or user interaction. Wait for a meaningful page condition rather than an arbitrary short pause, and extract only the fields you need. Browser automation consumes more resources and adds moving parts: browser versions, loading behavior, and page changes can all affect reliability.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Do not treat browser automation as a way around a challenge or login boundary. If the page requires an authorization step, use only access you are entitled to use, and do not defeat a site’s protective measures.
Example: inspect a page with Playwright
This minimal Python example opens an authorized page, waits for a selector, and prints its text. Replace the URL and selector with ones appropriate to a page you are permitted to access. Install Playwright and its Chromium browser before running it.
from playwright.sync_api import sync_playwright
URL = "https://example.com/"
SELECTOR = "h1"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
response = page.goto(URL, wait_until="domcontentloaded", timeout=30000)
if response is None:
raise RuntimeError("Navigation did not return a response")
print("status:", response.status)
page.locator(SELECTOR).wait_for(timeout=10000)
print(page.locator(SELECTOR).inner_text())
browser.close()
This example intentionally does not include stealth settings, CAPTCHA handling, or login automation. If navigation times out, inspect whether the site is slow, whether the selector exists, and whether the request is permitted before changing the wait condition.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo is a website screenshot API and MCP server; it is an alternative to maintaining browser capture infrastructure, not a general-purpose scraping API. A single GET request returns a PNG, JPEG, WebP, or PDF. For a full list of parameters and behavior, see the ScreenshotNeo API documentation.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #3
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.
Python and Node.js alternatives
These examples use the same endpoint and parameters; keep your API key private and do not commit it to source control.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write("shot.webp", bytes);
The Node.js example uses Bun’s file-writing API; in a Node.js project, write the returned bytes with the built-in fs/promises module. Screenshot capture is useful for visual records and rendered-page output, but it does not replace a scraper that must normalize fields, follow pagination, or produce structured datasets.
How should you handle robots.txt, pacing, and retries?
Read the site’s robots.txt and terms before crawling. Robots.txt is a crawl instruction, not a universal legal ruling or a mechanism that hides a page from search results. Google Search Central explicitly cautions against using it to hide pages from search results. Scrapy can apply its robots middleware, but its optimization documentation says that Crawl-delay and Request-rate directives are not automatically enforced; translate those directives into your own download-delay and concurrency settings.
Set conservative crawl controls
- Use a clear user agent and identify your crawler where appropriate.
- Set an explicit per-domain concurrency limit and delay; do not assume a framework will infer the site’s preferred pace.
- Cache responses during development and avoid downloading the same unchanged page repeatedly.
- Deduplicate URLs and records before issuing requests or writing results.
- Retry only transient failures, with a bounded retry count and exponential backoff. Stop or pause on persistent denials, challenge pages, or signs of overload.
Example: a restrained Scrapy configuration
For a project where you have permission to crawl, these settings show where to apply pacing and robots handling. Adjust the delay and concurrency to the site’s documented requirements; the values below are an illustrative conservative starting point, not a universal safe rate.
# settings.py
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
AUTOTHROTTLE_MAX_DELAY = 60
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
RETRY_ENABLED = True
RETRY_TIMES = 2
HTTPCACHE_ENABLED = True
Scrapy’s robots setting helps honor robots.txt, but it does not settle whether collection is legally permitted. Nor does a delay setting override the site’s access rules. If the site specifies a slower pace, use that pace; if it denies access, stop and seek an approved route.
Which approach should you choose?
Choose the least complex method that can obtain the needed information through an authorized route. More sophisticated tooling does not make an unauthorized collection appropriate, and browser rendering is not automatically more reliable than a direct request.
| Approach | Best fit | Trade-off |
|---|---|---|
| Direct HTTP client | Static HTML or an approved structured endpoint | Low runtime overhead; cannot execute browser-side rendering |
| Scrapy | Permissioned crawls with many pages, URL discovery, and pipeline needs | Offers crawl controls and structured processing; pacing and parsing still need configuration and maintenance |
| Playwright or another browser automation framework | Data that depends on rendered DOM behavior or page interaction | More complete browser behavior, with greater runtime cost and more operational dependencies |
| Managed scraping service | Teams that prefer a service over running parts of the collection infrastructure themselves | Can reduce infrastructure work, but does not remove the need to check permission, data quality, service cost, and access rules |
Compare options on JavaScript completeness, throughput, latency, infrastructure cost, maintenance burden, observability, data-quality controls, authentication handling, and how clearly the workflow respects the target site’s rules. Managed-service names alone do not establish that a particular service is suitable or permitted for a particular target.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →How do you prevent silent data corruption when a site changes?
A parser can keep running while extracting empty, duplicated, or misclassified data. Treat validation as a separate stage after retrieval, and alert on changes instead of silently accepting them.
Best Value
Validate fields and records
- Normalize values consistently, including whitespace, dates, currencies, and empty fields.
- Check required fields, allowed value formats, record counts, and uniqueness constraints before publishing results.
- Track missing-field rates and selector failures by page type. A sudden rise can indicate a redesign, consent screen, or unexpected response.
- Version parser logic and keep representative fixtures so a code change can be checked against known response shapes.
- Compare current output to prior runs where appropriate, and alert on implausible changes rather than automatically overwriting trusted data.
Keep collection logs and raw response samples only as long as needed, and handle any personal or confidential material according to applicable privacy, security, and retention requirements.
Is web scraping legal?
There is no single answer that applies to every site, dataset, purpose, and jurisdiction. Cornell Law School’s Legal Information Institute summarizes screen scraping as technically legal in general while noting that circumventing typical protective measures can raise Computer Fraud and Abuse Act concerns. That generalization is not a blanket authorization or a substitute for legal advice about a specific collection.
Before collecting or republishing data, consider public availability, the site’s terms, authentication boundaries, copyright, privacy obligations, and the laws that apply to you and the site. Public visibility does not by itself settle every question. Robots.txt provides crawl guidance but should not be treated as a universal legal prohibition or permission. Do not rely on a ruling concerning one dispute or jurisdiction as proof that every scraping project is safe.
Recommended Free Tools
Common troubleshooting cases
- HTTP 403 or a challenge page: Confirm that the route is authorized, lower your request load, check for a documented API, and stop if access remains restricted. Do not attempt to bypass the protection.
- HTTP 200 but no extracted fields: Inspect the response body. If the data is absent, check for a permitted underlying data request or use a browser only if rendering is necessary and allowed. If the data is present, repair and test the parser.
- Navigation timeout in a browser: Check the response status, whether the host is reachable, and whether the selector exists. Wait for a relevant page condition rather than adding an unbounded delay.
- Sudden drop in record counts: Compare response samples and validation metrics with the previous run. Check for layout drift, pagination changes, consent screens, and parser failures before trusting the output.
- Repeated timeouts or server errors: Reduce concurrency, honor the site’s stated limits, and use bounded backoff for transient failures. Pause if failures continue rather than increasing load.
- Duplicate records or runaway request volume: Normalize and deduplicate discovered URLs, cache repeat responses, and review pagination and link-following rules.
A practical sequence for a maintainable scraper
- Define the exact fields you need, why you need them, and whether you may collect and use them.
- Check terms, robots.txt, authentication requirements, and any API or export route supplied by the site.
- Inspect a browser’s network activity if the data is missing from initial HTML; request an approved data endpoint directly when appropriate.
- Use a browser only when rendered DOM behavior is genuinely necessary, and never to defeat an access restriction.
- Apply explicit delays, concurrency limits, caching, deduplication, and bounded retries.
- Validate and monitor outputs so missing data or schema drift raises an alert instead of passing unnoticed.
This sequence keeps the choice of tool subordinate to the actual problem: obtain only the data you are authorized to access, at a responsible pace, and make failures visible before they corrupt downstream work.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

