Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A “blocked” URL during capture can fail for very different reasons: your browser or network may be unable to connect, robots.txt may instruct a crawler not to fetch the path, or the site’s origin server, CDN, firewall, or WAF may reject the request. Start by recording the exact error and HTTP status, then test the same URL from another browser, device, and network. That evidence tells you which control layer to investigate instead of changing unrelated settings.

First, classify what “blocked” means

Do not treat every failed capture as a robots problem. Compare the scope, response, and control point:

What you observe Likely layer Useful next check
Only one browser fails Browser profile, extension, security software, or cache Private window, another browser, and extension/security-software checks
Every browser on one device fails System clock, DNS, local firewall, proxy, or endpoint security Resolve the hostname, verify time, and test another device
Devices on one network fail, other networks work Local DNS, router, ISP, or network filtering Use a separate network and compare DNS results
A crawler receives a policy message robots.txt Inspect the applicable user-agent and path rules
HTTP 403 or a branded challenge/block page Origin, CDN, firewall, or WAF Compare edge and origin logs and rule matches
HTTP 429 Rate limiting Check request volume, retry behavior, and rate-limit configuration
Timeout, DNS, or TLS error Availability, name resolution, certificate, or network path Test connectivity and server health before changing crawler policy

Google’s crawling guidance treats HTTP status, server availability, blocked resources, loading speed, and the rendered result as separate diagnostic evidence (Google Search Central). Record which category you have before making a fix.

1. Capture evidence before changing settings

  1. Copy the complete URL. Preserve the scheme, host, path, query string, and fragment (if relevant to your capture tool).
  2. Write down the time and frequency. Note the time zone, whether the failure is constant or intermittent, and whether only one path is affected.
  3. Save the visible error. Keep a screenshot or text of messages such as “DNS error,” “certificate error,” “403 Forbidden,” “429 Too Many Requests,” or a WAF challenge.
  4. Record the capture environment. Include browser or client name and version, device, operating system, network, proxy, user agent, and whether JavaScript was enabled.
  5. Capture the status and headers when available. A status code and response headers can distinguish a server response from a connection failure. If the tool exposes a rendered page, save that result too.

These details let a site owner correlate your request with edge and origin logs. They also prevent a common mistake: editing robots.txt when the real failure is DNS, TLS, a timeout, or a firewall rule.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
  • DUAL-BAND WIFI 6 ROUTER: Wi-Fi 6(802.11ax) technology achieves faster speeds, greater capacity and reduced network congestion compared to the previous gen. All WiFi routers require a separate modem. Dual-Band WiFi routers do not support the 6 GHz band.
  • AX1800: Enjoy smoother and more stable streaming, gaming, downloading with 1.8 Gbps total bandwidth (up to 1200 Mbps on 5 GHz and up to 574 Mbps on 2.4 GHz). Performance varies by conditions, distance to devices, and obstacles such as walls.
  • CONNECT MORE DEVICES: Wi-Fi 6 technology communicates more data to more devices simultaneously using revolutionary OFDMA technology
  • EXTENSIVE COVERAGE: Achieve the strong, reliable WiFi coverage with Archer AX1800 as it focuses signal strength to your devices far away using Beamforming technology, 4 high-gain antennas and an advanced front-end module (FEM) chipset
  • OUR CYBERSECURITY COMMITMENT: TP-Link is a signatory of the U.S. Cybersecurity and Infrastructure Security Agency’s (CISA) Secure-by-Design pledge. This device is designed, built, and maintained, with advanced security as a core requirement.

2. Test the browser, device, and network

Try another browser and a clean session

Open the URL in another current browser and, if possible, a private window. If only one browser fails, disable extensions that filter traffic, clear a corrupted site-data entry, and check local security software. Mozilla’s troubleshooting guidance also calls out incorrect system time, DNS failures, and security software as causes of pages that do not load (Mozilla Support).

Check the system clock and certificate path

An incorrect date or time can make a valid certificate appear expired or not-yet-valid. Set the device to automatic time synchronization, restart the browser, and retry. Do not suppress certificate warnings for a production capture; an apparently successful image could represent an unsafe or incomplete connection.

Check DNS and basic reachability

Confirm that the hostname resolves on the failing device and compare the result with another network. A DNS failure means the browser cannot find the server, so changing crawler rules will not help. If a proxy, VPN, corporate gateway, or endpoint firewall is present, test a permitted network without it and ask the administrator whether the destination is filtered.

Compare another device and network

If the URL works on a phone over cellular data but not on the office network, the scope points to local DNS or network policy. If it fails everywhere, continue with server-side checks. Keep the original environment details; changing networks without recording the difference removes useful evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Check robots.txt when you control the site

Understand what robots rules do

robots.txt is a crawler instruction file, not an access-control system. Google explains that a disallowed URL can still appear in search results because the URL may be discovered elsewhere even when its contents are not crawled (Google’s robots.txt guide). Use authentication, authorization, or another server-side control for private material; never rely on robots rules to protect secrets.

Rank #2
Sale
TP-Link BE6500 Dual-Band WiFi 7 Router (BE400)
  • 𝐅𝐮𝐭𝐮𝐫𝐞-𝐑𝐞𝐚𝐝𝐲 𝐖𝐢-𝐅𝐢 𝟕 - Designed with the latest Wi-Fi 7 technology, featuring Multi-Link Operation (MLO), Multi-RUs, and 4K-QAM. Achieve optimized performance on latest WiFi 7 laptops and devices, like the iPhone 16 Pro, and Samsung Galaxy S24 Ultra.
  • 𝟔-𝐒𝐭𝐫𝐞𝐚𝐦, 𝐃𝐮𝐚𝐥-𝐁𝐚𝐧𝐝 𝐖𝐢-𝐅𝐢 𝐰𝐢𝐭𝐡 𝟔.𝟓 𝐆𝐛𝐩𝐬 𝐓𝐨𝐭𝐚𝐥 𝐁𝐚𝐧𝐝𝐰𝐢𝐝𝐭𝐡 - Achieve full speeds of up to 5764 Mbps on the 5GHz band and 688 Mbps on the 2.4 GHz band with 6 streams. Enjoy seamless 4K/8K streaming, AR/VR gaming, and incredibly fast downloads/uploads.
  • 𝐖𝐢𝐝𝐞 𝐂𝐨𝐯𝐞𝐫𝐚𝐠𝐞 𝐰𝐢𝐭𝐡 𝐒𝐭𝐫𝐨𝐧𝐠 𝐂𝐨𝐧𝐧𝐞𝐜𝐭𝐢𝐨𝐧 - Get up to 2,400 sq. ft. max coverage for up to 90 devices at a time. 6x high performance antennas and Beamforming technology, ensures reliable connections for remote workers, gamers, students, and more.
  • 𝐔𝐥𝐭𝐫𝐚-𝐅𝐚𝐬𝐭 𝟐.𝟓 𝐆𝐛𝐩𝐬 𝐖𝐢𝐫𝐞𝐝 𝐏𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞 - 1x 2.5 Gbps WAN/LAN port, 1x 2.5 Gbps LAN port and 3x 1 Gbps LAN ports offer high-speed data transmissions.³ Integrate with a multi-gig modem for gigplus internet.
  • 𝐎𝐮𝐫 𝐂𝐲𝐛𝐞𝐫𝐬𝐞𝐜𝐮𝐫𝐢𝐭𝐲 𝐂𝐨𝐦𝐦𝐢𝐭𝐦𝐞𝐧𝐭 - TP-Link is a signatory of the U.S. Cybersecurity and Infrastructure Security Agency’s (CISA) Secure-by-Design pledge. This device is designed, built, and maintained, with advanced security as a core requirement.

The Internet Engineering Task Force’s Robots Exclusion Protocol states: “If a crawler successfully downloads a robots.txt file, the crawler MUST follow the parseable rules” (RFC 9309). That normative requirement applies to compliant crawlers, not to an ordinary human browser, and it does not force a server to permit access.

Inspect the exact user-agent and path

Fetch the site’s /robots.txt and identify the user-agent group used by your capture client. Check whether a broader Disallow rule matches the requested path, whether an allow rule is actually effective for that crawler, and whether the file is valid and parseable. Test the exact URL after correcting an unintended rule; do not assume that a rule for one host or path applies to another.

Use Search Console for Google’s report

For Google-specific diagnosis, open Search Console’s URL Inspection tool, enter the URL, and review the crawl result. Google documents this workflow for finding and correcting an unintended robots block (Unblock a page blocked by robots.txt). Retest after publishing the corrected file and allowing time for the crawler to fetch it again.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Investigate the origin server, CDN, firewall, and WAF

Separate edge responses from origin responses

A CDN or WAF can reject a request before it reaches your application. Compare the response body, headers, request ID, and timestamp at the edge with the origin server’s access log. A branded challenge page, a rule identifier, or a provider header suggests an edge decision; an application-generated 403 points farther downstream. Confirm the difference with the site operator rather than trying to evade the control.

Interpret common status codes

  • 403 Forbidden: permission, IP reputation, authentication, bot policy, or another rule may have denied the request. Inspect matched firewall/WAF rules and the request’s user agent, headers, cookies, and source address.
  • 429 Too Many Requests: the service is rate-limiting requests. Cloudflare documents 429 responses in its troubleshooting material (Cloudflare Error Pages troubleshooting). Reduce concurrency, honor any Retry-After value, and review limits rather than repeatedly retrying.
  • 5xx: an origin or edge availability problem may be preventing a complete response. Check deployment health, upstream dependencies, and timeout settings.
  • 3xx loops or unexpected redirects: verify canonical host, HTTPS redirects, authentication redirects, and whether the capture client follows redirects.

Review logs and rule matches

Inspect the relevant CDN, firewall, WAF, reverse-proxy, and origin logs for the recorded time. Look for rate limits, geo or ASN rules, bot scores, missing authentication, blocked methods, request-size limits, and resource-specific policies. Cloudflare’s crawl-error guidance recommends comparing the failure with availability and crawl observations to locate the failing layer (Cloudflare: Troubleshoot crawl errors).

Rank #3
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
  • Dual band router upgrades to 1200 Mbps high speed internet (300mbps for 2.4GHz plus 900Mbps for 5GHz), reducing buffering and ideal for 4K stream
  • Full Gigabit Ports - Gigabit Router with 4 Gigabit LAN ports, ideal for any internet plan and allow you to directly connect your wired devices
  • Boosted Coverage - Four external antennas equipped with Beamforming technology extend and concentrate the Wi-Fi signals
  • MU-MIMO technology - (5GHz band) allows high speeds for multiple devices simultaneously
  • Access Point Mode - Supports AP Mode to transform your wired connection into wireless network, an ideal wireless router for home

Check rendering and dependent resources

A document can return 200 while its capture is blank because scripts, styles, images, fonts, or API calls are blocked or too slow. Review browser-console and network evidence when available. Google lists blocked resources, slow loading, large resources, and rendering failures among crawl problems (Google Search Central). Allow only the resources required for the page and verify that the final rendered output contains the expected content.

5. Choose the fix based on who controls the failing layer

If you control the machine or network

  • Correct the system clock and DNS configuration.
  • Test without a conflicting extension, proxy, VPN, or endpoint rule, subject to your organization’s policy.
  • Ask the network administrator to review filtering logs for the URL and time.

If you control the website

  • Correct an unintended robots rule and validate the user-agent/path match.
  • Review origin, CDN, firewall, and WAF logs for the request.
  • Adjust rate limits, authentication, or bot policy only when the capture is authorized.
  • Retest from the same client and network, then from an independent one.

If you do not control the website

Send the owner or administrator the complete URL, timestamp and time zone, status or visible error, client and network, and a screenshot of the response. Ask for access or explicit capture permission where appropriate. Do not bypass a login, CAPTCHA, paywall, robots policy, firewall, or other access control; an access-denied page may be intentional.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup: ScreenshotNeo

ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Use the API documentation at screenshotneo.com/docs/ for authentication and options. A basic cURL request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For difficult pages, ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper size/margins/orientation/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay, or network idle, ad/tracker/request/resource blocking, custom headers, cookies, user agent and Authorization, timezone and geolocation, transparent backgrounds, image resizing, selectable cache TTL, signed public-image links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, usage API, OpenAPI, and compatibility with parameter names used by other screenshot APIs.

Plan Allowance and price
Free 1,000 shots/month, no card
Starter $5 for 3,000 shots
Growth $15 for 15,000 shots
Pro $39 for 60,000 shots
Scale $99 for 250,000 shots
Business $249 for 1,000,000 shots

Yearly billing gives two months free, and every feature is available on every plan. If you need a repeatable capture without local browser, consent, or popup setup, sign up for the free plan with 1,000 screenshots a month and no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
  • Dual-band Wi-Fi with 5 GHz speeds up to 867 Mbps and 2.4 GHz speeds up to 300 Mbps, delivering 1200 Mbps of total bandwidth¹. Dual-band routers do not support 6 GHz. Performance varies by conditions, distance to devices, and obstacles such as walls.
  • Covers up to 1,000 sq. ft. with four external antennas for stable wireless connections and optimal coverage.
  • Supports IGMP Proxy/Snooping, Bridge and Tag VLAN to optimize IPTV streaming
  • Access Point Mode - Supports AP Mode to transform your wired connection into wireless network, an ideal wireless router for home
  • Advanced Security with WPA3 - The latest Wi-Fi security protocol, WPA3, brings new capabilities to improve cybersecurity in personal networks

Troubleshooting checklist

  • Have you saved the exact URL, timestamp, environment, visible error, status, and rendered result?
  • Does another browser, device, or network change the outcome?
  • Is the failure DNS, TLS, timeout, HTTP status, robots policy, or rendering?
  • If it is robots-related, did you check the matching user-agent and path?
  • If it is 403 or 429, did you inspect edge and origin logs and rule matches?
  • Are blocked scripts or resources producing a blank or incomplete capture?
  • Do you have permission to capture the site, and have you avoided bypassing controls?

FAQ

Can a robots.txt rule block my normal browser?

No. Robots Exclusion Protocol rules direct compliant crawlers. Browser access can still be denied separately by authentication, server permissions, a CDN, firewall, or WAF.

Why does the URL open normally but the capture is blank?

The capture client may be missing JavaScript execution, waiting too briefly, or unable to load a required script, stylesheet, image, font, or API response. Inspect the rendered output and dependent requests rather than relying only on the document’s status code.

Should I keep retrying after a 429?

No. Reduce request rate, follow the server’s retry guidance, and ask the site operator to review authorized limits. Repeated immediate retries can prolong or intensify rate limiting.

Frequently Asked Questions

Can a robots.txt rule block my normal browser?

No. Robots rules direct compliant crawlers; separate authentication, server, CDN, firewall, or WAF controls can deny browser access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does the URL open normally but the capture is blank?

The capture client may not execute scripts, wait long enough, or load a required dependent resource. Inspect rendered output and network requests.

Should I keep retrying after a 429?

No. Reduce request rate, honor retry guidance, and have the site operator review authorized limits.

Quick Recap

SaleBestseller No. 1
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
VPN SERVER: Archer AX21 Supports both Open VPN Server and PPTP VPN Server
$59.98
Bestseller No. 3
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
MU-MIMO technology - (5GHz band) allows high speeds for multiple devices simultaneously
$44.99
Bestseller No. 4
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
Supports IGMP Proxy/Snooping, Bridge and Tag VLAN to optimize IPTV streaming
$34.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.