Use a staged crawler: discover approved URLs from robots.txt and sitemaps, request them with explicit limits, retain the final URL and response metadata, then run a task-specific content check. The templates below cover a small Python script, a Scrapy sitemap crawler, validation, reporting, JavaScript-heavy pages, and common failures. They are starting points, not permission to bypass authentication, access controls, or a site’s terms.
What a resource-checking scraper should do
A useful checker separates four jobs that are often mixed together:
- Input: an approved starting host or URL list, resource types or path patterns, concurrency and delay limits, and an output format.
- Discovery: read the host’s root
robots.txtand sitemap references, then expand sitemap indexes and URL sets. - Request: fetch only relevant URLs and preserve redirects, the final response URL, status, selected headers, and timing.
- Report: distinguish an HTTP result from the actual requirement—for example, “returns 200” is not the same as “contains a PDF link” or “has the expected canonical tag.”
Keep a requested URL and a final URL as separate fields. A redirect may be healthy, but it can also reveal a moved, misconfigured, or unexpectedly external resource. Record a timestamp so a later run can be compared with the earlier result.
Can I use robots.txt to tell a scraper what not to crawl?
Yes, as crawler guidance. Google defines it as a file that tells search-engine crawlers which URLs they can access; it is not authentication, an access-control list, or a reliable way to keep a private URL out of search results. A blocked URL can still appear in results without a crawlable description, and different crawlers can interpret syntax differently.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
The file belongs at the root of the relevant origin, such as https://example.com/robots.txt. Its scope is limited to that host, protocol, and port. Paths are case-sensitive. Use UTF-8 text, place crawler-specific rule groups deliberately, and use fully qualified sitemap locations. Do not assume a rule on www.example.com governs example.com, another port, or another protocol.
For your own checker, treat an inaccessible or malformed file as a condition to report and review, not as permission to crawl everything. Authentication and authorization must be handled by the site’s approved mechanism.
How do I find all URLs on a website?
Start with robots.txt and sitemaps
Fetch the root robots.txt, parse its Sitemap: lines, and follow both sitemap indexes and URL-set files. A sitemap encourages discovery; it does not constrain Google to crawl only listed URLs, and it is not a guarantee that every listed URL works.
Small, controlled Python discovery template
This script reads sitemap locations, handles a sitemap index, de-duplicates URLs, and checks only URLs whose path matches an optional pattern. It deliberately uses conservative timeouts and a delay.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →from __future__ import annotations
import csv
import re
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
import requests
from xml.etree import ElementTree as ET
START = "https://example.com"
PATH_RE = re.compile(r"^/(docs|downloads)/") # change or set to None
DELAY = 0.5
TIMEOUT = 20
HEADERS = {"User-Agent": "ResourceChecker/1.0 (+https://example.com/contact)"}
def local_name(tag: str) -> str:
return tag.rsplit("}", 1)[-1]
def sitemap_locations(root: str) -> list[str]:
robots_url = urljoin(root.rstrip("/") + "/", "robots.txt")
r = requests.get(robots_url, headers=HEADERS, timeout=TIMEOUT)
r.raise_for_status()
return [line.split(":", 1)[1].strip()
for line in r.text.splitlines()
if line.lower().startswith("sitemap:")]
def expand_sitemaps(locations: list[str]) -> list[str]:
seen_sitemaps, urls = set(), []
pending = list(locations)
while pending:
sitemap = pending.pop(0)
if sitemap in seen_sitemaps:
continue
seen_sitemaps.add(sitemap)
r = requests.get(sitemap, headers=HEADERS, timeout=TIMEOUT)
r.raise_for_status()
root = ET.fromstring(r.content)
kind = local_name(root.tag)
for node in root:
if local_name(node.tag) == "sitemapindex":
continue
loc = next((child.text.strip() for child in node
if local_name(child.tag) == "loc" and child.text), None)
if not loc:
continue
if kind == "sitemapindex":
pending.append(loc)
else:
urls.append(loc)
return list(dict.fromkeys(urls))
def check(url: str) -> dict:
checked = datetime.now(timezone.utc).isoformat()
try:
r = requests.get(url, headers=HEADERS, timeout=TIMEOUT,
allow_redirects=True)
content_type = r.headers.get("content-type", "")
result = "ok" if 200 <= r.status_code < 400 else "http_error"
if "text/html" in content_type and "<title" not in r.text.lower():
result = "missing_expected_title_marker"
return {
"requested_url": url, "final_url": r.url,
"status": r.status_code, "content_type": content_type,
"result": result, "checked_at": checked,
}
except requests.RequestException as exc:
return {"requested_url": url, "final_url": "", "status": "",
"content_type": "", "result": type(exc).__name__,
"checked_at": checked, "error": str(exc)}
if __name__ == "__main__":
candidates = expand_sitemaps(sitemap_locations(START))
if PATH_RE:
candidates = [u for u in candidates
if PATH_RE.search(urlparse(u).path)]
rows = []
for url in candidates:
rows.append(check(url))
time.sleep(DELAY)
with open("resource-report.csv", "w", newline="", encoding="utf-8") as f:
fields = ["requested_url", "final_url", "status", "content_type",
"result", "checked_at", "error"]
writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
writer.writerows(rows)
Install the dependency with python -m pip install requests. The XML parser above is intentionally small; production jobs should add limits for sitemap count, response size, and total URLs, and should reject URLs outside the approved host set.
Rank #2
How do I check if a website URL is working?
Interpret status and content separately
2xxnormally means the server returned a successful response, but the body may still be an error page or the wrong resource.3xxrecords a redirect; inspectfinal_url, redirect count, and whether the destination remains in scope.4xxindicates a client-side failure such as not found or forbidden. Do not retry aggressively.5xxindicates a server-side failure; retry with backoff only within an explicit budget.- A timeout, TLS error, DNS failure, or connection reset is a transport result, not an HTTP status.
Add checks that match the job: required text, a content type, a file signature, a canonical URL, a heading, or a JSON field. Keep the raw body only when policy and storage limits allow it; otherwise save a short hash or extracted evidence.
Scrapy template for larger URL sets
Scrapy’s SitemapSpider can locate sitemap URLs from robots.txt, process sitemap indexes, and route matching URL patterns to callbacks. This is a better fit when discovery, retries, throttling, deduplication, and structured output need to be maintained over time.
import scrapy
from scrapy.spiders import SitemapSpider
class ResourceSpider(SitemapSpider):
name = "resources"
sitemap_urls = ["https://example.com/robots.txt"]
sitemap_rules = [
(r"/docs/", "parse_resource"),
(r"/downloads/", "parse_resource"),
]
custom_settings = {
"USER_AGENT": "ResourceChecker/1.0 (+https://example.com/contact)",
"DOWNLOAD_DELAY": 0.5,
"CONCURRENT_REQUESTS_PER_DOMAIN": 4,
"ROBOTSTXT_OBEY": True,
"FEEDS": {"resource-report.jsonl": {"format": "jsonlines"}},
}
def parse_resource(self, response):
content_type = response.headers.get(b"Content-Type", b"").decode("latin-1")
body = response.body
yield {
"requested_url": response.request.url,
"final_url": response.url,
"status": response.status,
"content_type": content_type,
"title_marker": b"<title" in body.lower(),
"length": len(body),
"checked_at": response.headers.get(b"Date", b"").decode("latin-1"),
}
Run it with scrapy runspider resources.py. Scrapy exposes the response URL, status, headers, body, and extracted values to the callback. Use a separate item pipeline when reports must be written to a database, and configure retry rules narrowly so a broken site is not hammered.
Simple script or Scrapy: which should you choose?
| Requirement | Requests script | Scrapy |
|---|---|---|
| Dozens or a few hundred approved URLs | Low setup; easy to customize | More structure than necessary |
| Nested sitemaps and URL-pattern callbacks | Implement parsing and routing yourself | SitemapSpider provides these primitives |
| Retries, throttling, deduplication, feeds | Build and test each piece | Framework settings and extensions help |
| Rendered JavaScript | Requests sees the server response only | Still needs a browser integration or pre-rendered source |
| Long-term maintenance | Small codebase, but more custom policy | More conventions and dependencies |
Neither approach is universally best. Choose based on scale, page behavior, discovery needs, metadata requirements, and the output your team must maintain. A browser is necessary when the resource appears only after JavaScript runs; it also introduces higher cost, longer waits, cookie state, and new failure modes.
Validation and safety checklist
- Confirm the target host and URL list are approved.
- Fetch and parse root
robots.txt; report HTTP errors and syntax problems. - Keep host, protocol, port, and path scope explicit.
- Set connect/read timeouts, a concurrency limit, and a delay.
- Cap total URLs, sitemap bytes, response bytes, and retry attempts.
- Use a descriptive user agent and contact address where appropriate.
- Store requested URL, final URL, status, selected headers, timestamp, and a task-specific result.
- Review important resources for accessibility and rendering when diagnosing search-crawler behavior; a robots rule alone cannot prove that a page is private or secure.
Troubleshooting common failures
Robots file returns 404 or HTML
Verify the exact origin, protocol, port, and root path. A 404 or non-text response should be reported. Do not silently treat it as “allow all.”
Sitemap XML will not parse
Check the response content type and size, save the first bytes for diagnosis, and reject malformed or unexpectedly compressed content. Follow only HTTPS or other schemes your policy explicitly permits.
Everything is 200 but the report is wrong
Many sites return a branded “not found” page with status 200. Add body, title, canonical, or content-type checks and label the result separately from HTTP success.
Requests get blocked or rate-limited
Reduce concurrency, increase delay, honor published crawler guidance, cache results, and stop on repeated failures. Do not attempt to defeat a CAPTCHA, bot check, login, or access control.
Expected content is missing
The page may require JavaScript, a session, a location, or an authorization header. Use an approved browser or API workflow and document that the plain HTTP template cannot observe rendered content.
Redirects leave the approved host
Record the final URL, flag the out-of-scope destination, and decide whether to follow it before the next run. Never expand scope implicitly.
Or skip the browser setup
When the goal is a clean screenshot rather than HTML-level inspection, ScreenshotNeo provides a single website-screenshot API call. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools.
For API parameters and the complete option list, see the ScreenshotNeo documentation. A direct call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
The service supports PNG, JPEG, WebP, and PDF; full-page or CSS-selector captures; dark mode, device presets, arbitrary viewports, retina scale; PDF paper, margins, orientation, and page ranges; custom CSS and JavaScript; clicks, waits, blocked requests, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which helps with migrations.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to start.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Python, cURL, and Node.js screenshot calls
These examples use the same endpoint; adapt only the target URL. Keep the access key out of source control and consult the API documentation for optional parameters.
Free tools Windows power users keep installed
One-click scans. No signup required.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
How to make reports reliable
Version your checker and its rules. Keep a run identifier, URL-source identifier (robots, sitemap, or approved list), and configuration hash. Compare status, final URL, content type, and task-specific checks across runs rather than treating every change as an incident. Separate transient transport errors from durable content failures, and retain enough evidence to reproduce a decision without storing sensitive bodies unnecessarily.
Best Value
Frequently Asked Questions
How do I check a sitemap with Python?
Fetch the sitemap URL with a timeout, parse its XML namespace, distinguish a sitemap index from a URL set, recursively follow each index location, de-duplicate <loc> values, and then request only the URLs that match your approved patterns.
Does a sitemap mean Google will crawl every listed URL?
No. A sitemap encourages discovery; it does not restrict Google to the listed URLs or guarantee that each URL is crawled or valid.
Can a robots.txt file protect confidential files?
No. It is crawler guidance. Use authentication, authorization, and server-side controls for confidentiality.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhy does a browser show content that my Python request cannot find?
The content may be rendered by JavaScript, require cookies or authorization, vary by location, or be loaded after an API call. A plain HTTP client sees only the server response it receives.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

