Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a staged crawler: discover approved URLs from robots.txt and sitemaps, request them with explicit limits, retain the final URL and response metadata, then run a task-specific content check. The templates below cover a small Python script, a Scrapy sitemap crawler, validation, reporting, JavaScript-heavy pages, and common failures. They are starting points, not permission to bypass authentication, access controls, or a site’s terms.

What a resource-checking scraper should do

A useful checker separates four jobs that are often mixed together:

  • Input: an approved starting host or URL list, resource types or path patterns, concurrency and delay limits, and an output format.
  • Discovery: read the host’s root robots.txt and sitemap references, then expand sitemap indexes and URL sets.
  • Request: fetch only relevant URLs and preserve redirects, the final response URL, status, selected headers, and timing.
  • Report: distinguish an HTTP result from the actual requirement—for example, “returns 200” is not the same as “contains a PDF link” or “has the expected canonical tag.”

Keep a requested URL and a final URL as separate fields. A redirect may be healthy, but it can also reveal a moved, misconfigured, or unexpectedly external resource. Record a timestamp so a later run can be compared with the earlier result.

Can I use robots.txt to tell a scraper what not to crawl?

Yes, as crawler guidance. Google defines it as a file that tells search-engine crawlers which URLs they can access; it is not authentication, an access-control list, or a reliable way to keep a private URL out of search results. A blocked URL can still appear in results without a crawlable description, and different crawlers can interpret syntax differently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The file belongs at the root of the relevant origin, such as https://example.com/robots.txt. Its scope is limited to that host, protocol, and port. Paths are case-sensitive. Use UTF-8 text, place crawler-specific rule groups deliberately, and use fully qualified sitemap locations. Do not assume a rule on www.example.com governs example.com, another port, or another protocol.

For your own checker, treat an inaccessible or malformed file as a condition to report and review, not as permission to crawl everything. Authentication and authorization must be handled by the site’s approved mechanism.

How do I find all URLs on a website?

Start with robots.txt and sitemaps

Fetch the root robots.txt, parse its Sitemap: lines, and follow both sitemap indexes and URL-set files. A sitemap encourages discovery; it does not constrain Google to crawl only listed URLs, and it is not a guarantee that every listed URL works.

Small, controlled Python discovery template

This script reads sitemap locations, handles a sitemap index, de-duplicates URLs, and checks only URLs whose path matches an optional pattern. It deliberately uses conservative timeouts and a delay.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from __future__ import annotations

import csv
import re
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
import requests
from xml.etree import ElementTree as ET

START = "https://example.com"
PATH_RE = re.compile(r"^/(docs|downloads)/")  # change or set to None
DELAY = 0.5
TIMEOUT = 20
HEADERS = {"User-Agent": "ResourceChecker/1.0 (+https://example.com/contact)"}


def local_name(tag: str) -> str:
    return tag.rsplit("}", 1)[-1]


def sitemap_locations(root: str) -> list[str]:
    robots_url = urljoin(root.rstrip("/") + "/", "robots.txt")
    r = requests.get(robots_url, headers=HEADERS, timeout=TIMEOUT)
    r.raise_for_status()
    return [line.split(":", 1)[1].strip()
            for line in r.text.splitlines()
            if line.lower().startswith("sitemap:")]


def expand_sitemaps(locations: list[str]) -> list[str]:
    seen_sitemaps, urls = set(), []
    pending = list(locations)
    while pending:
        sitemap = pending.pop(0)
        if sitemap in seen_sitemaps:
            continue
        seen_sitemaps.add(sitemap)
        r = requests.get(sitemap, headers=HEADERS, timeout=TIMEOUT)
        r.raise_for_status()
        root = ET.fromstring(r.content)
        kind = local_name(root.tag)
        for node in root:
            if local_name(node.tag) == "sitemapindex":
                continue
            loc = next((child.text.strip() for child in node
                        if local_name(child.tag) == "loc" and child.text), None)
            if not loc:
                continue
            if kind == "sitemapindex":
                pending.append(loc)
            else:
                urls.append(loc)
    return list(dict.fromkeys(urls))


def check(url: str) -> dict:
    checked = datetime.now(timezone.utc).isoformat()
    try:
        r = requests.get(url, headers=HEADERS, timeout=TIMEOUT,
                         allow_redirects=True)
        content_type = r.headers.get("content-type", "")
        result = "ok" if 200 <= r.status_code < 400 else "http_error"
        if "text/html" in content_type and "<title" not in r.text.lower():
            result = "missing_expected_title_marker"
        return {
            "requested_url": url, "final_url": r.url,
            "status": r.status_code, "content_type": content_type,
            "result": result, "checked_at": checked,
        }
    except requests.RequestException as exc:
        return {"requested_url": url, "final_url": "", "status": "",
                "content_type": "", "result": type(exc).__name__,
                "checked_at": checked, "error": str(exc)}


if __name__ == "__main__":
    candidates = expand_sitemaps(sitemap_locations(START))
    if PATH_RE:
        candidates = [u for u in candidates
                      if PATH_RE.search(urlparse(u).path)]
    rows = []
    for url in candidates:
        rows.append(check(url))
        time.sleep(DELAY)
    with open("resource-report.csv", "w", newline="", encoding="utf-8") as f:
        fields = ["requested_url", "final_url", "status", "content_type",
                  "result", "checked_at", "error"]
        writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
        writer.writeheader()
        writer.writerows(rows)

Install the dependency with python -m pip install requests. The XML parser above is intentionally small; production jobs should add limits for sitemap count, response size, and total URLs, and should reject URLs outside the approved host set.

How do I check if a website URL is working?

Interpret status and content separately

  • 2xx normally means the server returned a successful response, but the body may still be an error page or the wrong resource.
  • 3xx records a redirect; inspect final_url, redirect count, and whether the destination remains in scope.
  • 4xx indicates a client-side failure such as not found or forbidden. Do not retry aggressively.
  • 5xx indicates a server-side failure; retry with backoff only within an explicit budget.
  • A timeout, TLS error, DNS failure, or connection reset is a transport result, not an HTTP status.

Add checks that match the job: required text, a content type, a file signature, a canonical URL, a heading, or a JSON field. Keep the raw body only when policy and storage limits allow it; otherwise save a short hash or extracted evidence.

Scrapy template for larger URL sets

Scrapy’s SitemapSpider can locate sitemap URLs from robots.txt, process sitemap indexes, and route matching URL patterns to callbacks. This is a better fit when discovery, retries, throttling, deduplication, and structured output need to be maintained over time.

import scrapy
from scrapy.spiders import SitemapSpider

class ResourceSpider(SitemapSpider):
    name = "resources"
    sitemap_urls = ["https://example.com/robots.txt"]
    sitemap_rules = [
        (r"/docs/", "parse_resource"),
        (r"/downloads/", "parse_resource"),
    ]
    custom_settings = {
        "USER_AGENT": "ResourceChecker/1.0 (+https://example.com/contact)",
        "DOWNLOAD_DELAY": 0.5,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 4,
        "ROBOTSTXT_OBEY": True,
        "FEEDS": {"resource-report.jsonl": {"format": "jsonlines"}},
    }

    def parse_resource(self, response):
        content_type = response.headers.get(b"Content-Type", b"").decode("latin-1")
        body = response.body
        yield {
            "requested_url": response.request.url,
            "final_url": response.url,
            "status": response.status,
            "content_type": content_type,
            "title_marker": b"<title" in body.lower(),
            "length": len(body),
            "checked_at": response.headers.get(b"Date", b"").decode("latin-1"),
        }

Run it with scrapy runspider resources.py. Scrapy exposes the response URL, status, headers, body, and extracted values to the callback. Use a separate item pipeline when reports must be written to a database, and configure retry rules narrowly so a broken site is not hammered.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Simple script or Scrapy: which should you choose?

Requirement Requests script Scrapy
Dozens or a few hundred approved URLs Low setup; easy to customize More structure than necessary
Nested sitemaps and URL-pattern callbacks Implement parsing and routing yourself SitemapSpider provides these primitives
Retries, throttling, deduplication, feeds Build and test each piece Framework settings and extensions help
Rendered JavaScript Requests sees the server response only Still needs a browser integration or pre-rendered source
Long-term maintenance Small codebase, but more custom policy More conventions and dependencies

Neither approach is universally best. Choose based on scale, page behavior, discovery needs, metadata requirements, and the output your team must maintain. A browser is necessary when the resource appears only after JavaScript runs; it also introduces higher cost, longer waits, cookie state, and new failure modes.

Validation and safety checklist

  • Confirm the target host and URL list are approved.
  • Fetch and parse root robots.txt; report HTTP errors and syntax problems.
  • Keep host, protocol, port, and path scope explicit.
  • Set connect/read timeouts, a concurrency limit, and a delay.
  • Cap total URLs, sitemap bytes, response bytes, and retry attempts.
  • Use a descriptive user agent and contact address where appropriate.
  • Store requested URL, final URL, status, selected headers, timestamp, and a task-specific result.
  • Review important resources for accessibility and rendering when diagnosing search-crawler behavior; a robots rule alone cannot prove that a page is private or secure.

Troubleshooting common failures

Robots file returns 404 or HTML

Verify the exact origin, protocol, port, and root path. A 404 or non-text response should be reported. Do not silently treat it as “allow all.”

Sitemap XML will not parse

Check the response content type and size, save the first bytes for diagnosis, and reject malformed or unexpectedly compressed content. Follow only HTTPS or other schemes your policy explicitly permits.

Everything is 200 but the report is wrong

Many sites return a branded “not found” page with status 200. Add body, title, canonical, or content-type checks and label the result separately from HTTP success.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Requests get blocked or rate-limited

Reduce concurrency, increase delay, honor published crawler guidance, cache results, and stop on repeated failures. Do not attempt to defeat a CAPTCHA, bot check, login, or access control.

Expected content is missing

The page may require JavaScript, a session, a location, or an authorization header. Use an approved browser or API workflow and document that the plain HTTP template cannot observe rendered content.

Redirects leave the approved host

Record the final URL, flag the out-of-scope destination, and decide whether to follow it before the next run. Never expand scope implicitly.

Or skip the browser setup

When the goal is a clean screenshot rather than HTML-level inspection, ScreenshotNeo provides a single website-screenshot API call. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For API parameters and the complete option list, see the ScreenshotNeo documentation. A direct call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

The service supports PNG, JPEG, WebP, and PDF; full-page or CSS-selector captures; dark mode, device presets, arbitrary viewports, retina scale; PDF paper, margins, orientation, and page ranges; custom CSS and JavaScript; clicks, waits, blocked requests, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which helps with migrations.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to start.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Python, cURL, and Node.js screenshot calls

These examples use the same endpoint; adapt only the target URL. Keep the access key out of source control and consult the API documentation for optional parameters.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

How to make reports reliable

Version your checker and its rules. Keep a run identifier, URL-source identifier (robots, sitemap, or approved list), and configuration hash. Compare status, final URL, content type, and task-specific checks across runs rather than treating every change as an incident. Separate transient transport errors from durable content failures, and retain enough evidence to reproduce a decision without storing sensitive bodies unnecessarily.

Frequently Asked Questions

How do I check a sitemap with Python?

Fetch the sitemap URL with a timeout, parse its XML namespace, distinguish a sitemap index from a URL set, recursively follow each index location, de-duplicate <loc> values, and then request only the URLs that match your approved patterns.

Does a sitemap mean Google will crawl every listed URL?

No. A sitemap encourages discovery; it does not restrict Google to the listed URLs or guarantee that each URL is crawled or valid.

Can a robots.txt file protect confidential files?

No. It is crawler guidance. Use authentication, authorization, and server-side controls for confidentiality.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does a browser show content that my Python request cannot find?

The content may be rendered by JavaScript, require cookies or authorization, vary by location, or be loaded after an API call. A plain HTTP client sees only the server response it receives.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.