Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A dependable link checker is a small crawler, not just a script that sends one HTTP request. It needs to fetch pages within a defined scope, extract and normalize links, probe destinations without overloading servers, follow and record redirects, and report exact outcomes so you can tell a broken link from a timeout or an access restriction. This guide builds that workflow in Python with Requests, then explains what to add before running it across a real site.

What a custom link checker should do

A checker usually has two jobs: discover links in pages and determine what happens when each destination is requested. Keep those results distinct. A URL can be reachable but point to the wrong page; an HTTP error can be caused by a temporary external outage; a timeout is not the same as a 404.

For a useful report, retain the source page, original link text or spelling, normalized destination, HTTP status or exception type, redirect chain, final URL, content type, and elapsed time. Include a suggested action, but do not reduce the underlying result to a single “valid” or “invalid” flag.

Set scope and safety limits first

Before making requests, decide what the checker is allowed to visit. This is especially important when the seed URL can come from a user: an unrestricted crawler can follow links outside the intended site or be abused to make requests to internal services.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Accept only http and https; reject other schemes before requesting them.
  • Set maximum pages, maximum links, concurrency, request timeout, redirect-hop limit, and a descriptive user-agent.
  • Choose whether to stay on the seed’s origin. An origin is the combination of scheme, hostname, and port.
  • Resolve and check every redirect destination as well as every discovered link. A same-origin check on the initial URL alone is not enough.
  • Respect robots.txt rules for the checker’s user-agent. The W3C Link Checker documentation says it honors robots exclusion rules and supports a W3C-checklink user-agent rule (W3C Link Checker documentation).

For a publicly reachable production crawler, also account for DNS and private-address protections: validate resolved addresses and re-check after redirects so a hostname cannot resolve or redirect to a private network target. Keep TLS certificate verification enabled; do not “fix” certificate errors by disabling verification.

Build a single-page checker in Python

This runnable starting point checks links found in one HTML page. Install Requests with python -m pip install requests, save the code as link_check.py, and run python link_check.py https://example.com. It uses a HEAD request first and falls back to a streamed GET when HEAD is unsupported or denied. The example intentionally checks one page only; it does not yet crawl the discovered pages or implement robots policy, rate limiting, or DNS safety controls.

import json
import sys
import time
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin, urlsplit, urlunsplit

import requests

TIMEOUT_SECONDS = 10
USER_AGENT = "ExampleLinkChecker/1.0 (contact: webmaster@example.com)"
ALLOWED_SCHEMES = {"http", "https"}


class LinkParser(HTMLParser):
    """Collect common navigational and embedded-resource references."""

    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.links = []

    def handle_starttag(self, tag, attrs):
        attrs = dict(attrs)
        attr = "href" if tag in {"a", "area", "link"} else "src"
        value = attrs.get(attr)
        if value and value.strip():
            self.links.append(value.strip())


def normalize(base_url, raw_reference):
    """Resolve a reference, discard its fragment, and reject unsupported schemes."""
    absolute = urljoin(base_url, raw_reference)
    absolute, _fragment = urldefrag(absolute)
    parts = urlsplit(absolute)
    if parts.scheme.lower() not in ALLOWED_SCHEMES or not parts.hostname:
        return None
    # Normalize scheme/host for comparison while retaining path and query spelling.
    netloc = parts.netloc.lower()
    return urlunsplit((parts.scheme.lower(), netloc, parts.path, parts.query, ""))


def probe(session, url):
    """Return status, redirect details, timing, and a distinct network error if any."""
    started = time.monotonic()
    try:
        response = session.head(
            url, allow_redirects=True, timeout=TIMEOUT_SECONDS
        )
        if response.status_code in {403, 405, 501}:
            response.close()
            response = session.get(
                url, allow_redirects=True, timeout=TIMEOUT_SECONDS, stream=True
            )
        result = {
            "status": response.status_code,
            "redirect_chain": [
                {"status": item.status_code, "url": item.url,
                 "location": item.headers.get("Location")}
                for item in response.history
            ],
            "final_url": response.url,
            "content_type": response.headers.get("Content-Type"),
        }
        response.close()
    except requests.RequestException as exc:
        result = {"error_class": type(exc).__name__, "detail": str(exc)}
    result["elapsed_seconds"] = round(time.monotonic() - started, 3)
    return result


def main(seed_url):
    seed = normalize(seed_url, seed_url)
    if seed is None:
        raise SystemExit("Seed must be an absolute http:// or https:// URL")

    session = requests.Session()
    session.headers.update({"User-Agent": USER_AGENT})
    try:
        page_response = session.get(seed, timeout=TIMEOUT_SECONDS)
        page_response.raise_for_status()
        content_type = page_response.headers.get("Content-Type", "")
        if "html" not in content_type.lower():
            raise SystemExit(f"Seed did not return HTML: {content_type or 'no Content-Type'}")
        page_url = page_response.url
        parser = LinkParser()
        parser.feed(page_response.text)
        page_response.close()

        seen = set()
        rows = []
        for raw in parser.links:
            normalized = normalize(page_url, raw)
            if normalized is None:
                rows.append({"source_page": page_url, "discovered": raw,
                             "error_class": "UnsupportedSchemeOrInvalidURL"})
                continue
            if normalized in seen:
                continue
            seen.add(normalized)
            rows.append({"source_page": page_url, "discovered": raw,
                         "normalized_url": normalized,
                         **probe(session, normalized)})
        print(json.dumps(rows, indent=2, ensure_ascii=False))
    finally:
        session.close()


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit(f"Usage: {sys.argv[0]} https://example.com")
    main(sys.argv[1])

The code preserves redirect status and URLs from Requests’ response history, not just the final destination. Requests documents its session, request methods, timeout, TLS verification, and redirect controls in its API reference (Requests API). The fallback set is a practical policy, not a guarantee that all unusual HEAD behavior is detected. A destination that returns a misleading 200 to HEAD but fails on GET may require a GET check based on resource type or an explicit validation mode.

Turn the one-page probe into a site crawler

A site-wide checker adds a queue and a visited set. Start with the seed URL, fetch each allowed HTML page, extract links, enqueue eligible same-scope pages, and probe destinations. Maintain separate limits for pages fetched and links probed: one page can contain thousands of links, and external links should not automatically become pages to crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Fetch robots rules. Retrieve the origin’s /robots.txt, parse rules for the chosen user-agent, and skip disallowed URLs. Treat an unavailable robots file according to an explicit policy and document that choice.
  2. Fetch a page. Use a Requests session with a descriptive user-agent, a finite timeout, and TLS verification enabled. Check status and content type before parsing; do not feed PDFs or images to the HTML parser.
  3. Extract references. Use HTMLParser.handle_starttag to collect href from a, area, and link, and src from configured resource elements such as img, script, and iframe. Extend the set deliberately for srcset, inline CSS, or other formats rather than assuming every URL is an HTML attribute.
  4. Resolve and validate. Resolve each reference against the page URL with urljoin, remove its fragment with urldefrag, then apply scheme, host, and scope checks. Python documents urljoin as combining a base URL with another URL (Python URL parsing documentation). Because an absolute reference can replace the base host, scope checks must occur after joining.
  5. Deduplicate carefully. A fragment identifies a location within a document, not a distinct HTTP fetch, so remove it before deduplication. Lowercase scheme and hostname for comparisons. Avoid aggressive normalization of paths or query strings: servers can treat their spelling, case, or parameter order as significant.
  6. Probe with policy. Try HEAD for ordinary resources, then use a GET fallback for unsupported or inconclusive responses. Keep timeout and redirect limits in force for both methods, and validate every redirect target against scheme, host, and safety rules.
  7. Throttle work. Use bounded workers, per-host delays, a maximum redirect-hop count, and backoff only for transient failures such as selected server errors or connection resets. Cache each normalized URL’s result for the duration of a run.
  8. Write actionable output. Save JSON or CSV with source page, discovered spelling, normalized URL, status, exception class, redirect chain, final URL, content type, elapsed time, and suggested action. Group errors by source page to make repairs easier.

Choose HEAD, GET, and redirect behavior deliberately

HEAD asks for response metadata without the response body. MDN describes it as requesting the headers a server would send for GET (MDN: HEAD). It can reduce bandwidth, especially for large resources, but some servers block it, implement it inconsistently, or return a result that does not represent a GET request.

Use GET when HEAD returns 405 or 501, when access controls produce a response that needs confirmation, or when the body itself matters. For link reachability, a streamed GET can avoid downloading the full body; close the response promptly. For content validation, read only what the validation requires and enforce a size limit. Do not interpret a 403 as a broken link automatically: it may require authentication or block automated clients.

Redirects are 3xx responses with a Location header, as described by MDN (MDN: redirections). Retain every hop and the final URL. A 301 or 308 is permanent; 302, 303, and 307 have differing temporary and method semantics, so do not collapse them into an unexplained “redirected” result. Requests follows redirects by default for GET and HEAD, and exposes prior responses through response.history (Requests API).

Interpret results without false alarms

  • 2xx: the server returned a successful HTTP response. It does not prove the expected content is present.
  • 3xx: a redirect occurred. Report the chain and final URL; flag loops, excessive hops, or a destination that leaves the allowed scope.
  • 4xx: the server returned a client-side error. A 404 or 410 may indicate a dead destination; 401 and 403 often mean authentication or access policy rather than a typo.
  • 5xx: the server returned a server-side error. Record it as observed and consider a later retry rather than declaring a permanent failure.
  • Network exception: report DNS, connection, TLS, timeout, or other exception classes separately. These are not HTTP status codes.
  • Parse or scheme issue: report malformed references and unsupported schemes separately from destination failures.

Python’s URL and HTTP libraries also distinguish HTTP errors from other failures; the exception class and context are valuable diagnostics rather than noise (Python urllib.error documentation). A successful HTTP result cannot establish that a JavaScript-rendered link works, that the intended page content is present, or that a signed-in user can access it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Improve the report for real maintenance work

Use a stable output schema so results can be diffed between runs. Store timestamps and checker version, but avoid including credentials or sensitive query values in shared reports. Include both the source page and the link’s original spelling: the normalized fetch target is useful for deduplication, while the source context tells an editor where to fix it.

Suggested actions should be evidence-based: “update internal href” for a confirmed local 404, “review redirect target” for a changed destination, or “retry/verify manually” for a timeout or external 5xx. A repeated failure across multiple source pages is usually one destination problem, not many independent link problems.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

HEAD says 403, but the page opens in a browser

The site may block HEAD or automated user-agents. Retry with GET under the same timeout and robots policy. If the request still receives 403, report access denied rather than labeling the URL dead; browser access may depend on cookies, login, or anti-bot checks.

Every request times out or TLS fails

Check the hostname, network access, system clock, and certificate chain. Keep certificate verification on. Increase the timeout only when a slower response is plausible; a long timeout multiplied across many links can stall an entire run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A relative link resolves to the wrong site

Inspect the effective page URL after redirects and resolve the reference against that URL, not necessarily the original seed. Then apply same-origin checks after joining. Scheme-relative references such as //cdn.example.org/file.js intentionally inherit the current scheme but change hosts.

The report contains duplicates with different fragments

Strip fragments before deduplication. Preserve the original discovered value in the report so anchors remain visible to editors, but make one HTTP probe for the shared document URL.

A redirect escapes the crawl boundary

Check every hop and final destination against allowed schemes and hosts. Do not assume an in-scope initial URL guarantees an in-scope redirect target. Stop at the configured hop limit and record where the chain ended.

The result changes between runs

External services can rate-limit, block bots, or have transient outages. Add per-host pacing, bounded concurrency, and limited backoff for transient conditions. Keep the observed status and timestamp rather than silently replacing errors with a generic success after retries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your next step is to capture a visual snapshot of a page while investigating a link or layout issue, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; see the ScreenshotNeo site and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are not billed. Its MCP server lets AI agents use screenshot tools. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, with no card required.

Frequently Asked Questions

Does a successful status code prove a link is good?

No. It confirms an HTTP response, not that the destination contains the intended information or works for an authenticated user.

Can a link checker validate links created by JavaScript?

A basic HTML parser sees links present in returned markup. It does not execute page JavaScript, so dynamically created links need a browser-based discovery step.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.