Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A dependable Python website monitor does more than request a URL and check for 200. It needs explicit timeouts, connection reuse, redirect and TLS handling, content checks, persistent state, transition-based alerts, and safe scheduling. This guide builds that monitor in small steps, then shows when a hosted service is a better fit.

What the monitor should do

For each target, the script should store an expected health policy, make a bounded HTTP request, classify the result, record useful evidence, compare it with the previous result, and notify only when the state changes. A practical policy can include accepted status codes and an optional text marker that must appear in the response.

  • Success: an accepted status code and, when configured, the expected marker are present.
  • Redirect: redirects are followed and the final URL is recorded; decide per site whether this remains healthy or is degraded.
  • Client or server error: 4xx and 5xx responses are retained as distinct failures.
  • Timeout, DNS, or TLS failure: the exception type and message are recorded.
  • Unexpected content: the HTTP request succeeds, but the required marker is missing or normalized content changed.

Do not treat a status code as the whole truth. A page can return 200 while displaying an outage message, redirecting to a login page, or serving an incomplete response.

Install Python and Requests

Use Python 3.9 or newer in a virtual environment:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
python -m pip install requests

Requests provides status-code access, timeouts, redirects, exceptions, TLS verification, and connection reuse through a persistent Session. Keep certificate verification enabled; disabling it hides real certificate problems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define targets and policies

Keep configuration separate from code so adding a URL does not require editing the monitor. This example uses JSON:

{
  "targets": [
    {
      "name": "Example home page",
      "url": "https://example.com/",
      "accepted_statuses": [200],
      "contains": "Example Domain",
      "timeout": [5, 20]
    },
    {
      "name": "Redirecting service",
      "url": "https://example.org/",
      "accepted_statuses": [200, 204],
      "contains": null,
      "timeout": [5, 20]
    }
  ],
  "state_file": "monitor-state.json",
  "log_file": "monitor-results.jsonl"
}

The two timeout values are connect and read seconds. A tuple prevents a stalled DNS connection or server from holding the loop forever.

Build a single check

Create monitor.py. The check records a timestamp, elapsed time, status, final URL, and a stable outcome label. It catches network failures without confusing them with HTTP responses.

from __future__ import annotations

import hashlib
import json
import os
import re
import time
from dataclasses import dataclass, asdict
from datetime import datetime, timezone
from pathlib import Path
from typing import Any

import requests
from requests import Session
from requests.exceptions import RequestException


@dataclass
class CheckResult:
    name: str
    url: str
    checked_at: str
    outcome: str
    status_code: int | None
    final_url: str | None
    elapsed_ms: int | None
    content_hash: str | None
    detail: str | None


def normalize_html(text: str) -> str:
    """Reduce harmless formatting differences before hashing content."""
    text = re.sub(r"<scriptb[^>]*>.*?</script>", " ", text,
                  flags=re.I | re.S)
    text = re.sub(r"<styleb[^>]*>.*?</style>", " ", text,
                  flags=re.I | re.S)
    text = re.sub(r"s+", " ", text)
    return text.strip()


def check_target(session: Session, target: dict[str, Any]) -> CheckResult:
    started = time.perf_counter()
    checked_at = datetime.now(timezone.utc).isoformat()
    timeout = tuple(target.get("timeout", [5, 20]))
    accepted = set(target.get("accepted_statuses", [200]))

    try:
        response = session.get(
            target["url"],
            timeout=timeout,
            allow_redirects=True,
        )
        elapsed_ms = round((time.perf_counter() - started) * 1000)
        normalized = normalize_html(response.text)
        digest = hashlib.sha256(normalized.encode("utf-8")).hexdigest()

        if response.status_code not in accepted:
            outcome = "client_error" if 400 <= response.status_code < 500 
                else "server_error" if response.status_code >= 500 
                else "unexpected_status"
            detail = f"HTTP {response.status_code}"
        elif target.get("contains") and target["contains"] not in response.text:
            outcome = "unexpected_content"
            detail = f"Missing marker: {target['contains']}"
        else:
            outcome = "success"
            detail = None

        return CheckResult(
            target["name"], target["url"], checked_at, outcome,
            response.status_code, response.url, elapsed_ms, digest, detail
        )
    except requests.exceptions.Timeout as exc:
        outcome, detail = "timeout", str(exc)
    except requests.exceptions.SSLError as exc:
        outcome, detail = "tls_failure", str(exc)
    except requests.exceptions.ConnectionError as exc:
        outcome, detail = "dns_or_connection_failure", str(exc)
    except RequestException as exc:
        outcome, detail = "request_failure", str(exc)

    elapsed_ms = round((time.perf_counter() - started) * 1000)
    return CheckResult(
        target["name"], target["url"], checked_at, outcome,
        None, None, elapsed_ms, None, detail
    )

The content hash is deliberately based on normalized content. For a real site, remove rotating timestamps, ad slots, counters, and other known dynamic regions before hashing. If the page is large, extract a specific stable element instead of hashing the entire document.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Persist state and alert on transitions

Without state, every failed poll sends another alert. Store the previous outcome and hash, then notify only for healthy-to-failed, failed-to-healthy, or meaningful content transitions.

def load_json(path: Path, default: Any) -> Any:
    try:
        return json.loads(path.read_text(encoding="utf-8"))
    except (FileNotFoundError, json.JSONDecodeError):
        return default


def save_json(path: Path, value: Any) -> None:
    temporary = path.with_suffix(path.suffix + ".tmp")
    temporary.write_text(json.dumps(value, indent=2), encoding="utf-8")
    temporary.replace(path)


def notify(message: str) -> None:
    # Replace this with email, an incident system, or a chat webhook.
    print("ALERT:", message)


def changed(previous: dict[str, Any] | None, current: CheckResult) -> bool:
    if previous is None:
        return True
    old_outcome = previous.get("outcome")
    if old_outcome != current.outcome:
        return True
    return (current.outcome == "success" and
            previous.get("content_hash") != current.content_hash)


def run_once(config_path: str = "monitor-config.json") -> None:
    config = load_json(Path(config_path), {})
    state_path = Path(config.get("state_file", "monitor-state.json"))
    log_path = Path(config.get("log_file", "monitor-results.jsonl"))
    state = load_json(state_path, {})

    with requests.Session() as session:
        session.headers.update({
            "User-Agent": "ite_guides-website-monitor/1.0"
        })
        for target in config.get("targets", []):
            result = check_target(session, target)
            old = state.get(target["name"])
            if changed(old, result):
                if old is not None or result.outcome != "success":
                    notify(f"{result.name}: {result.outcome} ({result.detail or result.status_code})")
            record = asdict(result)
            state[target["name"]] = record
            with log_path.open("a", encoding="utf-8") as stream:
                stream.write(json.dumps(record) + "n")

    save_json(state_path, state)


if __name__ == "__main__":
    run_once()

Run a check manually:

python monitor.py

The JSON Lines log preserves every observation for later analysis, while the state file contains only the latest result. The atomic replacement in save_json avoids leaving a half-written state file if the process stops during a write.

Schedule the monitor safely

Linux or macOS cron

Use the virtual environment’s interpreter and an absolute working directory. For example, edit crontab -e and run:

*/5 * * * * cd /opt/site-monitor && /opt/site-monitor/.venv/bin/python monitor.py >> cron.log 2>&1

Choose an interval that the site owner permits; there is no universally correct polling frequency. Use a service manager when you need restart behavior, structured logs, and environment-file secrets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Windows Task Scheduler

  1. Create a basic task with the desired trigger.
  2. Set Program to the full path of python.exe inside .venv.
  3. Set Arguments to the full path of monitor.py.
  4. Set Start in to the project directory so relative state and log paths resolve correctly.

Retries, pooling, and concurrency

A single transient packet loss should not necessarily page someone, but retries can multiply load and hide a genuine failure. Retry only transient connection failures and selected 5xx responses, use exponential backoff, and cap attempts. Do not retry deterministic 401, 403, 404, or content-marker failures.

Session reuses connections. At larger volumes, urllib3’s pooling, retry helpers, proxy support, compression, redirect handling, and thread-safe components can support bounded concurrency. Keep the worker count finite and record which attempt produced the final result. Concurrent checks also need per-host rate limits so a long target list does not become a burst.

Security and operational safeguards

  • Use an honest, identifiable User-Agent and respect authorization, terms, robots guidance, and rate limits.
  • Put webhook tokens, SMTP passwords, and API keys in environment variables or a secret manager, never in the JSON configuration committed to source control.
  • If users can submit URLs, allow only intended schemes such as HTTPS and block loopback, private, link-local, metadata-service, and other reserved destinations. This prevents server-side request forgery.
  • Restrict redirects when monitoring untrusted input, because a safe-looking URL can redirect to an internal address.
  • Limit response size where appropriate and avoid logging cookies, Authorization headers, or sensitive response bodies.
  • Keep clocks synchronized so timestamps and alert transitions are meaningful.

Diagnose common failures

Symptom Likely cause Fix
Every check ends in timeout Connect or read timeout is too short, DNS is slow, or the host is unreachable Test DNS and connectivity separately, then tune the two timeout values; never remove the timeout.
TLS failure Expired, mismatched, or untrusted certificate Fix the certificate chain or trust store. Do not set verify=False in production.
200 but marked failed Expected text is absent, often because content is client-rendered or the site returned a login/interstitial page Inspect the final URL and response body; monitor a server-rendered marker or use a browser-capable capture for JavaScript-heavy pages.
Alerts repeat every run Previous state is not writable or the script runs from different working directories Use absolute paths, check permissions, and verify that the state file changes after each run.
False content-change alerts Ads, timestamps, personalization, or rotating tokens change the hash Extract and normalize a stable region, or compare a required marker rather than the whole page.
Requests are blocked Polling is too aggressive, credentials are missing, or the site forbids automated access Slow down, authenticate only with permission, review the site’s policy, and stop if access is not authorized.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a script is no longer enough

Keep the custom monitor when you have a small URL list, a known schedule, and a narrowly defined rule. It gives maximum control over headers, cookies, parsing, and notification logic, but you own deployment, persistence, retries, dashboards, and maintenance.

A monitoring platform becomes attractive when you need concurrent probes, durable history, DNS/SSL/port checks, ping checks, content-change detection, alert routing, reports, metrics, or centralized SSRF controls. A useful comparison is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Capability Custom Python script Monitoring platform
Setup Install, deploy, schedule, and maintain code Configure targets in a hosted interface or API
Request and content rules Highest flexibility Depends on the product’s rule set
History and reports You build storage and visualization Usually included as a managed feature
Probe breadth Whatever you implement Often includes HTTP, API, DNS, SSL, port, and ping checks
Concurrency You tune workers, pooling, and rate limits Managed by the service
Security You must validate targets and protect secrets May provide built-in SSRF and access controls
Ongoing cost Infrastructure and engineering time Subscription and vendor dependency

Or skip the browser setup

If your goal is a clean visual capture rather than writing and maintaining a browser stack, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

One GET request returns PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for authentication and response details. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. There are 63 options, including full-page and CSS-selector captures, lazy-image loading, device presets, dark mode, retina scale, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification. Existing screenshot-API parameter names also work.

All features are available on every plan: 1,000 shots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

FAQ

Should redirects count as downtime?

That depends on the service contract. Record the final URL and define whether a redirect is healthy, degraded, or failed for that target.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How can I monitor an authenticated page?

Use an authorized session with carefully scoped cookies or headers, keep credentials outside source control, and avoid writing secrets or private response content to logs.

Can this monitor JavaScript-rendered pages?

Requests observes the server response, not a browser’s rendered DOM. For client-rendered content, expose a server-side health endpoint or use a permitted browser-capable capture service.

What should I retain from each check?

Retain the UTC timestamp, requested and final URLs, outcome, status code, elapsed time, exception details, and a content digest. That evidence is enough to investigate most transitions without storing entire sensitive pages.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.