Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable way to build an automated price tracker is to treat it as a data pipeline: identify a product variant, retrieve the page through a permitted method, extract and validate the price, save a timestamped observation, compare it with a baseline, and send an alert only when a defined rule is met. Start with an official retailer API or feed when one exists. If you must parse HTML, check the retailer’s terms and robots.txt, request conservatively, and make failures visible instead of recording a false price.

What the tracker must do

A useful tracker preserves context around every observation. A number such as 19.99 is not enough to reproduce or trust a decision later.

  • Stable identity: retailer, product ID, URL, and variant (size, color, storage, or pack quantity).
  • Source context: currency, country or region assumptions, and whether the value is a sale, member, or coupon price.
  • Time: the retrieval timestamp in UTC.
  • Quality state: successful extraction, unavailable, blocked, timed out, or changed markup.

Model the workflow as:

  1. Load product configuration.
  2. Check that the URL may be fetched and choose an approved source.
  3. Retrieve the response with a timeout and an identifying user agent.
  4. Parse the intended element and its surrounding product/variant context.
  5. Normalize and validate the money value.
  6. Insert a new observation; never overwrite history.
  7. Compare against a target or prior observation and deduplicate alerts.

Choose an allowed source before writing a scraper

Prefer an API or feed

Many retailers publish an API, product feed, or partner program. It is usually more stable and gives clearer permission than scraping page markup. Read the current terms for the specific retailer and use the documented authentication, rate limits, and fields.

Check robots.txt for the exact path

Python’s urllib.robotparser.RobotFileParser can answer whether a user agent may fetch a URL under the site’s published robots rules; see the Python documentation. The broader URL and request modules are documented in Python’s urllib documentation. Robots rules are an implementation signal, not a complete answer to contractual or legal permission. If the retailer disallows your path or your use is not permitted, stop and use an approved source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

def allowed_by_robots(url: str, user_agent: str) -> bool:
    parts = urlparse(url)
    robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
    rp = RobotFileParser(robots_url)
    try:
        rp.read()
    except Exception as exc:
        raise RuntimeError(f"Could not read {robots_url}: {exc}")
    return rp.can_fetch(user_agent, url)

A crawler design that retrieves robots.txt as part of setup is also described in AWS Prescriptive Guidance. Do not use a scraper to bypass a login, CAPTCHA, bot check, paywall, or other access control.

Define products and a storage schema

Keep configuration separate from code so a selector change does not change a product’s identity. This example uses SQLite, which is adequate for a small personal tracker.

PRODUCTS = [
    {
        "product_id": "acme-headphones-black",
        "retailer": "Example Store",
        "url": "https://shop.example/item/headphones?color=black",
        "currency": "USD",
        "price_selector": "meta[itemprop='price']",
        "target_price": 79.00,
    },
]
CREATE TABLE IF NOT EXISTS observations (
    id INTEGER PRIMARY KEY,
    product_id TEXT NOT NULL,
    observed_at TEXT NOT NULL,
    price_cents INTEGER NOT NULL,
    currency TEXT NOT NULL,
    source_url TEXT NOT NULL,
    status TEXT NOT NULL DEFAULT 'ok'
);
CREATE INDEX IF NOT EXISTS observations_product_time
  ON observations(product_id, observed_at);

CREATE TABLE IF NOT EXISTS alerts (
    product_id TEXT NOT NULL,
    price_cents INTEGER NOT NULL,
    sent_at TEXT NOT NULL,
    PRIMARY KEY(product_id, price_cents)
);

Store integer cents (or the smallest unit for the currency) rather than binary floating-point values. If a retailer returns a currency different from the configured one, mark the observation invalid and investigate instead of converting silently.

Complete Python tracker

Install the two parsing dependencies with python -m pip install requests beautifulsoup4. The script below handles robots checks, timeouts, extraction, validation, SQLite history, and a simple console alert. Replace the example URL, selector, and product settings with values you are permitted to use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#!/usr/bin/env python3
import re
import sqlite3
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation, ROUND_HALF_UP
from email.utils import parsedate_to_datetime
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup

DB = "prices.db"
USER_AGENT = "ExamplePriceTracker/1.0 (+mailto:you@example.com)"
PRODUCTS = [
    {
        "product_id": "acme-headphones-black",
        "retailer": "Example Store",
        "url": "https://shop.example/item/headphones?color=black",
        "currency": "USD",
        "price_selector": "meta[itemprop='price']",
        "target_price": Decimal("79.00"),
    },
]


def init_db():
    with sqlite3.connect(DB) as db:
        db.executescript("""
        CREATE TABLE IF NOT EXISTS observations (
            id INTEGER PRIMARY KEY,
            product_id TEXT NOT NULL,
            observed_at TEXT NOT NULL,
            price_cents INTEGER NOT NULL,
            currency TEXT NOT NULL,
            source_url TEXT NOT NULL,
            status TEXT NOT NULL
        );
        CREATE TABLE IF NOT EXISTS alerts (
            product_id TEXT NOT NULL,
            price_cents INTEGER NOT NULL,
            sent_at TEXT NOT NULL,
            PRIMARY KEY(product_id, price_cents)
        );
        """)


def robots_allow(url):
    p = urlparse(url)
    rp = RobotFileParser(f"{p.scheme}://{p.netloc}/robots.txt")
    rp.read()
    return rp.can_fetch(USER_AGENT, url)


def parse_money(raw, expected_currency):
    text = " ".join(raw.replace("xa0", " ").split())
    # Adapt this expression for the retailer's locale; do not guess separators.
    match = re.search(r"(?:USD|US\$|\$)\s*([0-9][0-9,]*(?:\.[0-9]{2})?)", text, re.I)
    if not match:
        raise ValueError(f"No {expected_currency} amount in {raw!r}")
    try:
        amount = Decimal(match.group(1).replace(",", ""))
    except InvalidOperation as exc:
        raise ValueError("Unparseable amount") from exc
    if amount < 0 or amount > Decimal("100000000"):
        raise ValueError("Amount outside configured sanity range")
    return int((amount * 100).quantize(Decimal("1"), rounding=ROUND_HALF_UP))


def fetch_price(product):
    if not robots_allow(product["url"]):
        raise PermissionError("robots.txt disallows this URL for the configured user agent")
    response = requests.get(
        product["url"],
        headers={"User-Agent": USER_AGENT, "Accept": "text/html"},
        timeout=(10, 30),
    )
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    node = soup.select_one(product["price_selector"])
    if node is None:
        raise ValueError("Price selector matched no element")
    raw = node.get("content") or node.get_text(" ", strip=True)
    cents = parse_money(raw, product["currency"])
    return cents


def latest_price(db, product_id):
    row = db.execute(
        "SELECT price_cents FROM observations WHERE product_id=? "
        "ORDER BY observed_at DESC LIMIT 1", (product_id,)
    ).fetchone()
    return row[0] if row else None


def record(product, cents):
    now = datetime.now(timezone.utc).isoformat()
    with sqlite3.connect(DB) as db:
        previous = latest_price(db, product["product_id"])
        db.execute("INSERT INTO observations VALUES (NULL,?,?,?,?,?,?)", (
            product["product_id"], now, cents, product["currency"], product["url"], "ok"
        ))
        target_cents = int((product["target_price"] * 100).quantize(Decimal("1")))
        if cents <= target_cents and not db.execute(
            "SELECT 1 FROM alerts WHERE product_id=? AND price_cents=?",
            (product["product_id"], cents)).fetchone():
            print(f"ALERT: {product['product_id']} is {cents/100:.2f} {product['currency']}")
            db.execute("INSERT INTO alerts VALUES (?,?,?)", (product["product_id"], cents, now))
        if previous is not None and cents != previous:
            print(f"CHANGE: {previous/100:.2f} -> {cents/100:.2f} {product['currency']}")
        db.commit()


def main():
    init_db()
    for product in PRODUCTS:
        try:
            record(product, fetch_price(product))
        except Exception as exc:
            # Log a failure; never insert zero or a stale value as a new price.
            print(f"FAILED {product['product_id']}: {type(exc).__name__}: {exc}")

if __name__ == "__main__":
    main()

Run it with python price_tracker.py. The example intentionally treats a failed selector, blocked request, timeout, or malformed currency as a failure event. In production, write structured logs and a separate failure table so you can see when markup changes.

Make extraction trustworthy

Select the price and its identity together

Prefer a stable attribute, such as structured data or a retailer-provided product element, over a brittle positional selector. Confirm that the selected node belongs to the requested variant. A page can contain a “from” price, a subscription price, a crossed-out old price, shipping, tax, or prices for several variants.

Validate before comparison

  • Require the expected currency and a plausible range.
  • Reject empty, negative, NaN, and unexpectedly huge values.
  • Check the product title, SKU, or variant marker when available.
  • Record stock and promotion status if they affect the decision.
  • Keep the raw response or a redacted diagnostic sample where your policy permits, so a parser change can be audited.

Handle client-rendered pages honestly

requests sees server-delivered HTML only. If the price appears after JavaScript runs, use an approved retailer API/feed or a browser automation method that complies with the site’s rules. Do not “fix” a missing price by treating a bot-check page as a product page.

Scheduling, comparison, and alerts

There is no universal polling interval. Choose the least frequent schedule that meets your need and the retailer’s permitted request volume. A cron entry such as 15 * * * * /usr/bin/python3 /opt/tracker/price_tracker.py >> /var/log/price-tracker.log 2>&1 runs hourly at 15 minutes past the hour. A systemd timer, CI scheduler, or managed task can do the same.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare a validated observation with one of three explicit baselines:

  • Target: alert when price is at or below a user-defined amount.
  • Previous: alert on a decrease, with a minimum absolute or percentage change.
  • Reference: compare with a seven-day or other chosen historical value computed from stored rows.

Deduplicate alerts by recording the condition that was sent. Add a reset rule, such as allowing a new alert only after the price rises above the target and later falls again. Send email, a webhook, or a chat notification from the same decision point, but keep network failures from deleting the observation.

Common failures and fixes

Symptom Likely cause Fix
robots.txt denies the URL Your user agent or path is disallowed Stop; use an approved API/feed or request permission. Do not evade the rule.
HTTP 403, CAPTCHA, or “checking your browser” HTML Access control or bot detection Do not bypass it. Reduce requests only if permitted, or switch to an authorized source.
Selector matches nothing Markup changed or price is JavaScript-rendered Inspect current permitted HTML, update the selector with tests, or use an official API/browser workflow.
Price is zero or wildly high Wrong node, locale separators, or sale/old-price confusion Reject it, capture diagnostics, and add currency, range, and variant checks.
Every run sends an alert No alert state or reset logic Persist sent conditions using a unique key and define when they may re-arm.
History has gaps Scheduler stopped or failures were discarded Monitor exit status and logs; store failure events separately from successful observations.

Performance, reliability, and cost decisions

  • Use a session, connection reuse, and bounded timeouts for multiple permitted requests.
  • Space requests and cap concurrency; a faster scraper is not automatically an allowed scraper.
  • Cache unchanged pages only when the retailer’s rules permit it, and retain the observation timestamp even when content is cached.
  • Use retries sparingly for transient network errors, with exponential backoff and a maximum attempt count. Never retry permission failures or bot checks.
  • For many products, queue work, isolate one product’s failure from the batch, and expose metrics for success, parse failures, latency, and alert delivery.
  • Prices vary by location, currency, taxes, inventory, account, and promotion. Treat each row as an observation, not a guaranteed checkout total.

Or skip the browser setup

When a product page requires a rendered browser and you need a clean visual record for review, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF; it can wait for a selector or network idle, use custom headers and cookies, choose a device or viewport, and capture a selected element. It is not a substitute for permission to collect a retailer’s data, and a screenshot alone is not a structured price feed.

For a visual checkpoint after your tracker identifies a changed page, call:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for parameters and response headers. In plain terms, it accepts cookie/consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and headers report the page verdict and billing state. Its MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, and every feature is included on every plan.

Create a free ScreenshotNeo account to get the 1,000 no-card shots and add rendered-page evidence to your tracker.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Monetization warning for Amazon pages

If you plan to publish the tracker on a site using Amazon Associates, read the current Amazon Associates Operating Policies first. The policy states: “Unless otherwise agreed by Amazon, your Site must not have price tracking and/or price alerting functionality.” It also restricts use of Program Content and data-mining or similar extraction tools. An affiliate link or access to product content is not, by itself, permission to build a tracker. Obtain any required agreement or choose another permitted source.

Optional further reading

A sample for the physical book Website Scraping with Python Using BeautifulSoup is available from PocketBook. Verify the current edition, seller listing, and terms before buying; the sample does not establish a current listing or affiliate eligibility.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I scrape any page that is publicly visible?

No. Public visibility does not establish permission. Check the retailer’s terms, robots.txt, and any API or feed agreement for your intended use.

Should I store the whole HTML page for every observation?

Not necessarily. Store the normalized observation and diagnostics needed to audit parsing, while following the source’s terms and your privacy policy.

How often should the script run?

Use the least frequent interval that meets your alert requirement and stays within the retailer’s permitted request volume; no universal interval is established.

Why did my tracker record a sale price that was not available at checkout?

Prices can depend on variant, region, account, stock, tax, shipping, or promotion. Record those dimensions and treat scraped values as time-specific observations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.