Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape every product in an e-commerce category, first identify the product-card HTML, then follow a real pagination URL or permitted data endpoint until the catalog is exhausted. Use an HTTP client with BeautifulSoup or Scrapy when cards are in the initial response; use Playwright only when JavaScript is required. At every step, respect robots.txt, the site’s terms, rate limits, privacy obligations, and applicable law.

A reliable scraper records a stable product URL and identifier alongside title, price, currency, availability, image URL, category path, and crawl timestamp. It normalizes prices, deduplicates records, logs failures, and stops at a configured page limit rather than crawling without bounds.

Define the dataset and crawl boundary first

Write down exactly what one product record contains and where the crawl starts and ends. A typical category-page record includes:

  • Canonical product URL
  • Title or product name
  • SKU, product ID, or another exposed stable identifier
  • Price as a numeric value and its currency
  • Availability or stock label
  • Primary image URL
  • Category and subcategory path
  • UTC crawl timestamp

Also set the category URLs, maximum page count, refresh interval, and whether variants are separate records. A page limit is a safety control for broken pagination and unexpectedly large catalogs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check access rules before sending requests

Fetch and read the store’s robots.txt. Configure your crawler to obey it and identify yourself with a descriptive user agent. Google’s guidance explains that robots.txt manages crawler traffic and is not a way to hide URLs from search results; it is also not a complete permission grant. Read the Google robots.txt guide before collecting data.

Separately review terms of service, authentication barriers, rate limits, privacy requirements, copyright and database rights, and any contract that governs the data. Do not bypass logins, CAPTCHAs, bot checks, paywalls, or other access controls. Collect only what you need, store it securely, and obtain permission when the publisher requires it.

Choose the least complex method that works

Situation Approach Trade-off
Cards and next links are in the initial HTML HTTP client plus BeautifulSoup, lxml, or Scrapy selectors Fast and inexpensive, but it cannot see content inserted by JavaScript
Many categories, retries, and scheduled refreshes Scrapy spider with item pipelines and persistent job state Strong crawl control, but more framework setup
Prices or cards appear after JavaScript actions Find a permitted JSON endpoint first; otherwise use Playwright or another browser renderer Higher fidelity, with more CPU, memory, and latency
A sitemap or merchant feed lists complete product URLs Discover URLs from the sitemap or feed, then request product pages selectively Efficient discovery, although feed fields may differ from page fields

Start with a normal category URL and inspect its response. If the product cards and a real <a href> for the next page are present, a browser is unnecessary.

Build a bounded Python scraper with Requests and BeautifulSoup

Install the dependencies:

python -m pip install requests beautifulsoup4

The following script follows ordinary next-page links, retries transient failures, obeys robots.txt, normalizes basic fields, and deduplicates by canonical URL. Replace the selectors for the store you are allowed to crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from __future__ import annotations

import json
import re
import time
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from urllib.parse import urljoin, urlparse, urldefrag
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

START_URL = "https://example.com/category/widgets"
MAX_PAGES = 50
USER_AGENT = "ExampleCatalogResearchBot/1.0 (+https://example.com/contact)"

# Change these selectors to match the permitted site.
CARD = "article.product-card"
TITLE = ".product-title"
PRICE = ".price"
AVAILABILITY = ".availability"
IMAGE = "img"
NEXT = "a[rel='next'], a.next"


def text_or_none(node):
    return node.get_text(" ", strip=True) if node else None


def canonical_url(base, href):
    if not href:
        return None
    absolute = urljoin(base, href)
    absolute, _ = urldefrag(absolute)
    parsed = urlparse(absolute)
    # Keep query parameters only if this store uses them for pagination or identity.
    return parsed._replace(fragment="").geturl()


def parse_price(raw):
    if not raw:
        return None, None
    value = re.sub(r"[^0-9,. -]", "", raw).strip().replace(" ", "")
    # Adapt this rule for the store's locale; this handles common 1,234.56 and 1.234,56 forms.
    if value.count(",") == 1 and value.count(".") == 0:
        value = value.replace(",", ".")
    elif value.count(",") and value.count("."):
        value = value.replace(".", "").replace(",", ".")
    else:
        value = value.replace(",", "")
    try:
        amount = str(Decimal(value))
    except InvalidOperation:
        amount = None
    currency = None
    symbol = re.search(r"(USD|EUR|GBP|CAD|AUD|[$€£])", raw, re.I)
    if symbol:
        currency = {"$": "USD", "€": "EUR", "£": "GBP"}.get(symbol.group(1), symbol.group(1).upper())
    return amount, currency


def robots_allows(url):
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    rp = RobotFileParser(robots_url)
    try:
        rp.read()
        return rp.can_fetch(USER_AGENT, url)
    except Exception:
        # A failed robots fetch is a reason to stop or obtain a policy decision,
        # not a reason to crawl aggressively.
        return False


def make_session():
    retry = Retry(
        total=4,
        backoff_factor=1,
        status_forcelist=(429, 500, 502, 503, 504),
        allowed_methods=frozenset(["GET"]),
        respect_retry_after_header=True,
    )
    session = requests.Session()
    session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
    session.mount("https://", HTTPAdapter(max_retries=retry))
    session.mount("http://", HTTPAdapter(max_retries=retry))
    return session


def scrape_category(start_url):
    if not robots_allows(start_url):
        raise RuntimeError("robots.txt does not allow this URL or could not be read")
    session = make_session()
    current = start_url
    seen_urls = set()
    seen_products = set()
    records = []

    for page_number in range(1, MAX_PAGES + 1):
        if not current or current in seen_urls:
            break
        seen_urls.add(current)
        response = session.get(current, timeout=(10, 30))
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")
        cards = soup.select(CARD)
        if not cards:
            raise RuntimeError(f"No product cards found on page {current}; check selectors or JavaScript rendering")

        page_ids = set()
        for card in cards:
            link = card.select_one("a[href]")
            product_url = canonical_url(current, link.get("href") if link else None)
            if not product_url or product_url in seen_products:
                continue
            seen_products.add(product_url)
            product_id = card.get("data-product-id") or card.get("data-sku")
            title = text_or_none(card.select_one(TITLE))
            raw_price = text_or_none(card.select_one(PRICE))
            price, currency = parse_price(raw_price)
            image = card.select_one(IMAGE)
            image_url = canonical_url(current, image.get("src") or image.get("data-src")) if image else None
            record = {
                "url": product_url,
                "id": product_id,
                "title": title,
                "price": price,
                "currency": currency,
                "availability": text_or_none(card.select_one(AVAILABILITY)),
                "image_url": image_url,
                "category_url": start_url,
                "crawl_timestamp": datetime.now(timezone.utc).isoformat(),
            }
            records.append(record)
            page_ids.add(product_url)

        next_node = soup.select_one(NEXT)
        next_url = canonical_url(current, next_node.get("href")) if next_node else None
        if not next_url or next_url in seen_urls or not page_ids:
            break
        time.sleep(1.0)  # Set a delay appropriate to the publisher's policy.
        current = next_url

    return records


if __name__ == "__main__":
    data = scrape_category(START_URL)
    with open("products.json", "w", encoding="utf-8") as output:
        json.dump(data, output, ensure_ascii=False, indent=2)
    print(f"Wrote {len(data)} unique products")

The selectors are deliberately store-specific. Inspect one card in your browser’s developer tools, identify a stable class or data-* attribute, and test against a saved HTML fixture before crawling the full category. Do not use a selector that depends on a visual position or a frequently changing CSS class when a product ID or semantic attribute is available.

Discover every page without guessing

Follow ordinary pagination links

Prefer a real next-page link or a documented request pattern. Continue until the link disappears, product IDs stop changing, or your configured maximum is reached. Canonicalize URLs and track every page already visited so a malformed “next” link cannot create a loop. Google recommends unique URLs for paginated sequences and warns that URL fragments are not reliable page numbers; see its pagination guidance.

Inspect load-more requests

For a “Load more” control, open browser developer tools, select the Network panel, and activate the control once. Determine whether the page requests JSON, HTML fragments, or a GraphQL operation. If the endpoint is publicly accessible and permitted by the site’s rules, request it directly with the same required parameters, then increment its cursor or page value. Validate that each response contains new product IDs.

Handle infinite scroll

Infinite scroll is a presentation pattern, not a data source. Look for the underlying request and a cursor in the response. If no stable endpoint exists and products appear only after JavaScript execution, use a browser renderer as a slower fallback. Google’s documentation notes that crawlers do not click buttons and generally do not trigger JavaScript functions requiring user actions.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Scrapy when the crawl becomes a project

Scrapy is appropriate when you need many categories, persistent jobs, pipelines, throttling, and structured retries. A minimal spider looks like this:

import scrapy

class CategorySpider(scrapy.Spider):
    name = "category"
    start_urls = ["https://example.com/category/widgets"]
    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "USER_AGENT": "ExampleCatalogResearchBot/1.0 (+https://example.com/contact)",
        "DOWNLOAD_DELAY": 1.0,
        "AUTOTHROTTLE_ENABLED": True,
    }

    def parse(self, response):
        for card in response.css("article.product-card"):
            href = card.css("a[href]::attr(href)").get()
            yield {
                "url": response.urljoin(href) if href else None,
                "title": card.css(".product-title ::text").get(default="").strip(),
                "price": card.css(".price ::text").get(default="").strip(),
                "availability": card.css(".availability ::text").get(default="").strip(),
            }
        next_href = response.css("a[rel='next']::attr(href)").get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

Scrapy describes spiders as components that generate requests, parse responses, and return structured items. Its spider documentation and selector documentation cover item pipelines, feeds, and XPath or CSS selectors.

When prices are JavaScript-rendered

First determine whether the initial HTML contains a JSON-LD block, embedded state object, or permitted JSON request with the price. Parsing that structured response is usually faster and more stable than rendering a page. Confirm that the endpoint is allowed, preserve required pagination cursors, and do not replay private or authenticated requests without authorization.

If a browser is necessary, use Playwright to open the category, wait for the product-card selector, scroll or activate the permitted control, and extract the rendered DOM. Set a finite number of scroll cycles and a maximum item count. Browser automation consumes substantially more memory and time, so reserve it for pages that cannot be collected through HTML or a permitted endpoint.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize, deduplicate, and validate the output

Normalize values

  • Store a numeric price and an explicit ISO-style currency code; retain the original text for audit.
  • Resolve relative URLs and remove fragments. Keep query parameters only when they identify a product or pagination state.
  • Map availability labels such as “in stock,” “sold out,” and “preorder” to your own controlled vocabulary while retaining the source label.
  • Keep variant IDs when color, size, or pack options represent separate offers.

Deduplicate safely

Use SKU or another stable product ID when it is exposed. Otherwise use the canonical product URL. Do not deduplicate on title alone: two products can share a name, and one product can have multiple legitimate variants.

Measure scraper health

Log page URLs, HTTP status codes, response times, retry counts, and parser exceptions. Track missing-field rates, duplicate rates, page counts, and the number of new product IDs per page. Save a small fixture of representative category pages so selector changes can be tested before deployment. A sudden zero-product page, a large increase in missing prices, or a changed card count should stop or quarantine the run rather than silently producing incomplete data.

Performance, reliability, and cost controls

  • Use connection reuse, finite timeouts, exponential backoff, and a delay or concurrency limit that fits the publisher’s policy.
  • Cache responses during development and scheduled refreshes; avoid downloading unchanged pages repeatedly.
  • Persist crawl state so a failed run resumes from the last confirmed page instead of restarting.
  • Set hard limits for pages, products, response bytes, and browser scrolls.
  • Prefer sitemap or feed discovery when a complete catalog is published; it can avoid crawling every category page.
  • Store raw response metadata for a limited retention period so parser failures can be diagnosed without retaining unnecessary personal data.

Common failures and fixes

“No product cards found”

The selector may be wrong, the template may have changed, or cards may be injected by JavaScript. Save the response, inspect its HTML, and compare it with a browser’s rendered DOM. If the response is only a shell, find a permitted JSON endpoint or switch to Playwright.

The scraper repeats the same page

Some sites emit a next link with an unchanged URL or a tracking parameter that creates a loop. Canonicalize URLs, keep a visited set, and stop when the next URL has already been seen or product IDs no longer change.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prices are blank or incorrect

Prices may be in a data-price attribute, JSON-LD, a locale-specific string, or a variant selector. Extract the authoritative value, retain currency, and test thousands and decimal separators for each locale you crawl.

HTTP 403 or 429 responses

Stop increasing concurrency. Check the site’s rules, reduce request frequency, honor Retry-After, and request permission if required. Do not attempt to evade an access control or bot challenge.

Only the first batch appears

“Load more” and infinite-scroll pages commonly return a cursor in a network response. Capture that request and follow the cursor rather than repeatedly downloading the first HTML document. If the endpoint is not permitted or cannot be made stable, use a bounded browser workflow or omit the category.

Duplicate products appear across categories

This is expected when a product belongs to multiple categories. Keep a category-to-product relation, but deduplicate the product entity by SKU or canonical URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It can capture a category page after accepting the cookie or consent banner and removing more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It is useful when you need visual evidence of each category page rather than parsed product fields.

One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work.

See the ScreenshotNeo documentation for the complete option list. Replace the example URL with a category URL you are allowed to capture.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

An MCP server provides take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients, so an AI agent can inspect pages without you maintaining browser code. Pricing is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Allowance Price
Free 1,000 shots/month $0, no card
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

Frequently Asked Questions

Can I scrape a category page that requires a login?

Only if you have explicit authorization and the site’s terms permit the collection. Keep credentials out of logs, follow the publisher’s access controls, and do not bypass authentication.

How often should a category scraper run?

Choose a refresh interval from the business need and the site’s published limits. Start with the least frequent schedule that keeps your dataset useful, then adjust using observed change rates and server responses.

Should product variants be separate records?

Make that decision at the schema stage: use separate records when variants have distinct identifiers, prices, or availability; otherwise retain variant attributes under one product while preserving the source identifiers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

Use Requests plus BeautifulSoup or Scrapy for server-rendered category pages, follow real pagination or permitted JSON cursors, and reserve Playwright for JavaScript-only content. Reliable results come from bounded crawling, respectful access, stable identifiers, normalization, deduplication, and validation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.