Recommended Free Tools
To scrape every product in an e-commerce category, first identify the product-card HTML, then follow a real pagination URL or permitted data endpoint until the catalog is exhausted. Use an HTTP client with BeautifulSoup or Scrapy when cards are in the initial response; use Playwright only when JavaScript is required. At every step, respect robots.txt, the site’s terms, rate limits, privacy obligations, and applicable law.
A reliable scraper records a stable product URL and identifier alongside title, price, currency, availability, image URL, category path, and crawl timestamp. It normalizes prices, deduplicates records, logs failures, and stops at a configured page limit rather than crawling without bounds.
Define the dataset and crawl boundary first
Write down exactly what one product record contains and where the crawl starts and ends. A typical category-page record includes:
- Canonical product URL
- Title or product name
- SKU, product ID, or another exposed stable identifier
- Price as a numeric value and its currency
- Availability or stock label
- Primary image URL
- Category and subcategory path
- UTC crawl timestamp
Also set the category URLs, maximum page count, refresh interval, and whether variants are separate records. A page limit is a safety control for broken pagination and unexpectedly large catalogs.
Check access rules before sending requests
Fetch and read the store’s robots.txt. Configure your crawler to obey it and identify yourself with a descriptive user agent. Google’s guidance explains that robots.txt manages crawler traffic and is not a way to hide URLs from search results; it is also not a complete permission grant. Read the Google robots.txt guide before collecting data.
Separately review terms of service, authentication barriers, rate limits, privacy requirements, copyright and database rights, and any contract that governs the data. Do not bypass logins, CAPTCHAs, bot checks, paywalls, or other access controls. Collect only what you need, store it securely, and obtain permission when the publisher requires it.
Choose the least complex method that works
| Situation | Approach | Trade-off |
|---|---|---|
| Cards and next links are in the initial HTML | HTTP client plus BeautifulSoup, lxml, or Scrapy selectors | Fast and inexpensive, but it cannot see content inserted by JavaScript |
| Many categories, retries, and scheduled refreshes | Scrapy spider with item pipelines and persistent job state | Strong crawl control, but more framework setup |
| Prices or cards appear after JavaScript actions | Find a permitted JSON endpoint first; otherwise use Playwright or another browser renderer | Higher fidelity, with more CPU, memory, and latency |
| A sitemap or merchant feed lists complete product URLs | Discover URLs from the sitemap or feed, then request product pages selectively | Efficient discovery, although feed fields may differ from page fields |
Start with a normal category URL and inspect its response. If the product cards and a real <a href> for the next page are present, a browser is unnecessary.
Build a bounded Python scraper with Requests and BeautifulSoup
Install the dependencies:
python -m pip install requests beautifulsoup4
The following script follows ordinary next-page links, retries transient failures, obeys robots.txt, normalizes basic fields, and deduplicates by canonical URL. Replace the selectors for the store you are allowed to crawl.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutefrom __future__ import annotations
import json
import re
import time
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from urllib.parse import urljoin, urlparse, urldefrag
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
START_URL = "https://example.com/category/widgets"
MAX_PAGES = 50
USER_AGENT = "ExampleCatalogResearchBot/1.0 (+https://example.com/contact)"
# Change these selectors to match the permitted site.
CARD = "article.product-card"
TITLE = ".product-title"
PRICE = ".price"
AVAILABILITY = ".availability"
IMAGE = "img"
NEXT = "a[rel='next'], a.next"
def text_or_none(node):
return node.get_text(" ", strip=True) if node else None
def canonical_url(base, href):
if not href:
return None
absolute = urljoin(base, href)
absolute, _ = urldefrag(absolute)
parsed = urlparse(absolute)
# Keep query parameters only if this store uses them for pagination or identity.
return parsed._replace(fragment="").geturl()
def parse_price(raw):
if not raw:
return None, None
value = re.sub(r"[^0-9,. -]", "", raw).strip().replace(" ", "")
# Adapt this rule for the store's locale; this handles common 1,234.56 and 1.234,56 forms.
if value.count(",") == 1 and value.count(".") == 0:
value = value.replace(",", ".")
elif value.count(",") and value.count("."):
value = value.replace(".", "").replace(",", ".")
else:
value = value.replace(",", "")
try:
amount = str(Decimal(value))
except InvalidOperation:
amount = None
currency = None
symbol = re.search(r"(USD|EUR|GBP|CAD|AUD|[$€£])", raw, re.I)
if symbol:
currency = {"$": "USD", "€": "EUR", "£": "GBP"}.get(symbol.group(1), symbol.group(1).upper())
return amount, currency
def robots_allows(url):
parsed = urlparse(url)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
rp = RobotFileParser(robots_url)
try:
rp.read()
return rp.can_fetch(USER_AGENT, url)
except Exception:
# A failed robots fetch is a reason to stop or obtain a policy decision,
# not a reason to crawl aggressively.
return False
def make_session():
retry = Retry(
total=4,
backoff_factor=1,
status_forcelist=(429, 500, 502, 503, 504),
allowed_methods=frozenset(["GET"]),
respect_retry_after_header=True,
)
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))
return session
def scrape_category(start_url):
if not robots_allows(start_url):
raise RuntimeError("robots.txt does not allow this URL or could not be read")
session = make_session()
current = start_url
seen_urls = set()
seen_products = set()
records = []
for page_number in range(1, MAX_PAGES + 1):
if not current or current in seen_urls:
break
seen_urls.add(current)
response = session.get(current, timeout=(10, 30))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
cards = soup.select(CARD)
if not cards:
raise RuntimeError(f"No product cards found on page {current}; check selectors or JavaScript rendering")
page_ids = set()
for card in cards:
link = card.select_one("a[href]")
product_url = canonical_url(current, link.get("href") if link else None)
if not product_url or product_url in seen_products:
continue
seen_products.add(product_url)
product_id = card.get("data-product-id") or card.get("data-sku")
title = text_or_none(card.select_one(TITLE))
raw_price = text_or_none(card.select_one(PRICE))
price, currency = parse_price(raw_price)
image = card.select_one(IMAGE)
image_url = canonical_url(current, image.get("src") or image.get("data-src")) if image else None
record = {
"url": product_url,
"id": product_id,
"title": title,
"price": price,
"currency": currency,
"availability": text_or_none(card.select_one(AVAILABILITY)),
"image_url": image_url,
"category_url": start_url,
"crawl_timestamp": datetime.now(timezone.utc).isoformat(),
}
records.append(record)
page_ids.add(product_url)
next_node = soup.select_one(NEXT)
next_url = canonical_url(current, next_node.get("href")) if next_node else None
if not next_url or next_url in seen_urls or not page_ids:
break
time.sleep(1.0) # Set a delay appropriate to the publisher's policy.
current = next_url
return records
if __name__ == "__main__":
data = scrape_category(START_URL)
with open("products.json", "w", encoding="utf-8") as output:
json.dump(data, output, ensure_ascii=False, indent=2)
print(f"Wrote {len(data)} unique products")
The selectors are deliberately store-specific. Inspect one card in your browser’s developer tools, identify a stable class or data-* attribute, and test against a saved HTML fixture before crawling the full category. Do not use a selector that depends on a visual position or a frequently changing CSS class when a product ID or semantic attribute is available.
Discover every page without guessing
Follow ordinary pagination links
Prefer a real next-page link or a documented request pattern. Continue until the link disappears, product IDs stop changing, or your configured maximum is reached. Canonicalize URLs and track every page already visited so a malformed “next” link cannot create a loop. Google recommends unique URLs for paginated sequences and warns that URL fragments are not reliable page numbers; see its pagination guidance.
Inspect load-more requests
For a “Load more” control, open browser developer tools, select the Network panel, and activate the control once. Determine whether the page requests JSON, HTML fragments, or a GraphQL operation. If the endpoint is publicly accessible and permitted by the site’s rules, request it directly with the same required parameters, then increment its cursor or page value. Validate that each response contains new product IDs.
Handle infinite scroll
Infinite scroll is a presentation pattern, not a data source. Look for the underlying request and a cursor in the response. If no stable endpoint exists and products appear only after JavaScript execution, use a browser renderer as a slower fallback. Google’s documentation notes that crawlers do not click buttons and generally do not trigger JavaScript functions requiring user actions.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use Scrapy when the crawl becomes a project
Scrapy is appropriate when you need many categories, persistent jobs, pipelines, throttling, and structured retries. A minimal spider looks like this:
import scrapy
class CategorySpider(scrapy.Spider):
name = "category"
start_urls = ["https://example.com/category/widgets"]
custom_settings = {
"ROBOTSTXT_OBEY": True,
"USER_AGENT": "ExampleCatalogResearchBot/1.0 (+https://example.com/contact)",
"DOWNLOAD_DELAY": 1.0,
"AUTOTHROTTLE_ENABLED": True,
}
def parse(self, response):
for card in response.css("article.product-card"):
href = card.css("a[href]::attr(href)").get()
yield {
"url": response.urljoin(href) if href else None,
"title": card.css(".product-title ::text").get(default="").strip(),
"price": card.css(".price ::text").get(default="").strip(),
"availability": card.css(".availability ::text").get(default="").strip(),
}
next_href = response.css("a[rel='next']::attr(href)").get()
if next_href:
yield response.follow(next_href, callback=self.parse)
Scrapy describes spiders as components that generate requests, parse responses, and return structured items. Its spider documentation and selector documentation cover item pipelines, feeds, and XPath or CSS selectors.
When prices are JavaScript-rendered
First determine whether the initial HTML contains a JSON-LD block, embedded state object, or permitted JSON request with the price. Parsing that structured response is usually faster and more stable than rendering a page. Confirm that the endpoint is allowed, preserve required pagination cursors, and do not replay private or authenticated requests without authorization.
If a browser is necessary, use Playwright to open the category, wait for the product-card selector, scroll or activate the permitted control, and extract the rendered DOM. Set a finite number of scroll cycles and a maximum item count. Browser automation consumes substantially more memory and time, so reserve it for pages that cannot be collected through HTML or a permitted endpoint.
Normalize, deduplicate, and validate the output
Normalize values
- Store a numeric price and an explicit ISO-style currency code; retain the original text for audit.
- Resolve relative URLs and remove fragments. Keep query parameters only when they identify a product or pagination state.
- Map availability labels such as “in stock,” “sold out,” and “preorder” to your own controlled vocabulary while retaining the source label.
- Keep variant IDs when color, size, or pack options represent separate offers.
Deduplicate safely
Use SKU or another stable product ID when it is exposed. Otherwise use the canonical product URL. Do not deduplicate on title alone: two products can share a name, and one product can have multiple legitimate variants.
Measure scraper health
Log page URLs, HTTP status codes, response times, retry counts, and parser exceptions. Track missing-field rates, duplicate rates, page counts, and the number of new product IDs per page. Save a small fixture of representative category pages so selector changes can be tested before deployment. A sudden zero-product page, a large increase in missing prices, or a changed card count should stop or quarantine the run rather than silently producing incomplete data.
Performance, reliability, and cost controls
- Use connection reuse, finite timeouts, exponential backoff, and a delay or concurrency limit that fits the publisher’s policy.
- Cache responses during development and scheduled refreshes; avoid downloading unchanged pages repeatedly.
- Persist crawl state so a failed run resumes from the last confirmed page instead of restarting.
- Set hard limits for pages, products, response bytes, and browser scrolls.
- Prefer sitemap or feed discovery when a complete catalog is published; it can avoid crawling every category page.
- Store raw response metadata for a limited retention period so parser failures can be diagnosed without retaining unnecessary personal data.
Common failures and fixes
“No product cards found”
The selector may be wrong, the template may have changed, or cards may be injected by JavaScript. Save the response, inspect its HTML, and compare it with a browser’s rendered DOM. If the response is only a shell, find a permitted JSON endpoint or switch to Playwright.
The scraper repeats the same page
Some sites emit a next link with an unchanged URL or a tracking parameter that creates a loop. Canonicalize URLs, keep a visited set, and stop when the next URL has already been seen or product IDs no longer change.
Free tools Windows power users keep installed
One-click scans. No signup required.
Prices are blank or incorrect
Prices may be in a data-price attribute, JSON-LD, a locale-specific string, or a variant selector. Extract the authoritative value, retain currency, and test thousands and decimal separators for each locale you crawl.
HTTP 403 or 429 responses
Stop increasing concurrency. Check the site’s rules, reduce request frequency, honor Retry-After, and request permission if required. Do not attempt to evade an access control or bot challenge.
Only the first batch appears
“Load more” and infinite-scroll pages commonly return a cursor in a network response. Capture that request and follow the cursor rather than repeatedly downloading the first HTML document. If the endpoint is not permitted or cannot be made stable, use a bounded browser workflow or omit the category.
Duplicate products appear across categories
This is expected when a product belongs to multiple categories. Keep a category-to-product relation, but deduplicate the product entity by SKU or canonical URL.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBest Value
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It can capture a category page after accepting the cookie or consent banner and removing more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It is useful when you need visual evidence of each category page rather than parsed product fields.
One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work.
See the ScreenshotNeo documentation for the complete option list. Replace the example URL with a category URL you are allowed to capture.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
An MCP server provides take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients, so an AI agent can inspect pages without you maintaining browser code. Pricing is:
| Plan | Allowance | Price |
|---|---|---|
| Free | 1,000 shots/month | $0, no card |
| Starter | 3,000 shots | $5 |
| Growth | 15,000 shots | $15 |
| Pro | 60,000 shots | $39 |
| Scale | 250,000 shots | $99 |
| Business | 1,000,000 shots | $249 |
Yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
Frequently Asked Questions
Can I scrape a category page that requires a login?
Only if you have explicit authorization and the site’s terms permit the collection. Keep credentials out of logs, follow the publisher’s access controls, and do not bypass authentication.
How often should a category scraper run?
Choose a refresh interval from the business need and the site’s published limits. Start with the least frequent schedule that keeps your dataset useful, then adjust using observed change rates and server responses.
Should product variants be separate records?
Make that decision at the schema stage: use separate records when variants have distinct identifiers, prices, or availability; otherwise retain variant attributes under one product while preserving the source identifiers.
The Bottom Line
Use Requests plus BeautifulSoup or Scrapy for server-rendered category pages, follow real pagination or permitted JSON cursors, and reserve Playwright for JavaScript-only content. Reliable results come from bounded crawling, respectful access, stable identifiers, normalization, deduplication, and validation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

