Scrape product catalogs in two stages: parse category or search-result pages to collect product URLs and summary fields, then visit each product’s detail page for richer attributes. Follow the site’s actual pagination, use separate parsers for listing and detail pages, and verify that missing or blocked responses are not mistaken for valid product data. Before crawling, define the permitted scope and review the target site’s current terms and crawl guidance; technical accessibility alone does not establish permission.
Plan the crawl and inspect the pages
Start with the smallest allowed set of category or search-result URLs. Decide which fields you need, why you need them, and how often the data must be refreshed. Permission depends on the target site, intended use, and applicable jurisdiction; the technical steps here do not determine those questions.
Inspect one representative listing page and one product detail page. Compare the HTML returned to an HTTP client with what appears in the browser. Identify stable product links or identifiers, repeated card structure, pagination controls, and fields that appear only after scripts run. Selectors are site-specific, so do not assume a generic product-card selector will work across sites.
Discover URLs and parse listing pages
Listing pages typically provide candidate product URLs and summary fields such as a displayed name or price. Use the product URL or a stable identifier to join listing records to detail records and deduplicate products.
#1 Best Overall
You can begin with explicitly selected category URLs or use a sitemap as a discovery input. Scrapy’s SitemapSpider supports sitemap URLs, can discover sitemap locations from robots.txt, and can route product and category URL patterns to different callbacks. Sitemap inclusion helps discover candidate URLs; it does not establish permission for every use.
For each listing response, select the repeated product card, extract the product link and any needed summary values, then resolve the link to an absolute URL. Follow the page’s real next-page link and stop when it is absent. Guard against repeated URLs or pagination loops. Scrapy’s tutorial demonstrates extracting items and yielding a request for the next page.
Parse product detail pages separately
Send each discovered product URL to a detail-page callback with its own selectors. Extract only fields present and relevant to your use, such as the product name, brand, SKU, description, price, availability, or variant choices. Store unavailable values as missing rather than inferring them. Keep a stable URL or identifier so detail records can be matched to listing records.
Scrapy selectors integrate with responses and use CSS or XPath against HTML; see the Scrapy selectors guide. Validate selectors on multiple representative pages, since a single product may not expose every field or variant found elsewhere in the catalog.
Recommended Free Tools
Rank #3
Choose the extraction method for dynamic fields
If a field is visible in a browser but absent from the downloaded response, inspect the page source and the browser’s network requests. The data may be embedded in JavaScript or supplied by a separate request. Scrapy’s dynamic-content guidance recommends finding the underlying data source; reproducing that request can return structured data without parsing a rendered page.
| Method | Use it when | Trade-off |
|---|---|---|
| Parse the initial HTML response | The required fields are already present in the response. | CSS or XPath selectors are direct, but they cannot extract data that the response does not contain. |
| Reproduce the data request | An embedded source or separate request supplies the missing field, and its URL, method, headers, body, or form parameters can be identified. | It can provide structured data with less parsing and transfer than a browser, but the request details may need maintenance. |
| Use a headless browser | The required state exists only in a rendered DOM, depends on interaction, or reproducing the source request is impractical. | It adds browser setup and resource use compared with parsing HTML or requesting a data source directly. |
Prefer the least complex method that returns the required fields. A screenshot can show what was rendered, but it is not a substitute for extracting structured product data when an underlying response or request provides it.
Validate results and handle failures
Before treating a crawl as complete, check listing counts, URL uniqueness, required-field presence, representative extracted values, and whether each detail record matches a listing record. Keep URLs and timestamps when useful to the project, and make the output structure explicit. Scrapy callbacks can yield items and further requests, with items handled through pipelines or feed exports, as described in its spider documentation.
- A field is missing: Confirm whether it is absent from the response or only missing because the selector does not match. If it appears in the browser, inspect the page source and network requests before moving to browser automation.
- The page appears empty or blocked: Do not record the response as a valid product. Inspect the response and determine whether it is a challenge, failed load, or genuinely empty page; follow the site’s permitted access guidance.
- Pagination repeats or does not end: Resolve next links to absolute URLs, track visited listing URLs, and stop when the next link is absent or already visited.
- Listing and detail fields disagree: Keep the page type and collection time associated with each record, then decide which source is authoritative for each field rather than silently overwriting values.
Or skip the browser setup
For rendered-page screenshots rather than structured catalog extraction, ScreenshotNeo provides a website screenshot API. It can also help inspect how a listing or detail page renders, but a screenshot does not replace extracting product fields from HTML or a data request.
Best Value
One GET request returns a screenshot or PDF; for example, this cURL request saves a WebP screenshot of a product listing URL. See the ScreenshotNeo API documentation for parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/category -o shot.webp
Before capture, ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers screenshot and PDF tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month with no card.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

