Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape every product from an e-commerce site, first define exactly which catalog you mean, then build an inventory from the site’s product sitemap and any documented, authorized catalog API. Use category pagination to find gaps, fetch pages respectfully, and reconcile every discovered URL against the records you parsed. A crawler that merely follows links—or stops when it has “a lot” of products—cannot demonstrate completeness.

Only crawl pages you own or are authorized to access. Check the site’s terms, robots.txt rules, authentication boundaries, and rate limits before sending requests. The workflow below is for permitted catalog collection; it is not a way to bypass access controls.

Define what “every product” means

“Every product” needs a boundary before it can be measured. A store may have separate regional catalogs, languages, storefronts, product variants, or pages that are unpublished or unavailable. Record the scope before crawling, or a technically successful run can still produce the wrong dataset.

  • Host and paths: Specify the exact domain and which URL paths count as product pages. Decide whether subdomains or separate storefronts are included.
  • Region and language: State which locale, currency, and market you are collecting. A single storefront crawl does not establish coverage of other regional catalogs.
  • Variant policy: Decide whether each size, color, or other purchasable variant is a separate record, a child of a parent product, or both.
  • Availability policy: Decide whether out-of-stock, discontinued, preorder, and temporarily unavailable products belong in the inventory.
  • Stop condition: Choose a checkable endpoint, such as processing all in-scope sitemap URLs plus API pages through the documented total, or reaching the final pagination token.

Keep those decisions with the crawl run. They let you distinguish “we fetched everything in scope” from “we fetched everything the site exposes through one particular path.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Check authorization and crawler rules first

Confirm that you have permission to collect the pages and fields you need. Review the site’s terms and any applicable contractual or legal requirements, and do not use credentials or private endpoints outside their authorized purpose. Respect robots.txt and the site’s stated crawl limits. Robots directives help communicate crawler access rules; they are not, by themselves, proof of permission to collect or reuse data.

AWS says its Bedrock Web Crawler should be used only for pages the user owns or is authorized to crawl, and that it defaults to disallow when robots.txt is missing. Amazon’s Vendor Central documentation says AmazonProductDiscoverybot respects robots.txt and that changes to directives for that bot may take up to 24 hours to update. Those statements concern those named crawlers; do not assume every crawler or site behaves the same way.

Choose discovery sources in order of coverage

Use an inventory source before resorting to link walking. A sitemap can enumerate product URLs without requiring the crawler to guess which navigation paths lead to them. A documented catalog API can provide structured records and explicit pagination. Category and search pages are useful fallback discovery channels, but they may omit products because of filters, sorting, merchandising, or pagination behavior.

Source Best use What to verify
Product sitemap Build an initial list of public product URLs. Whether it is current, whether it includes all in-scope locales and product types, and whether it links to nested sitemaps.
Documented catalog API Retrieve structured records and traverse a defined collection. Authorization, scope, total count or continuation token, page-size limits, and whether variants are represented separately.
Category pagination Find products missing from the primary inventory, or use as a fallback when no complete inventory source is available. Every category, filter state, page token, and final-page condition. Avoid treating a repeated or empty page as proof of completeness unless that is the site’s documented behavior.
Search results and internal links Discover additional URLs when permitted and when the other sources leave gaps. Search indexing and navigation may not expose the whole catalog; record this route as a discovery method rather than treating it as a complete inventory.

Scrapy’s official SitemapSpider supports sitemap rules, nested sitemap handling, and sitemap discovery from robots.txt. Use those capabilities to route product URLs to the product parser, then compare the resulting URL set with other authorized inventory sources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Build a crawler that can account for its work

The following Scrapy starter reads a sitemap and parses product pages whose paths contain /product/. Replace the example host, sitemap, path rule, and selectors for the permitted site you are crawling. Sites vary: some use different URL patterns, server-rendered HTML, or structured product data. This code does not discover every possible site-specific pagination scheme or prove the sitemap is complete.

  1. Install Scrapy: Use a supported Python environment and install the package with python -m pip install scrapy.
  2. Create a project: Run scrapy startproject catalog_crawl, then save the spider below as catalog_crawl/spiders/products.py.
  3. Configure scope: Replace the example domain and sitemap, and set an identifying user agent. Start with low concurrency and a delay; follow any stricter site-specific limits.
  4. Run and retain the output: From the project directory, run scrapy crawl products -O products.jsonl. Keep the output tied to the run’s scope and timestamp.
import scrapy
from scrapy.spiders import SitemapSpider

class ProductsSpider(SitemapSpider):
    name = "products"
    allowed_domains = ["www.example.com"]
    sitemap_urls = ["https://www.example.com/sitemap.xml"]
    sitemap_rules = [
        (r"/product/", "parse_product"),
    ]

    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "USER_AGENT": "CatalogResearchBot/1.0 (contact: crawler@example.com)",
        "CONCURRENT_REQUESTS": 2,
        "DOWNLOAD_DELAY": 0.5,
    }

    def parse_product(self, response):
        canonical = response.css('link[rel="canonical"]::attr(href)').get()
        yield {
            "canonical_url": response.urljoin(canonical) if canonical else response.url,
            "product_id_or_sku": response.css('[itemprop="sku"]::attr(content)').get(),
            "title": response.css('h1::text').get(),
            "brand": response.css('[itemprop="brand"] [itemprop="name"]::text').get(),
            "price": response.css('[itemprop="price"]::attr(content)').get(),
            "currency": response.css('[itemprop="priceCurrency"]::attr(content)').get(),
            "availability": response.css('[itemprop="availability"]::attr(href)').get(),
            "image_url": response.css('[itemprop="image"]::attr(src)').get(),
            "http_status": response.status,
        }

The selector examples use common HTML attributes, not a universal e-commerce standard. If a field is absent, inspect an authorized response and adjust the parser; do not silently treat a missing value as evidence that the product lacks that attribute. Prefer a documented API or server-rendered product data when available. If the product details only appear after client-side execution, use browser rendering selectively and record which URLs needed it.

Traverse pagination without silently stopping early

For an authorized API, follow its documented page size, total, cursor, or next-page token until the documented end condition is reached. Scrapy.io documents offset-and-limit pagination with a maximum limit of 100 and an example that advances offsets until all rows have been collected. That limit applies to Scrapy.io’s documented API, not to every e-commerce API.

For category or search pages, keep a record of the category, filters, sort order, page or cursor, and response outcome. Continue until the site’s documented total or continuation mechanism says there is no next page. If no reliable total is exposed, treat a blank page or repeated set as a warning to investigate—not automatic proof of full coverage. Some sites cap results, alter pagination under filters, or return different products for different sort orders.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Extract a stable product record

Choose a schema before scaling the crawl. Keep identifiers and provenance alongside display fields so records can be compared across runs and traced back to their source.

  • Identity: canonical URL, product ID or SKU, and any parent or variant identifiers needed by your variant policy.
  • Product details: title, brand, category breadcrumbs, and image URLs.
  • Offer state: price, currency, and availability as observed. Keep their values as returned; do not assume prices in different currencies are directly comparable.
  • Provenance: source timestamp, HTTP status, discovery source, rendering path, and a reference to the raw response or payload.

Retaining raw responses lets you revise a parser without downloading every page again, where storage and site rules permit. Keep each observation with its capture time; prices and availability can change between requests.

Deduplicate, validate, and prove coverage

Normalize URLs consistently before comparing them: use canonical URLs where appropriate and remove tracking parameters that do not change product identity. Deduplicate on both canonical URL and product ID or SKU when available. Do not merge distinct variants merely because they share a parent URL or similar title.

Completeness is an accounting exercise. Track these sets for each run:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
  • Discovered: all in-scope product URLs from sitemaps, authorized APIs, and fallback discovery.
  • Fetched: URLs for which a request completed, with status and retry history.
  • Parsed: responses that produced records meeting required-field checks.
  • Failed or unresolved: URLs still blocked by errors, timeouts, disallowed paths, or parser gaps.

Reconcile the sets and report counts by source, status, and failure reason. Compare the final product identifiers against any documented inventory total, while accounting for the scope and variant policy you chose. A mismatch is a signal to investigate, not a number to hide by dropping failed URLs.

Schedule recrawls without losing change history

For recurring collection, store append-only crawl events as well as a current product view. Checkpoint after batches so an interrupted run can resume without starting from zero, and log retries rather than silently discarding failures. Set a recrawl schedule that matches your permitted use and how quickly the catalog changes. Define what counts as a change—such as price, stock status, title, or image URL—and retain the observation time for each change.

Hosted crawling services can reduce the work of running jobs and exporting datasets. Scrapy.io documents tool discovery, synchronous and asynchronous runs, polling, dataset retrieval, and schedules. AWS Bedrock Web Crawler documents sitemap seeds, authentication, crawl limits, scope controls, and incremental synchronization for teams using that service. Check each provider’s current documentation and access requirements before choosing it; those capabilities do not remove your responsibility to authorize the crawl and validate coverage.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you already have a list of product URLs and need visual captures for review, ScreenshotNeo can return a screenshot or PDF from one GET request. It is a capture API, not a catalog-discovery or product-data extraction system, so use your crawler or authorized API to find URLs and collect product fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

The request below captures one page as WebP. See the ScreenshotNeo API documentation for the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Troubleshoot common crawl failures

Symptom Likely cause What to check or do
Some products never appear in the output. The sitemap is incomplete or stale, pagination stopped early, or a locale/category was left out of scope. Compare discovery sources and page/cursor logs; verify sitemap coverage and the declared catalog boundary.
Pages return access errors or no useful content. The URL may be disallowed, require authorization, or be outside the permitted scope. Check the site’s rules and your authorization. Do not try to evade access controls; exclude the URL or request authorized access.
Product fields are blank although the page loads. The selectors do not match the site, fields are embedded in a different structure, or client-side code creates them. Inspect an authorized response, revise and validate the parser, or render only the pages that require browser execution.
The same product appears multiple times. Tracking parameters, alternate URLs, category paths, or variant URLs point to overlapping records. Normalize URLs, compare canonical URLs and identifiers, and apply the stated variant policy before merging.
The crawl stalls, times out, or triggers rate limits. Request volume may be too high, or the site may be slow or temporarily unavailable. Lower concurrency, honor stated limits, use backoff for transient failures, and checkpoint so completed batches are retained.
Record counts do not match a site total. The total may refer to a different region, filter, date, or variant definition—or the crawl has gaps. Align scope and counting rules, then reconcile discovered, fetched, parsed, and unresolved records before drawing a conclusion.

Frequently Asked Questions

Does a robots.txt file give me permission to reuse product data?

No. It communicates crawler access preferences; it does not replace authorization, the site’s terms, or any legal review applicable to your intended use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should variants count as separate products?

There is no universal answer. Set the rule to match your task and preserve parent and variant identifiers so your records can support either grouped or variant-level analysis.

Can screenshots substitute for product extraction?

No. A screenshot is a visual representation of a page. Use structured page data or an authorized API for fields such as SKU, price, and availability.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.