iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
When a store redesign breaks price extraction, first find out whether the price moved, the response changed, or the request is failing. Inspect the response your scraper receives, trace where the price is supplied, and then choose the matching extraction method: parse the response directly, reproduce a data request, or render the page in a browser. Validate the extracted prices after every run so a successful crawl cannot hide missing or incorrect data.
Diagnose the failure before changing selectors
A page that looks correct in a browser may not deliver the same content to your scraper. Start with the response received by the crawler, rather than assuming the retailer merely changed a CSS class. Scrapy’s guide to dynamically loaded content recommends inspecting the response and locating the source of data before choosing an extraction approach.
- Save and inspect the response as seen by the scraper. Scrapy’s
fetchcommand can save that response for examination. - Compare it with the browser page. Check whether the price is in the initial HTML, embedded JavaScript, or a separate network request.
- If a regular HTTP client receives the price but Scrapy does not, compare request details such as the user agent and headers.
- Check for redirects, server errors, inconsistent responses, or request blocking before treating the issue as a layout change. Scrapy notes that intermittent expected responses can indicate server problems, overload, or banning.
Choose an extraction method that matches the page
| What the page does | First approach | Trade-off |
|---|---|---|
| Price appears in the initial HTML | Use CSS or XPath selectors on the response. | Lightweight, but dependent on the response containing the value and the selector identifying the intended price. |
| Price arrives in a separate JSON or HTML request | Reproduce the request and parse its response. | Can provide structured data with less parsing and transfer, but may require the correct method, URL, headers, body, or form parameters. |
| Price appears only in the rendered page, or the underlying request is impractical to reproduce | Use a headless browser, such as Playwright. | Can access the rendered DOM, at the cost of browser execution and integration overhead. |
| The crawl completes but fields may be missing or wrong | Add field and record validation, monitoring, and alerts. | Helps detect silent failures; checks must reflect the retailer and its price formats. |
When the price is in the initial response
Use selectors against the response your scraper actually receives. Scrapy supports CSS and XPath selectors, and its selectors documentation describes an interactive shell for inspecting responses and trying expressions. Test the price field against current responses, including pages with sale prices, struck-through list prices, or other competing values.
Recommended Free Tools
Be careful with first-match extraction. Scrapy’s .get() returns the first match or None when there is no match; .getall() returns all matches. A selector that finds several prices should not silently return the first unless that is demonstrably the intended value. Check both zero matches and unexpected multiple matches.
#1 Best Overall
When a separate request supplies the price
Use browser network tools to identify the request that carries the price, then reproduce the necessary request details and parse its response as HTML, JSON, or another appropriate format. Scrapy’s documentation says that on pages fetching data from additional requests, reproducing the request containing the desired data is the preferred approach. This can avoid parsing a rendered page and may return more structured, complete data.
When browser rendering is necessary
Use a headless browser if the required value exists only in the rendered DOM or reproducing the relevant request is too difficult to do efficiently. Scrapy’s dynamic-content guidance discusses Playwright and recommends scrapy-playwright for better integration with Scrapy components than using Playwright in a way that bypasses components such as middleware and duplicate filtering.
For browser automation, Playwright recommends user-facing attributes and explicit locator contracts such as accessible roles. That guidance can help make locators more meaningful, but retailer markup is outside your control: a role-based locator is not guaranteed to expose a price uniquely or semantically. Narrow the locator with page context and validate the value it returns.
Build checks that catch silent price failures
A crawler can finish without an exception while extracting empty fields or the wrong price. Scrapy’s extensions page warns that “Spiders fail quietly in production” and describes Spidermon for monitoring, validation, and alerts. Add checks suited to each store and product set:
- Fail or alert when a required price is absent, cannot be parsed, or appears an unexpected number of times.
- Track product counts and the share of records with valid prices. Alert when they drop meaningfully from that crawler’s established baseline; there is no universal threshold.
- Check currency and plausible changes against the product and its earlier observations, rather than accepting any parseable number.
- Keep the URL, timestamp, and a sample response or diagnostic artifact where permitted, so a failure can be reproduced.
- Keep each retailer’s selectors and price transformations isolated from request scheduling and storage logic. Scrapy describes
scrapy-poetpage objects as a way to separate extraction from parsing for testing and reuse.
Compare candidate methods by completeness and correctness, maintenance effort, runtime and network cost, and observability. No cited source benchmarks them on a particular retailer, so the right choice depends on that site’s response and your crawler’s needs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Verify store-specific constraints
Technical documentation cannot establish whether a particular retailer exposes a suitable endpoint, what its current page structure is, or whether your planned crawling is permitted under that site’s terms. Verify those details for the target store before implementation; the general guidance here is not permission to scrape a specific site.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

