Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can scrape search-result HTML with Python’s requests and BeautifulSoup only when the site permits automated access. Check Amazon’s robots.txt and terms first, identify yourself, send requests slowly, stop on 403, 429, 503, CAPTCHA, or robot-check responses, and keep a strict page limit. The example below uses example.com and generic selectors deliberately; Amazon’s customer-facing markup changes and its bot documentation does not grant permission to collect search pages.

1. Confirm that automated access is allowed

Before writing a scraper, decide whether you have a lawful, permissioned source. Read the target locale’s terms and its robots.txt. If a path is disallowed or the terms prohibit automated access, stop and use an official API, a permitted export, or a licensed data provider instead. Amazon’s documentation describes Amazonbot, Amzn-SearchBot, and Amzn-User and how those crawlers follow robots and page directives; those rules describe Amazon’s own crawlers, not a blanket authorization for customer-facing search scraping.

Use a practice target first

Develop your parser against a site that explicitly permits automated requests. The code in this guide targets https://example.com/search with generic selectors so that you can test control flow without presenting unstable Amazon selectors as a guarantee. Replace the URL and selectors only after confirming permission.

Check robots.txt in code

A robots check is one input to your decision, not a substitute for contractual or legal review. This function downloads the file, handles request failures, and asks whether your declared user agent may fetch a URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import urlparse
import requests
from urllib.robotparser import RobotFileParser

def allowed_by_robots(target_url, user_agent):
    parsed = urlparse(target_url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    try:
        response = requests.get(robots_url, headers={"User-Agent": user_agent}, timeout=10)
        response.raise_for_status()
    except requests.RequestException as exc:
        raise RuntimeError(f"Could not verify robots.txt: {exc}") from exc

    parser = RobotFileParser()
    parser.set_url(robots_url)
    parser.parse(response.text.splitlines())
    return parser.can_fetch(user_agent, target_url)

ua = "ResearchExampleBot/1.0 (contact: you@example.com)"
url = "https://example.com/search?k=python+book"
if not allowed_by_robots(url, ua):
    raise SystemExit("robots.txt disallows this URL for the declared user agent")

2. Build a low-rate, stoppable HTTP collector

Use one requests.Session so connection reuse reduces overhead. Send an honest identifying user agent with a contact address, set a finite timeout, cap retries, and add a delay between pages. A 403, 429, or 503 is a stop signal. So is a response containing a CAPTCHA or robot-check page. Do not rotate identities, defeat challenges, or increase concurrency to evade those controls.

import csv
import hashlib
import time
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

BASE_URL = "https://example.com/search"
QUERY = "python book"
MAX_PAGES = 3
DELAY_SECONDS = 2
USER_AGENT = "ResearchExampleBot/1.0 (contact: you@example.com)"

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})

STOP_STATUSES = {403, 429, 503}

def looks_blocked(response):
    sample = response.text[:200_000].lower()
    markers = ("captcha", "robot check", "automated access", "verify you are human")
    return response.status_code in STOP_STATUSES or any(marker in sample for marker in markers)

def get_with_limited_retries(params, attempts=2):
    for attempt in range(attempts + 1):
        try:
            response = session.get(BASE_URL, params=params, timeout=15)
        except requests.RequestException as exc:
            if attempt == attempts:
                raise
            time.sleep(2 ** attempt)
            continue
        if response.status_code == 200:
            return response
        if response.status_code in STOP_STATUSES:
            return response
        if attempt == attempts:
            return response
        time.sleep(2 ** attempt)
    raise RuntimeError("unreachable")

rows = []
seen_urls = set()
raw_hashes = []

for page in range(1, MAX_PAGES + 1):
    params = {"k": QUERY, "page": page}
    response = get_with_limited_retries(params)
    retrieved_at = datetime.now(timezone.utc).isoformat()
    print({"page": page, "status": response.status_code, "url": response.url})

    if response.status_code != 200 or looks_blocked(response):
        print("Stopping: blocked, challenged, or unsuccessful response")
        break

    raw_hashes.append({"page": page, "sha256": hashlib.sha256(response.content).hexdigest()})
    soup = BeautifulSoup(response.text, "html.parser")
    new_on_page = 0

    for card in soup.select("article.product"):
        title_node = card.select_one(".title")
        link_node = card.select_one("a[href]")
        if not title_node or not link_node:
            continue
        product_url = urljoin(response.url, link_node["href"])
        if product_url in seen_urls:
            continue
        seen_urls.add(product_url)
        new_on_page += 1
        rows.append({
            "url": product_url,
            "title": title_node.get_text(" ", strip=True),
            "price_text": (card.select_one(".price") or {}).get_text(" ", strip=True) if card.select_one(".price") else "",
            "rating_text": (card.select_one(".rating") or {}).get_text(" ", strip=True) if card.select_one(".rating") else "",
            "review_count_text": (card.select_one(".reviews") or {}).get_text(" ", strip=True) if card.select_one(".reviews") else "",
            "retrieved_at": retrieved_at,
        })

    if new_on_page == 0:
        print("Stopping: page yielded no new products")
        break
    time.sleep(DELAY_SECONDS)

with open("products.csv", "w", newline="", encoding="utf-8") as file:
    writer = csv.DictWriter(file, fieldnames=["url", "title", "price_text", "rating_text", "review_count_text", "retrieved_at"])
    writer.writeheader()
    writer.writerows(rows)

print(f"Wrote {len(rows)} unique products")
print("Raw-page hashes:", raw_hashes)

The example deliberately records text rather than converting prices or ratings to numbers. Currency symbols, decimal separators, localized labels, and missing fields vary by locale. Keep the original text and add a separate normalized value only after defining rules for each locale.

3. Handle pagination without guessing

Prefer a verified next link

If the permitted site exposes a real next-page link, parse and validate that link instead of assuming a page-number parameter. Check that the next URL remains on the approved host and path, then stop when no link exists.

next_node = soup.select_one("a[rel='next'][href]")
if not next_node:
    break
next_url = urljoin(response.url, next_node["href"])

When a page parameter is documented

Use the documented parameter, set a hard maximum such as MAX_PAGES, deduplicate by canonical product URL or ASIN, and stop when a page produces no new records. Never let an unexpected loop determine request volume.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not assume clicks or infinite scroll are ordinary links

Crawlers can miss URLs created through clicks, JavaScript, infinite scroll, or other interaction-driven navigation. If the allowed target requires a browser to reveal more results, treat browser automation as a separate, permissioned approach and retain the same rate limits and stop signals.

4. Select only the fields you need

BeautifulSoup parses the HTML; Requests retrieves it. Restrict extraction to the fields that answer your question: product URL, title, displayed price, displayed rating, review-count text, and retrieval time are a useful minimum. Prefer stable attributes supplied by the target, such as a documented product identifier, over long chains of presentation classes.

Expect missing and changing markup

  • Use select_one and record a parsing miss rather than failing the whole page.
  • Keep raw HTML or a cryptographic hash, status code, response URL, and timestamp for every page.
  • Log the selector and page number whenever an expected field is absent.
  • Do not treat a successful HTTP 200 as proof that product cards were present; a consent, challenge, or alternate locale page can also return 200.

5. Validate the output before using it

Inspect a sample of rows and compare counts with the page you retrieved. Look for duplicate URLs, empty titles, impossible currency conversions, and a sudden drop to zero products. Preserve the displayed strings so a later reviewer can see exactly what the page contained at retrieval time. If the data will drive pricing, inventory, or ranking decisions, add a review step rather than silently filling missing values.

6. Rate limits, reliability, and scale

Low request volume reduces load and makes failures easier to diagnose. Use a single worker until you have written permission for more. Honor crawl-delay guidance when present, monitor response headers, and keep filters narrow. At larger volumes, an industry guide reports that 503 blocking and TLS/JA3 fingerprinting can affect automated clients; that is a reason to evaluate a compliant alternative, not a reason to bypass controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an approach deliberately

Approach Permission and reliability Extraction and maintenance Best fit
Requests plus BeautifulSoup Simple, low overhead; still subject to terms, throttling, and blocks Excellent for server-rendered HTML; selectors require maintenance Small, permissioned prototypes
Browser automation Handles rendered content but consumes more resources and still must obey access rules Can follow interaction-driven navigation; browser scripts are more fragile Permitted pages whose data appears after interaction
Official API or export Usually the clearest contract and more predictable quotas Structured fields; availability and coverage depend on the provider Production data collection when offered
Managed scraping or data API Provider handles infrastructure, but you must review its authorization and terms May offer normalized data; recurring service cost and schema dependency Material volume when an official source is unavailable

7. Troubleshooting common failures

Symptom Likely cause Action
503 Service Unavailable Throttling, overload, or a bot defense Stop the run, preserve the response, review permission and crawl rate, then use an official or managed source if appropriate.
403 or 429 Access denied or rate limit exceeded Do not retry aggressively. End the run and ask the site owner or provider for an approved method.
CAPTCHA or robot-check HTML Automated access challenge Treat it as a terminal signal; do not attempt to solve or evade it.
HTTP 200 but zero cards Selector drift, consent page, challenge page, or locale variation Save the HTML, inspect its title and key text, then update selectors only for an allowed target.
Pages repeat forever Unvalidated next link or ignored page parameter Track visited URLs, cap pages, and stop when no new product URL appears.
Requests time out Slow response or network problem Keep the finite timeout, retry only a small number of times with backoff, and lower request volume.
Prices or ratings look wrong Locale-specific currency, decimal, or label formats Store source text, record the locale, and normalize with explicit locale rules.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

8. Or skip the browser setup

If your requirement is a visual record rather than structured product fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it can accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports X-Page-Verdict and X-Billed headers.

One request, using the documented API at https://screenshotneo.com/docs/:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.amazon.com/s?k=python+book -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.amazon.com/s?k=python+book"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.amazon.com/s?k=python+book' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Use the URL only where you are authorized to capture it; a screenshot is not a substitute for permission to collect structured Amazon data. ScreenshotNeo also offers take_screenshot, get_page_info, and capture_pdf through MCP clients such as Claude or Cursor. Every plan includes its features. The Free plan includes 1,000 shots per month without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000, with two months free on yearly billing.

Create a free ScreenshotNeo account to use the 1,000 monthly shots without entering a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I use the Amazonbot user-agent string for my script?

No. Amazon documents that string for its own crawler systems. Identify your own client honestly and obtain permission for the pages you request.

Should a scraper save screenshots as well as CSV data?

Only when visual evidence is part of your requirement. Screenshots preserve appearance; they do not provide reliable structured fields, so keep the permitted HTML and parsed records for data work.

When is a managed data API preferable to a custom parser?

Consider one when request volume is material, markup maintenance is consuming engineering time, or an official API/export is unavailable. Review the provider’s authorization, coverage, locale support, quotas, and cost before switching.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.