You can scrape search-result HTML with Python’s requests and BeautifulSoup only when the site permits automated access. Check Amazon’s robots.txt and terms first, identify yourself, send requests slowly, stop on 403, 429, 503, CAPTCHA, or robot-check responses, and keep a strict page limit. The example below uses example.com and generic selectors deliberately; Amazon’s customer-facing markup changes and its bot documentation does not grant permission to collect search pages.
1. Confirm that automated access is allowed
Before writing a scraper, decide whether you have a lawful, permissioned source. Read the target locale’s terms and its robots.txt. If a path is disallowed or the terms prohibit automated access, stop and use an official API, a permitted export, or a licensed data provider instead. Amazon’s documentation describes Amazonbot, Amzn-SearchBot, and Amzn-User and how those crawlers follow robots and page directives; those rules describe Amazon’s own crawlers, not a blanket authorization for customer-facing search scraping.
Use a practice target first
Develop your parser against a site that explicitly permits automated requests. The code in this guide targets https://example.com/search with generic selectors so that you can test control flow without presenting unstable Amazon selectors as a guarantee. Replace the URL and selectors only after confirming permission.
Check robots.txt in code
A robots check is one input to your decision, not a substitute for contractual or legal review. This function downloads the file, handles request failures, and asks whether your declared user agent may fetch a URL.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
from urllib.parse import urlparse
import requests
from urllib.robotparser import RobotFileParser
def allowed_by_robots(target_url, user_agent):
parsed = urlparse(target_url)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
try:
response = requests.get(robots_url, headers={"User-Agent": user_agent}, timeout=10)
response.raise_for_status()
except requests.RequestException as exc:
raise RuntimeError(f"Could not verify robots.txt: {exc}") from exc
parser = RobotFileParser()
parser.set_url(robots_url)
parser.parse(response.text.splitlines())
return parser.can_fetch(user_agent, target_url)
ua = "ResearchExampleBot/1.0 (contact: you@example.com)"
url = "https://example.com/search?k=python+book"
if not allowed_by_robots(url, ua):
raise SystemExit("robots.txt disallows this URL for the declared user agent")
2. Build a low-rate, stoppable HTTP collector
Use one requests.Session so connection reuse reduces overhead. Send an honest identifying user agent with a contact address, set a finite timeout, cap retries, and add a delay between pages. A 403, 429, or 503 is a stop signal. So is a response containing a CAPTCHA or robot-check page. Do not rotate identities, defeat challenges, or increase concurrency to evade those controls.
import csv
import hashlib
import time
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
BASE_URL = "https://example.com/search"
QUERY = "python book"
MAX_PAGES = 3
DELAY_SECONDS = 2
USER_AGENT = "ResearchExampleBot/1.0 (contact: you@example.com)"
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
STOP_STATUSES = {403, 429, 503}
def looks_blocked(response):
sample = response.text[:200_000].lower()
markers = ("captcha", "robot check", "automated access", "verify you are human")
return response.status_code in STOP_STATUSES or any(marker in sample for marker in markers)
def get_with_limited_retries(params, attempts=2):
for attempt in range(attempts + 1):
try:
response = session.get(BASE_URL, params=params, timeout=15)
except requests.RequestException as exc:
if attempt == attempts:
raise
time.sleep(2 ** attempt)
continue
if response.status_code == 200:
return response
if response.status_code in STOP_STATUSES:
return response
if attempt == attempts:
return response
time.sleep(2 ** attempt)
raise RuntimeError("unreachable")
rows = []
seen_urls = set()
raw_hashes = []
for page in range(1, MAX_PAGES + 1):
params = {"k": QUERY, "page": page}
response = get_with_limited_retries(params)
retrieved_at = datetime.now(timezone.utc).isoformat()
print({"page": page, "status": response.status_code, "url": response.url})
if response.status_code != 200 or looks_blocked(response):
print("Stopping: blocked, challenged, or unsuccessful response")
break
raw_hashes.append({"page": page, "sha256": hashlib.sha256(response.content).hexdigest()})
soup = BeautifulSoup(response.text, "html.parser")
new_on_page = 0
for card in soup.select("article.product"):
title_node = card.select_one(".title")
link_node = card.select_one("a[href]")
if not title_node or not link_node:
continue
product_url = urljoin(response.url, link_node["href"])
if product_url in seen_urls:
continue
seen_urls.add(product_url)
new_on_page += 1
rows.append({
"url": product_url,
"title": title_node.get_text(" ", strip=True),
"price_text": (card.select_one(".price") or {}).get_text(" ", strip=True) if card.select_one(".price") else "",
"rating_text": (card.select_one(".rating") or {}).get_text(" ", strip=True) if card.select_one(".rating") else "",
"review_count_text": (card.select_one(".reviews") or {}).get_text(" ", strip=True) if card.select_one(".reviews") else "",
"retrieved_at": retrieved_at,
})
if new_on_page == 0:
print("Stopping: page yielded no new products")
break
time.sleep(DELAY_SECONDS)
with open("products.csv", "w", newline="", encoding="utf-8") as file:
writer = csv.DictWriter(file, fieldnames=["url", "title", "price_text", "rating_text", "review_count_text", "retrieved_at"])
writer.writeheader()
writer.writerows(rows)
print(f"Wrote {len(rows)} unique products")
print("Raw-page hashes:", raw_hashes)
The example deliberately records text rather than converting prices or ratings to numbers. Currency symbols, decimal separators, localized labels, and missing fields vary by locale. Keep the original text and add a separate normalized value only after defining rules for each locale.
3. Handle pagination without guessing
Prefer a verified next link
If the permitted site exposes a real next-page link, parse and validate that link instead of assuming a page-number parameter. Check that the next URL remains on the approved host and path, then stop when no link exists.
next_node = soup.select_one("a[rel='next'][href]")
if not next_node:
break
next_url = urljoin(response.url, next_node["href"])
When a page parameter is documented
Use the documented parameter, set a hard maximum such as MAX_PAGES, deduplicate by canonical product URL or ASIN, and stop when a page produces no new records. Never let an unexpected loop determine request volume.
Free tools Windows power users keep installed
One-click scans. No signup required.
Do not assume clicks or infinite scroll are ordinary links
Crawlers can miss URLs created through clicks, JavaScript, infinite scroll, or other interaction-driven navigation. If the allowed target requires a browser to reveal more results, treat browser automation as a separate, permissioned approach and retain the same rate limits and stop signals.
4. Select only the fields you need
BeautifulSoup parses the HTML; Requests retrieves it. Restrict extraction to the fields that answer your question: product URL, title, displayed price, displayed rating, review-count text, and retrieval time are a useful minimum. Prefer stable attributes supplied by the target, such as a documented product identifier, over long chains of presentation classes.
Rank #3
Expect missing and changing markup
- Use
select_oneand record a parsing miss rather than failing the whole page. - Keep raw HTML or a cryptographic hash, status code, response URL, and timestamp for every page.
- Log the selector and page number whenever an expected field is absent.
- Do not treat a successful HTTP 200 as proof that product cards were present; a consent, challenge, or alternate locale page can also return 200.
5. Validate the output before using it
Inspect a sample of rows and compare counts with the page you retrieved. Look for duplicate URLs, empty titles, impossible currency conversions, and a sudden drop to zero products. Preserve the displayed strings so a later reviewer can see exactly what the page contained at retrieval time. If the data will drive pricing, inventory, or ranking decisions, add a review step rather than silently filling missing values.
6. Rate limits, reliability, and scale
Low request volume reduces load and makes failures easier to diagnose. Use a single worker until you have written permission for more. Honor crawl-delay guidance when present, monitor response headers, and keep filters narrow. At larger volumes, an industry guide reports that 503 blocking and TLS/JA3 fingerprinting can affect automated clients; that is a reason to evaluate a compliant alternative, not a reason to bypass controls.
Choose an approach deliberately
| Approach | Permission and reliability | Extraction and maintenance | Best fit |
|---|---|---|---|
| Requests plus BeautifulSoup | Simple, low overhead; still subject to terms, throttling, and blocks | Excellent for server-rendered HTML; selectors require maintenance | Small, permissioned prototypes |
| Browser automation | Handles rendered content but consumes more resources and still must obey access rules | Can follow interaction-driven navigation; browser scripts are more fragile | Permitted pages whose data appears after interaction |
| Official API or export | Usually the clearest contract and more predictable quotas | Structured fields; availability and coverage depend on the provider | Production data collection when offered |
| Managed scraping or data API | Provider handles infrastructure, but you must review its authorization and terms | May offer normalized data; recurring service cost and schema dependency | Material volume when an official source is unavailable |
7. Troubleshooting common failures
| Symptom | Likely cause | Action |
|---|---|---|
| 503 Service Unavailable | Throttling, overload, or a bot defense | Stop the run, preserve the response, review permission and crawl rate, then use an official or managed source if appropriate. |
| 403 or 429 | Access denied or rate limit exceeded | Do not retry aggressively. End the run and ask the site owner or provider for an approved method. |
| CAPTCHA or robot-check HTML | Automated access challenge | Treat it as a terminal signal; do not attempt to solve or evade it. |
| HTTP 200 but zero cards | Selector drift, consent page, challenge page, or locale variation | Save the HTML, inspect its title and key text, then update selectors only for an allowed target. |
| Pages repeat forever | Unvalidated next link or ignored page parameter | Track visited URLs, cap pages, and stop when no new product URL appears. |
| Requests time out | Slow response or network problem | Keep the finite timeout, retry only a small number of times with backoff, and lower request volume. |
| Prices or ratings look wrong | Locale-specific currency, decimal, or label formats | Store source text, record the locale, and normalize with explicit locale rules. |
8. Or skip the browser setup
If your requirement is a visual record rather than structured product fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it can accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports X-Page-Verdict and X-Billed headers.
One request, using the documented API at https://screenshotneo.com/docs/:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.amazon.com/s?k=python+book -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.amazon.com/s?k=python+book"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.amazon.com/s?k=python+book' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Use the URL only where you are authorized to capture it; a screenshot is not a substitute for permission to collect structured Amazon data. ScreenshotNeo also offers take_screenshot, get_page_info, and capture_pdf through MCP clients such as Claude or Cursor. Every plan includes its features. The Free plan includes 1,000 shots per month without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000, with two months free on yearly billing.
Create a free ScreenshotNeo account to use the 1,000 monthly shots without entering a card.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteFrequently Asked Questions
Can I use the Amazonbot user-agent string for my script?
No. Amazon documents that string for its own crawler systems. Identify your own client honestly and obtain permission for the pages you request.
Best Value
Should a scraper save screenshots as well as CSV data?
Only when visual evidence is part of your requirement. Screenshots preserve appearance; they do not provide reliable structured fields, so keep the permitted HTML and parsed records for data work.
When is a managed data API preferable to a custom parser?
Consider one when request volume is material, markup maintenance is consuming engineering time, or an official API/export is unavailable. Review the provider’s authorization, coverage, locale support, quotas, and cost before switching.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

