Recommended Free Tools
Web scraping is a pipeline: send an HTTP request, receive a response, parse the returned HTML or XML, extract the fields you need, and store them. Use a browser automation tool only when the data is created after JavaScript runs. A reliable scraper also respects the target site’s published instructions, controls request load, handles failures, and expects page markup to change.
This guide uses “Web Scraping Cookbook: Practical Recipes for Real-World Sites” as an editorial title. The clearly identified related publication is Python Web Scraping Cookbook by Lazar Telebak, Michael Heydt, and Mei Lu, published by Packt in 2018 (364 pages, ISBN 9781787285217). Its examples and library versions are historical, so verify APIs against current documentation before copying them into production.
What counts as a web-scraping request?
A request is an HTTP transaction your client sends to a web server. A normal page download usually involves one request for the document, followed by additional requests for stylesheets, scripts, images, fonts, API calls, analytics, and advertisements. A simple Python scraper generally makes one deliberate request with requests.get(); a browser may make dozens while rendering one visible page.
Redirects and retries can create additional server-side traffic. A failed request still consumes resources, even if no page is returned. Cache hits, where your own program reuses a previously saved response, avoid a new request to the origin. There is no universal “safe” requests-per-second number: choose a conservative schedule based on the site’s instructions, response behavior, crawl size, and your operational need.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Before writing code: define the crawl
Identify the data and URL scope
- Write down the exact fields, such as title, price, publication date, or product URL.
- Start with a small, finite URL list or a narrowly defined link-following rule.
- Record the source URL and retrieval time with every extracted record.
- Decide how duplicates, missing fields, pagination, and deleted pages will be represented.
Read site instructions and access conditions
Check the site’s terms, API documentation, authentication requirements, and /robots.txt. RFC 9309, the IETF’s September 2022 Robots Exclusion Protocol specification, describes rules that crawlers are requested to honor. It explicitly says: “These rules are not a form of access authorization.” Robots.txt does not grant permission, replace authentication, or override other access controls. Whether a particular crawl is lawful depends on the site’s conditions and the applicable jurisdiction; a generic scraper tutorial cannot settle that question.
Plan load and storage
Use delays, bounded concurrency, caching, and a clear stop condition. The related cookbook treats crawling with delays and caching as separate practical techniques. Do not assume a fixed delay is appropriate for every service: begin conservatively, observe responses, and reduce load when the site slows, returns errors, or asks you to stop.
Recipe 1: fetch a static page with Requests
When the required content is in the initial HTML response, a regular HTTP client is simpler and lighter than a browser. Install the current Requests package in your virtual environment, then use a timeout and explicit error handling:
from pathlib import Path
import requests
url = "https://example.com/catalog"
headers = {
"User-Agent": "catalog-research/1.0 (contact: you@example.com)"
}
try:
response = requests.get(url, headers=headers, timeout=30)
response.raise_for_status()
except requests.RequestException as exc:
raise SystemExit(f"Request failed: {exc}")
Path("catalog.html").write_bytes(response.content)
print(response.status_code, response.url, len(response.content))
A timeout prevents a worker from waiting forever. raise_for_status() turns 4xx and 5xx responses into exceptions, while response.url reveals the final URL after redirects. Keep the response bytes when possible; decoding can be revisited if a page declares an unusual character set.
Rank #2
Recipe 2: parse HTML with Beautiful Soup
Beautiful Soup parses HTML or XML; it does not fetch pages by itself. Pass it the response content, select elements, and normalize missing values deliberately:
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.content, "html.parser")
rows = []
for card in soup.select("article.product"):
title_node = card.select_one("h2")
price_node = card.select_one(".price")
link_node = card.select_one("a")
rows.append({
"title": title_node.get_text(" ", strip=True) if title_node else None,
"price": price_node.get_text(" ", strip=True) if price_node else None,
"url": link_node.get("href") if link_node else None,
})
for row in rows:
print(row)
The official Beautiful Soup documentation currently shows version 4.14.3 at the time represented by the source material; check the documentation and your installed version for current behavior. Prefer stable semantic selectors over generated class names. Test selectors against fixtures saved from the site, and log when an expected element disappears instead of silently writing empty records.
Recipe 3: follow pagination without creating a runaway crawler
Represent pagination as an explicit loop with a maximum page count, a visited-URL set, and a delay you can adjust. Resolve relative links with urllib.parse.urljoin and keep the crawl on the intended host.
import time
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
session = requests.Session()
session.headers["User-Agent"] = "catalog-research/1.0 (contact: you@example.com)"
url = "https://example.com/catalog"
seen = set()
records = []
max_pages = 20
for _ in range(max_pages):
if url in seen or urlparse(url).netloc != "example.com":
break
seen.add(url)
response = session.get(url, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")
for card in soup.select("article.product"):
records.append(card.get_text(" ", strip=True))
next_node = soup.select_one("a[rel='next']")
if not next_node or not next_node.get("href"):
break
url = urljoin(response.url, next_node["href"])
time.sleep(2)
The two-second pause above is an example, not a universal recommendation. Tune it to the service’s published guidance and observed behavior. For larger jobs, persist the queue and results so a process restart does not repeat completed pages.
Rank #3
Recipe 4: handle JavaScript-rendered content
A plain HTTP response may contain only an application shell; the records appear after client-side JavaScript calls an API. First inspect the response and browser developer tools. If the site exposes a documented data endpoint, using that endpoint may be more efficient and predictable than rendering a full browser, subject to its access conditions.
When browser execution is required, Selenium is one approach identified by the related cookbook. Browser automation costs more CPU and memory, is slower, and introduces driver and browser-version maintenance. Keep it scoped to pages that actually need rendering:
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
options = Options()
options.add_argument("--headless=new")
options.add_argument("--disable-gpu")
with webdriver.Chrome(options=options) as driver:
driver.get("https://example.com/catalog")
WebDriverWait(driver, 20).until(
EC.presence_of_element_located((By.CSS_SELECTOR, "article.product"))
)
html = driver.page_source
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
print([x.get_text(" ", strip=True) for x in soup.select("article.product")])
Use current Selenium and browser-driver documentation for installation and capabilities. Waiting for a selector is more reliable than sleeping for an arbitrary duration, but it still needs a timeout and an error path. Do not assume Beautiful Soup can execute JavaScript; it only parses the markup it receives.
Choosing an approach
| Situation | Preferred starting point | Trade-off |
|---|---|---|
| Data is in initial HTML; small crawl | Requests plus Beautiful Soup | Simple and controllable, but no JavaScript execution |
| Many URLs, retries, scheduling, and persistence | A crawler framework such as Scrapy | More structure and configuration to learn |
| Content appears after browser execution | Selenium or another supported browser automation tool | Higher resource use and browser maintenance |
| Repeated snapshots of unchanged pages | HTTP caching and stored responses | Lower load, but data can become stale |
Choose based on where the data is generated, crawl size, operational control, and the target service’s instructions—not on a single “best” library.
Rank #4
Extraction that survives page changes
Validate every record
- Require key fields before writing a record.
- Check that URLs use the expected scheme and host.
- Preserve raw HTML or a content hash for debugging, subject to storage and privacy requirements.
- Count records per page and alert on sudden zero-result pages.
Normalize carefully
Strip surrounding whitespace, convert relative URLs, and parse numbers only after removing known currency and thousands separators. Keep the original text when a conversion fails. Dates require an explicit timezone and format assumption; do not silently reinterpret ambiguous dates.
Separate extraction from transport
Write one function that fetches bytes and another that parses a document. This lets you test selectors against saved fixtures without repeatedly contacting a live site, and it makes a later switch from Requests to Selenium less invasive.
Reliability, retries, and caching
Retry only transient failures, such as a connection reset or a service response that explicitly indicates temporary overload. Do not blindly retry authentication failures, forbidden responses, or a page that tells you to stop. Use exponential backoff with a cap, add jitter when multiple workers run, and enforce a maximum attempt count. Log status code, URL, elapsed time, attempt number, and exception type without recording secrets.
Cache successful responses keyed by URL and relevant request parameters. A cache prevents duplicate downloads during development and resumptions, but it can serve stale content. Store a retrieval timestamp and define when a refresh is required. For a large crawl, persist a queue, completed URLs, failures, and output checkpoints so interruption does not force a full restart.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 200 response but no products | Content is rendered by JavaScript or selectors changed | Inspect raw HTML, identify the data endpoint or use browser automation, then update and test selectors |
| 403 or 429 responses | Access policy, authentication, or excessive request load | Stop and read the site’s instructions; authenticate where authorized, reduce concurrency, add caching, and do not attempt to bypass controls |
| Intermittent timeouts | Slow origin, network instability, or an overly short timeout | Set separate connect/read timeouts, retry bounded transient errors with backoff, and record failures for later review |
| Duplicate records | Pagination links, redirects, or URL variants repeat content | Canonicalize URLs where appropriate, maintain a visited set, and deduplicate using a stable record key |
| Browser session fails to start | Missing or incompatible browser/driver installation | Use the current Selenium setup instructions, verify versions, and reproduce with a minimal page before adding extraction logic |
| Parser raises encoding errors | Incorrectly assumed character set | Parse response bytes, inspect headers and document declarations, and preserve undecodable input for diagnosis |
Testing and deployment checklist
- Save representative HTML fixtures, including an empty page, a changed layout, and an error response.
- Unit-test each selector and assert required fields.
- Run a small, bounded crawl before expanding the URL set.
- Configure timeouts, bounded retries, delays, logging, and a stop mechanism.
- Keep credentials in environment variables or a secret manager, never in source files.
- Schedule jobs with a persistent queue and cache; monitor error rates and record counts.
- Review access conditions whenever the target site changes its terms, robots file, API, or authentication flow.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It accepts cookie and consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each cleanup step off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the result identified by X-Page-Verdict and X-Billed headers.
For a rendered visual capture, make one request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, custom viewports and retina scale, PDF output, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs and webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. Sign up for the free ScreenshotNeo plan.
Further reading
Python Web Scraping Cookbook (Packt, 2018; ISBN 9781787285217) is a related beginner-to-intermediate reference covering Requests, Beautiful Soup, Scrapy, Selenium, dynamic pages, crawling conduct, delays, caching, and deployment. Treat its examples as historical and verify current package documentation before production use.
Frequently Asked Questions
Does robots.txt make scraping legal?
No. RFC 9309 describes requested crawler rules and states that they are not access authorization. Review the site’s conditions and the law applicable to your situation.
Should I always use Selenium?
No. Use Requests and a parser when the required data is in the initial HTML. Add browser automation only for content that genuinely requires JavaScript execution.
How fast can my scraper run?
There is no universal safe rate. Follow site guidance, start conservatively, cache responses, limit concurrency, and slow down when the service shows errors or overload.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →

