Start with Requests for ordinary server-rendered HTML, add Beautiful Soup to parse it, move to Scrapy when you need a controlled multi-page crawl, and use Playwright only when a real browser must execute JavaScript or perform interactions. This progression keeps a crawler faster, less fragile and easier to operate.
Choose the smallest tool that solves the page
A Python crawler has two separate jobs: downloading a response and extracting data from it. Requests handles the HTTP transport; Beautiful Soup handles HTML/XML navigation. Scrapy adds scheduling and crawl operations. Playwright drives a browser, so it is the right escalation when the useful content appears only after JavaScript runs.
| Tool | What it executes | Best fit | Main trade-off |
|---|---|---|---|
| Requests | HTTP requests only | One page or server-rendered sites | Does not execute JavaScript |
| Beautiful Soup | No downloading; parses supplied HTML/XML | Reliable text, links and attribute extraction | You must provide the response yourself |
| Scrapy | Asynchronous HTTP crawling plus selectors | Many pages, pagination, exports, pipelines and scheduled jobs | More project structure to learn |
| Playwright | Chromium, Firefox or WebKit browser execution | JavaScript-rendered pages, waits, dialogs and user-like actions | Higher CPU/memory use and UI fragility |
Do not begin with browser automation simply because a page looks modern. Check for a documented API, an export, or an HTML response containing the data first.
Before writing code: access, identity and scope
Prefer an API or export
A public API, bulk export or search endpoint is usually more stable and cheaper than scraping rendered pages. It also gives the site owner a clear way to control access. If none exists, define exactly which URLs and fields you need before starting.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Read robots.txt and site rules
robots.txt is one input to your decision, not a substitute for reviewing terms, privacy obligations, authentication boundaries and applicable law. Python’s standard-library urllib.robotparser can test whether your declared user agent is allowed to fetch a URL:
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
site = 'https://example.com'
agent = 'ItechGuidesCrawler/1.0 (+https://example.com/contact)'
robots_url = f'{urlparse(site).scheme}://{urlparse(site).netloc}/robots.txt'
parser = RobotFileParser(robots_url)
parser.read()
url = site + '/catalog/'
if not parser.can_fetch(agent, url):
raise PermissionError(f'Robots policy disallows {url}')
Identify and limit the crawler
Use a descriptive User-Agent with a contact address, enforce a per-domain concurrency limit, add a delay, and stop or slow down when 429 (rate limited) or 503 (temporarily unavailable) responses increase. Rising latency, repeated retries or ban pages are also signals that your rate is too high.
Step 1: fetch a static page with Requests
Requests is the correct first layer for server-rendered HTML. The example below validates the scheme, uses a descriptive identity, applies a timeout, retries only transient statuses with bounded exponential backoff, and records the final URL after redirects. Use a site intended for practice or one you are authorized to crawl.
import time
from urllib.parse import urlparse
import requests
USER_AGENT = 'ItechGuidesCrawler/1.0 (+https://example.com/contact)'
TRANSIENT = {429, 500, 502, 503, 504}
def fetch(url, attempts=3):
parsed = urlparse(url)
if parsed.scheme not in {'http', 'https'} or not parsed.netloc:
raise ValueError(f'Invalid URL: {url}')
headers = {'User-Agent': USER_AGENT, 'Accept': 'text/html,application/xhtml+xml'}
for attempt in range(attempts):
try:
response = requests.get(url, headers=headers, timeout=(10, 30), allow_redirects=True)
if response.status_code == 200:
return response
if response.status_code not in TRANSIENT:
response.raise_for_status()
except requests.RequestException:
if attempt == attempts - 1:
raise
if attempt < attempts - 1:
time.sleep(min(30, 2 ** attempt))
raise RuntimeError(f'No response after {attempts} attempts: {url}')
response = fetch('https://example.com/')
print('status:', response.status_code)
print('requested URL:', response.request.url)
print('response URL:', response.url)
print('bytes:', len(response.content))
print(response.text[:200])
Use response.content when you need raw bytes and response.text when you want Requests' decoded text. A timeout is essential: without one, a stalled connection can hold a worker indefinitely. Do not retry every error; a 401, 403 or 404 generally requires a different decision rather than more traffic.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Step 2: parse the response with Beautiful Soup
Beautiful Soup does not download pages. It turns the HTML you already fetched into a navigable tree. Select stable attributes, normalize whitespace, and treat missing fields as normal because templates change.
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.text, 'html.parser')
items = []
for card in soup.select('article.product-card'):
title_node = card.select_one('h2, h3')
link_node = card.select_one('a[href]')
price_node = card.select_one('[data-price], .price')
if not title_node or not link_node:
continue
items.append({
'title': ' '.join(title_node.get_text(' ', strip=True).split()),
'url': link_node.get('href'),
'price': price_node.get_text(' ', strip=True) if price_node else None,
})
print(items)
CSS selectors such as article.product-card are concise, but prefer semantic elements, data-* attributes or stable IDs over classes that exist only for styling. Keep parsing separate from downloading so you can test extraction against saved HTML without making more requests.
Step 3: add a respectful multi-page crawl
A small queue with Requests and Beautiful Soup
For a modest, single-domain job, a queue, normalized URLs, a visited set and a depth limit are enough. The crawler below follows links only inside the starting host, waits between requests, and stops after a configurable page count.
import time
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
from bs4 import BeautifulSoup
start = 'https://example.com/catalog/'
host = urlparse(start).netloc
queue = deque([(start, 0)])
visited = set()
max_depth = 2
max_pages = 100
delay = 1.0
while queue and len(visited) < max_pages:
current, depth = queue.popleft()
current, _ = urldefrag(current)
parsed = urlparse(current)
if parsed.netloc != host or current in visited or depth > max_depth:
continue
visited.add(current)
try:
page = fetch(current)
except Exception as exc:
print('fetch failed:', current, exc)
continue
soup = BeautifulSoup(page.text, 'html.parser')
for article in soup.select('article.product-card'):
title = article.select_one('h2, h3')
if title:
print({'url': current, 'title': title.get_text(' ', strip=True)})
for anchor in soup.select('a[href]'):
next_url = urldefrag(urljoin(page.url, anchor['href']))[0]
if urlparse(next_url).netloc == host and next_url not in visited:
queue.append((next_url, depth + 1))
time.sleep(delay)
Pagination without an infinite loop
Prefer a site's explicit next-page link. Normalize it with urljoin, stop when the link is absent, and keep the visited set so a malformed “next” link cannot cycle forever. If pagination uses numbered URLs, impose both a maximum page number and a maximum item/page count. Log every fetch failure rather than silently dropping records.
Recommended Free Tools
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Step 4: move to Scrapy for breadth and operations
Scrapy is an application framework for crawling websites and extracting structured data. It supplies spiders, an asynchronous scheduler, duplicate-request filtering, selectors, retries, feed exports, pipelines, middleware, caching options, robots.txt support and crawl-depth controls. Choose it when the crawl spans many pages or domains, must run repeatedly, or needs operational outputs rather than a one-off script.
Install and create a spider
python -m pip install scrapy
scrapy startproject catalog_crawler
cd catalog_crawler
scrapy genspider products example.com
Replace the generated spider with a focused parser. This example follows product pages and an explicit pagination link; Scrapy filters duplicate requests for you.
import scrapy
class ProductsSpider(scrapy.Spider):
name = 'products'
allowed_domains = ['example.com']
start_urls = ['https://example.com/catalog/']
def parse(self, response):
for card in response.css('article.product-card'):
yield {
'title': card.css('h2::text, h3::text').get(),
'url': response.urljoin(card.css('a::attr(href)').get()),
'price': card.css('[data-price]::text, .price::text').get(),
}
next_page = response.css('a[rel="next"]::attr(href)').get()
if next_page:
yield response.follow(next_page, callback=self.parse)
Run an export with a deliberate rate:
scrapy crawl products -O products.json
-s USER_AGENT='ItechGuidesCrawler/1.0 (+https://example.com/contact)'
-s CONCURRENT_REQUESTS_PER_DOMAIN=2
-s DOWNLOAD_DELAY=1.0
-s ROBOTSTXT_OBEY=True
CONCURRENT_REQUESTS caps simultaneous downloads globally, CONCURRENT_REQUESTS_PER_DOMAIN limits pressure on one domain, and DOWNLOAD_DELAY sets the minimum gap. Translate a site's Crawl-delay or Request-rate guidance into settings when applicable. Increase concurrency gradually only while latency and error rates remain acceptable.
Step 5: use Playwright when a browser is genuinely required
Playwright controls a real browser from Python. Use it when content appears only after JavaScript execution, when navigation requires clicks or dialogs, or when a meaningful selector is absent until client-side rendering finishes. Wait for a business-relevant element rather than an arbitrary sleep. If the page fetches JSON, capture that response directly and avoid scraping a fragile rendered tree.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Install and render a page
python -m pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto('https://example.com/app', wait_until='domcontentloaded', timeout=60_000)
page.locator('[data-loaded="true"]').wait_for(state='visible', timeout=30_000)
rows = page.locator('article.product-card').evaluate_all(
"els => els.map(e => ({title: e.querySelector('h2,h3')?.innerText.trim(), href: e.querySelector('a')?.href}))"
)
print(rows)
browser.close()
Capture an underlying API response
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
api_data = []
page.on('response', lambda r: api_data.append(r.json())
if '/api/products' in r.url and r.ok else None)
page.goto('https://example.com/app', wait_until='networkidle', timeout=60_000)
print(api_data)
browser.close()
In production, use a locator that expresses the required state, handle cookie or consent dialogs according to the site's rules, reuse browser contexts where safe, and close pages promptly. Browser sessions consume substantially more resources than HTTP requests and can break when labels or layout structure changes.
Or skip the browser setup
If your goal is a clean image or PDF of a URL rather than DOM data, ScreenshotNeo provides a single website screenshot API call. It accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDF paper and page controls, custom CSS/JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', image);
There is an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to get an access key.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsReliability, performance and cost controls
- Bound the work: set depth, page and item limits; canonicalize fragments and trailing slashes; reject off-domain URLs unless explicitly approved.
- Store checkpoints: write items incrementally and persist the queue or visited set so a crash does not restart the entire crawl.
- Measure health: record status, response URL, latency, bytes, retries and parser failures. Alert on rising 429/503 counts, ban pages or latency.
- Cache during development: replay saved responses or enable framework caching instead of repeatedly hitting the live site.
- Scale in stages: begin with one worker and conservative per-domain concurrency; increase only after observing the site's response.
- Choose the least expensive layer: Requests and Scrapy HTTP downloads use fewer resources than browser contexts. Reserve Playwright for pages that prove they need it.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML contains no visible data | Data is inserted by JavaScript | Inspect the response for a JSON endpoint; use that endpoint if documented, otherwise wait for a selector in Playwright. |
| 403 or a challenge page | Access controls, missing identity or prohibited automation | Stop, review permission and terms, identify your agent, and prefer an API. Do not attempt to bypass a CAPTCHA. |
| 429 or 503 responses | Request rate is too high or the service is overloaded | Reduce concurrency, increase delay, honor retry-after when supplied, and use bounded backoff. |
| Spider revisits the same URLs | Fragments, tracking parameters or relative links are not normalized | Use urldefrag, urljoin, a canonicalization policy and a visited set or Scrapy's duplicate filter. |
| Playwright times out | Wrong readiness condition, slow resource or failed navigation | Wait for a meaningful selector or network response, inspect console/network logs, and set a finite navigation timeout. |
| Fields suddenly become empty | Markup changed or a selector depended on styling classes | Save the failing HTML, choose semantic or data attributes, and add parser tests with representative fixtures. |
| Memory keeps growing | Pages, contexts or response bodies remain open | Close pages promptly, reuse a bounded browser context, stream exports and cap queue size. |
A practical decision checklist
- Can an API, export or search endpoint provide the data? Use it.
- Is the required content in the initial HTML? Fetch with Requests and parse with Beautiful Soup.
- Do you need a bounded crawl across many pages, exports, retries or pipelines? Build a Scrapy spider.
- Does the page require JavaScript, clicks, dialogs or browser-only state? Use Playwright, waiting on meaningful selectors.
- At every stage, check robots.txt and site rules, identify your crawler, limit concurrency, add delays and monitor 429/503 responses.
Frequently Asked Questions
Can Requests execute JavaScript?
No. Requests retrieves the HTTP response; it does not run the page's JavaScript. Use an available data endpoint or escalate to Playwright when browser execution is necessary.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Do I need Beautiful Soup if I use Scrapy?
Not necessarily. Scrapy includes selectors for extraction. Beautiful Soup remains useful for small scripts or when you want a standalone parser separated from downloading.
When should I replace a custom queue with Scrapy?
Move when the crawl needs many pages, asynchronous scheduling, duplicate filtering, exports, pipelines, retries, caching, robots handling or repeatable deployment.
Is browser automation always more accurate?
It can reveal browser-rendered content, but it consumes more resources and is sensitive to UI changes. A documented API or direct JSON response is usually more stable.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

