What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
To scrape a website with Python, use the simplest permitted method that can reach the data: Requests to fetch HTML, Beautiful Soup to parse it, and explicit normalization, validation, and storage around the result. Move to Scrapy for a multi-page crawl, reproduce an underlying data request for JavaScript pages when practical, and use a headless browser such as Playwright only when the rendered DOM is genuinely required.
The examples below are illustrative. Run them only against a site you own, have permission to access, or that explicitly supports your intended use.
1. Choose a permitted target and define the output
Start with a site that authorizes your use. Check for an official API or documented feed before writing a scraper, read the site’s terms, inspect robots.txt, and collect only the fields you need. Robots rules communicate crawler preferences; they are not authorization and do not settle the legal position for a particular jurisdiction, dataset, access method, or intended use.
Define fields before selectors
Suppose each record needs a title, author, and detail-page URL. Write that schema first:
Recommended Free Tools
#1 Best Overall
title: a non-empty stringauthor: a string when present, otherwise a flagged missing valueurl: an absolute HTTP(S) URL on an allowed host
This prevents a selector from quietly producing an attractive but unusable dataset. Decide how to represent missing values, duplicates, pagination, and failed records before collection starts.
2. Install the small static-page stack
Create an isolated environment and install the two libraries used for a one-page or small static collection:
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4
Requests performs HTTP fetching; Beautiful Soup parses and searches the returned markup. Pin versions in your project once you have a repeatable build, and review their current documentation when you deploy.
3. Fetch a static page safely
Use a finite timeout and make HTTP failures visible. The timeout bounds how long the client waits; raise_for_status() turns 4xx and 5xx responses into exceptions instead of allowing bad input into the parser.
import requests
from bs4 import BeautifulSoup
url = 'https://example.com/page' # illustrative; replace with an authorized target
response = requests.get(url, timeout=15)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
print(soup.title.get_text(strip=True) if soup.title else 'No title')
Inspect the returned HTML in your browser’s developer tools or by saving a fixture. Replace the illustrative URL only after confirming that the target permits the request. A successful HTTP response can still contain an error page, a consent wall, or an empty application shell, so inspect the content you actually received.
4. Parse, normalize, validate, and store records
A maintainable scraper has separate stages:
- Fetch: request a URL with a timeout and status handling.
- Parse: locate elements with stable CSS selectors or XPath.
- Normalize: trim text, resolve relative links, and standardize values.
- Validate: check required fields, schemes, hosts, and types.
- Store: write structured output only after validation.
A complete resilient example
from __future__ import annotations
import csv
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
START_URL = 'https://example.com/page' # illustrative authorized target
ALLOWED_HOSTS = {'example.com'}
def clean_text(node):
return node.get_text(' ', strip=True) if node else None
def allowed_http_url(value: str, base_url: str) -> str | None:
absolute = urljoin(base_url, value)
parsed = urlparse(absolute)
if parsed.scheme not in {'http', 'https'}:
return None
if parsed.hostname not in ALLOWED_HOSTS:
return None
return absolute
def extract_records(html: str, page_url: str) -> list[dict[str, str | None]]:
soup = BeautifulSoup(html, 'html.parser')
records = []
for card in soup.select('article'): # replace after inspecting authorized markup
title_node = card.select_one('h2, h3, .title')
author_node = card.select_one('.author, [rel="author"]')
link_node = card.select_one('a[href]')
href = link_node.get('href') if link_node else None
record = {
'title': clean_text(title_node),
'author': clean_text(author_node),
'url': allowed_http_url(href, page_url) if href else None,
}
if record['title'] and record['url']:
records.append(record)
return records
def main() -> None:
response = requests.get(
START_URL,
timeout=15,
headers={'User-Agent': 'ExampleResearchBot/1.0 (contact: you@example.com)'},
)
response.raise_for_status()
records = extract_records(response.text, response.url)
# A simple output check catches a changed selector or an unexpected page.
if not records:
raise RuntimeError('No valid records found; inspect the response and selectors')
with open('records.csv', 'w', newline='', encoding='utf-8') as output:
writer = csv.DictWriter(output, fieldnames=['title', 'author', 'url'])
writer.writeheader()
writer.writerows(records)
if __name__ == '__main__':
main()
The selectors are deliberately placeholders. Do not index an assumed first match such as select('article')[0]; a missing element should produce a flagged or skipped record, not crash an entire crawl. Resolve relative URLs against the response URL, reject unexpected schemes and hosts, and deduplicate records by a stable key such as the canonical URL.
Rank #2
Regression fixtures
Save a small authorized HTML response as a fixture and test extraction against it. A fixture-based check detects a changed class name before a scheduled job silently exports empty columns. Also log the number of fetched pages, parsed records, rejected records, and validation failures.
5. Add pagination and move to Scrapy when the job grows
For several pages, link following, retries, structured crawl state, and exports, use Scrapy rather than hand-rolling queues and crawl bookkeeping. Its workflow centers on a spider, initial requests, callbacks such as parse(), selectors, and yielded items.
Free tools Windows power users keep installed
One-click scans. No signup required.
Create a project and spider
python -m pip install scrapy
scrapy startproject quote_crawl
cd quote_crawl
scrapy genspider quotes example.com
The generated domain is only a scaffold. Replace it with an authorized practice target and its real host before running the spider.
import scrapy
class QuotesSpider(scrapy.Spider):
name = 'quotes'
allowed_domains = ['example.com']
start_urls = ['https://example.com/quotes'] # illustrative
def parse(self, response):
for card in response.css('article'):
title = card.css('h2::text').get()
author = card.css('.author::text').get()
href = card.css('a::attr(href)').get()
if title and href:
yield {
'title': title.strip(),
'author': author.strip() if author else None,
'url': response.urljoin(href),
}
next_href = response.css('a.next::attr(href)').get()
if next_href:
yield response.follow(next_href, callback=self.parse)
Run an export after replacing selectors and the target:
scrapy crawl quotes -O records.json
Use the Scrapy shell to inspect a response and refine CSS or XPath selectors:
scrapy shell 'https://example.com/quotes'
response.css('article h2::text').getall()
response.xpath('//article//h2/text()').getall()
Prefer .get() or .getall() and test for empty results instead of assuming a selector always matches. Scrapy’s tutorial makes the same resilience point: extraction should continue producing useful data when some page elements are absent.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsRobots and crawl controls in Scrapy
Identify your crawler with a descriptive User-Agent and configure Scrapy’s robots middleware to follow the site’s instructions. Keep concurrency and request volume proportionate to the task, add delays where appropriate, and stop when access is denied or disallowed. Following robots.txt is responsible engineering, but it is not a legal authorization.
6. Handle JavaScript-rendered pages without guessing
First, find the data request
If the initial HTML lacks the records, open the browser’s network panel, reload the page, and identify the request that returns the data. When an accessible JSON or HTML endpoint supplies the needed fields, reproduce that request with Requests or Scrapy, subject to the site’s permission and terms. This is usually simpler, cheaper, and more reliable than rendering every page.
Use a browser only when rendering is required
When the data exists only after browser execution or interaction, a headless browser such as Playwright for Python can render the page. Browser automation is a technical fallback, not a method for defeating bot checks, CAPTCHAs, authentication boundaries, or other restrictions.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto('https://example.com/page', wait_until='domcontentloaded', timeout=30_000)
page.wait_for_selector('article', timeout=10_000)
titles = page.locator('article h2').all_text_contents()
print([title.strip() for title in titles])
browser.close()
Install the browser binaries according to Playwright’s current Python documentation. Set finite navigation and selector timeouts, and close the browser in a finally block in production code so failed jobs do not leak processes.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →7. Be polite, secure, and predictable
- Identify yourself: send a descriptive User-Agent with a contact address where practical.
- Honor instructions: follow robots.txt and the site’s published usage limits; stop when access is denied.
- Limit load: fetch only required pages, avoid unnecessary assets, and use measured concurrency.
- Bound failures: set connect/read timeouts, handle non-success status codes, and record retries rather than looping forever.
- Protect secrets: keep API keys, cookies, and authorization headers out of source control and logs.
- Validate URLs: if URLs come from users or scraped content, allow only expected HTTP(S) schemes and hosts to reduce SSRF risk. Do not expose crawler control endpoints to untrusted networks.
Whether a particular collection is lawful depends on the target, data, jurisdiction, access method, contracts, and intended use. This tutorial is technical guidance, not jurisdiction-specific legal advice.
8. Which Python scraping library should you use?
| Need | Starting point | Reason |
|---|---|---|
| One or a few static pages | Requests + Beautiful Soup | Separate HTTP fetching from HTML parsing with a small setup. |
| Multi-page crawl and structured workflow | Scrapy | Spiders, requests, callbacks, selectors, link following, and exports provide crawl state. |
| Dynamic page with an identifiable data source | Reproduce the relevant request | Request-level extraction avoids unnecessary browser rendering. |
| Browser-only behavior or rendered-DOM data | Playwright or a Scrapy browser integration | Use automation when request-level extraction is not practical. |
Choose by page complexity, crawl scale, control over requests, setup time, and operational risk—not by an assumed universal speed ranking. Start small, measure response and parse failures, and promote the job to Scrapy or browser automation only when the requirements justify it.
9. Troubleshooting common failures
403, 429, or an access-denied page
Cause: the site rejected the request, rate-limited it, or requires an approved access path. Fix: stop or slow the job, verify permission and the published API, identify your crawler, and do not try to bypass the control.
200 response but no records
Cause: you received an application shell, consent page, changed markup, or the wrong URL. Fix: save the response, inspect its title and body, check network requests for a data endpoint, and update selectors only after confirming the new structure.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteTimeouts and intermittent connection errors
Cause: slow servers, overloaded concurrency, or an unbounded wait. Fix: use finite connect/read or navigation timeouts, modest retries with backoff, lower concurrency, and metrics that distinguish fetch failures from parse failures.
Relative or unsafe links
Cause: href values are fragments, protocol-relative URLs, or links to unexpected hosts. Fix: resolve with urljoin, allow only HTTP(S), enforce an allowed-host set, and discard or review anything outside it.
Fields disappear after a redesign
Cause: selectors depended on presentation classes or a fixed DOM position. Fix: prefer semantic attributes and stable containers, keep HTML fixtures, validate row counts and required fields, and alert on sudden changes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.10. Performance, reliability, and cost planning
Request-level extraction normally uses fewer resources than launching a browser per page. Reuse HTTP sessions where appropriate, avoid downloading assets you do not parse, cache only when the site’s rules permit it, and keep a bounded queue. For Scrapy, tune concurrency conservatively and monitor memory, retries, status codes, and item counts. For browser jobs, reuse a browser process carefully, limit pages, and always close contexts.
Best Value
Reliability comes from observable checkpoints: record the requested URL, final URL, status, elapsed time, parser version, extracted count, and validation errors. Store raw responses or hashes when your retention policy allows it so a failed parse can be reproduced without immediately re-requesting the site. There is no meaningful universal speed or success statistic here; the target, network, markup, and access policy determine results.
Or skip the browser setup
If your goal is a clean screenshot rather than parsed fields, ScreenshotNeo provides a website screenshot API and MCP server. Its request accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it can accept the cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Use the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom JavaScript and CSS, pre-capture clicks, hide selectors, selector/delay/network-idle waits, ad/tracker/request/resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, 100-URL bulk capture, usage reporting, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work.
One-call examples
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing gives two months free. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Create a free ScreenshotNeo account to try it without a card.
Frequently Asked Questions
How should I store scraped data for later analysis?
Use a stable schema and an explicit encoding such as UTF-8. CSV is convenient for flat records; JSON preserves nested values. Keep the source URL and collection timestamp with each record so downstream users can trace provenance.
How can I test a scraper without repeatedly contacting a live site?
Save a permitted response as an HTML fixture and run your parser against that file in automated tests. Add cases for missing fields, malformed links, duplicate records, and a changed layout.
When should a scraper become a scheduled job?
Schedule it only after selectors, validation checks, rate limits, error handling, and alerting are in place. A job that exports an empty file successfully is still a failed job.
Can I scrape authenticated pages?
Only with explicit authorization and a permitted access method. Treat session cookies and Authorization headers as secrets, avoid logging them, and respect the account’s terms and data boundaries.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

