A reusable web-scraping template is a small Python program with clear stages: check the site’s rules, configure a URL and selectors, fetch the page, parse named fields, validate them, and save structured output. It is a starting point, not a universal scraper: each site has its own markup, access rules, and failure modes. The example below is designed for permitted public pages whose needed content is present in the initial HTML response.
What a web-scraping template should do
A useful template separates decisions that change from code that stays stable. You should be able to update the target URL, selectors, request headers, output path, and pacing without rewriting the fetch-and-save workflow. Before automating a site, check its terms, technical instructions, and any official API or developer documentation. Prefer an appropriate official API when one is available.
A successful response is not proof that access is permitted, that the page markup is stable, or that the extracted values are correct. Build checks into the program so a blocked request or site redesign does not silently create bad records.
The reusable workflow
- Configure: identify the URL, selectors, output file, and a conservative delay consistent with site instructions.
- Check: inspect the robots.txt file for the correct origin and read the site’s terms and technical guidance.
- Fetch: request the page with an appropriate user agent, follow redirects deliberately, and check HTTP status.
- Parse: extract fields by name using selectors suited to the actual HTML.
- Validate: flag missing fields, malformed values, duplicates, and unexpected page changes.
- Save: write structured output and keep enough context to diagnose failures.
Check site rules before sending requests
Robots.txt is crawler guidance, not a security boundary or permission grant. Google explains that crawler instructions in robots.txt cannot enforce crawler behavior, and a disallowed URL may still be indexed if other pages link to it. Do not use robots.txt to protect private information. Read Google’s robots.txt introduction for its description of that distinction.
#1 Best Overall
Scope matters: Google says robots.txt rules apply to the host, protocol, and port where the file is hosted. A subdomain’s file does not automatically govern the parent domain. For Google’s crawler behavior, the specification is UTF-8 plain text, has a 500 KiB size limit, and does not support crawl-delay. Those are details of Google’s interpretation, not a guarantee that every crawler handles the file identically. See Google’s robots.txt specification.
Check the robots.txt URL at the same origin as the pages you intend to request—for example, the protocol and hostname in your target URL—and check the site’s terms and published access rules separately. If access is restricted or the site’s instructions are unclear, stop and seek permission. The legal status of a particular use depends on its facts and jurisdiction; this guide does not make that determination.
A Python template for one permitted page
This example uses requests to retrieve a page and Beautiful Soup to parse it. It deliberately handles one URL at a time; it does not guess selectors, bypass access controls, or attempt to evade bot checks. Install its dependencies with python -m pip install requests beautifulsoup4.
Replace the example URL and CSS selectors with values appropriate to the site. The selectors below are illustrative, not selectors verified for a particular website.
from __future__ import annotations
import csv
import json
import logging
import time
from pathlib import Path
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
# ---- Configuration: change these for the site and fields you need. ----
TARGET_URL = "https://example.com/articles/sample"
USER_AGENT = "ExampleResearchBot/1.0 (contact: you@example.com)"
OUTPUT_CSV = Path("records.csv")
OUTPUT_JSON = Path("records.json")
REQUEST_DELAY_SECONDS = 2.0 # Set this to comply with the site's instructions.
SELECTORS = {
"title": "h1",
"author": "[rel='author']",
"published": "time[datetime]",
}
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
def robots_allows(url: str, user_agent: str) -> bool:
"""Check the target origin's robots.txt for this crawler user agent."""
parsed = urlparse(url)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
parser = RobotFileParser()
parser.set_url(robots_url)
try:
parser.read()
except (OSError, ValueError) as exc:
raise RuntimeError(f"Could not read {robots_url}: {exc}") from exc
return parser.can_fetch(user_agent, url)
def fetch_html(url: str, session: requests.Session) -> str:
response = session.get(
url,
headers={"User-Agent": USER_AGENT, "Accept": "text/html"},
timeout=(10, 30), # connect timeout, then read timeout, in seconds
allow_redirects=True,
)
logging.info("GET %s -> %s", response.url, response.status_code)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "html" not in content_type.lower():
raise ValueError(f"Expected HTML, received Content-Type: {content_type!r}")
return response.text
def text_for(soup: BeautifulSoup, selector: str) -> str:
element = soup.select_one(selector)
return element.get_text(" ", strip=True) if element else ""
def parse_record(html: str, source_url: str) -> dict[str, str]:
soup = BeautifulSoup(html, "html.parser")
record = {name: text_for(soup, selector)
for name, selector in SELECTORS.items()}
# Preserve a machine-readable attribute where it is useful.
published = soup.select_one(SELECTORS["published"])
if published:
record["published_datetime"] = published.get("datetime", "")
record["source_url"] = source_url
return record
def validate(record: dict[str, str]) -> None:
missing = [name for name in ("title",) if not record.get(name)]
if missing:
raise ValueError(f"Required field(s) missing: {', '.join(missing)}")
# A basic guard against an unexpected or empty extraction.
if len(record["title"]) > 500:
raise ValueError("Title is unexpectedly long; check the selector or page markup")
def save_records(records: list[dict[str, str]]) -> None:
if not records:
logging.warning("No records to save")
return
fields = list(dict.fromkeys(key for record in records for key in record))
with OUTPUT_CSV.open("w", newline="", encoding="utf-8") as handle:
writer = csv.DictWriter(handle, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
writer.writerows(records)
OUTPUT_JSON.write_text(
json.dumps(records, ensure_ascii=False, indent=2), encoding="utf-8"
)
logging.info("Saved %d record(s) to %s and %s",
len(records), OUTPUT_CSV, OUTPUT_JSON)
def main() -> None:
if not robots_allows(TARGET_URL, USER_AGENT):
raise SystemExit(f"robots.txt disallows this URL for {USER_AGENT}")
# For a single URL this is a no-op; keep it when extending to a URL list.
time.sleep(REQUEST_DELAY_SECONDS)
with requests.Session() as session:
try:
html = fetch_html(TARGET_URL, session)
record = parse_record(html, TARGET_URL)
validate(record)
save_records([record])
except (requests.RequestException, ValueError, RuntimeError) as exc:
logging.error("Could not produce a valid record for %s: %s",
TARGET_URL, exc)
raise SystemExit(1) from exc
if __name__ == "__main__":
main()
The code writes both a CSV file and a JSON array in the current working directory. CSV is handy for a spreadsheet; JSON preserves named fields and is often more convenient for another program. The timeout tuple limits connection establishment and response reading, but it does not make a slow or unstable site reliable. A failed request is logged and stops the run instead of producing a plausible-looking empty record.
Rank #2
Adapt selectors against the actual page
Inspect the page’s HTML using your browser’s developer tools or a saved response, then choose selectors tied to the fields you need. soup.select_one() returns the first match or no match; soup.select() returns all matches. For repeated items such as search results, select each item container first and parse fields within that container so titles and links from different cards are not accidentally paired.
Prefer stable attributes, semantic elements, or documented structured data when available. A selector based on a long chain of layout containers can break after a minor redesign. Validate expected fields and, for list extraction, verify that the number of parsed items is plausible for the page rather than accepting an empty result.
Turn the one-page template into a small crawl
For multiple permitted URLs, put the URLs in a configuration file or a list, process them one at a time, and keep the delay between requests. Record the requested URL, final response URL, status code, timestamp, and error alongside failed items. This makes redirects, changed pages, and temporary outages distinguishable from parser failures.
- Deduplicate: normalize URLs consistently and track identifiers or canonical URLs if the site supplies them.
- Make output recoverable: append or checkpoint records for long runs rather than keeping the only copy in memory until the end.
- Keep failures visible: distinguish network exceptions, non-success HTTP statuses, missing selectors, and validation errors.
- Re-check changes: compare a small sample of fields with the source page when the site changes or the output suddenly shifts.
- Keep request rates conservative: follow the site’s stated limits and stop if it signals that automated access is unwanted.
Do not infer that a robots.txt allowance means unlimited request volume. Nor should a parser respond to blocking by rotating identities or trying to defeat access controls; seek an approved access method instead.
Should you use Scrapy or Playwright?
Choose based on what the job actually needs; there is no supported blanket winner or head-to-head speed, cost, or reliability figure here.
| Approach | Use it when | Trade-off to consider |
|---|---|---|
| Requests plus an HTML parser | The needed content is already in the fetched HTML and the task is small or straightforward. | You must build any larger crawl’s scheduling, retry policy, logging, and operational safeguards yourself. |
| Scrapy | You have repeated crawling work and want a framework with downloader middleware and request-management features. | Learn its project and settings model; verify that policy-related settings are enabled as intended. |
| Playwright | The workflow depends on browser-rendered interactions or browser-issued network activity. | Running a browser adds operational overhead compared with a direct HTTP request and parse. |
What Scrapy handles
Scrapy’s downloader middleware can filter requests disallowed by robots.txt when the middleware and ROBOTSTXT_OBEY setting are enabled. Scrapy’s documentation identifies Protego as the default robots.txt parser. Check the settings for your project and version rather than assuming the policy check is active. See the Scrapy downloader middleware documentation.
What Playwright adds
Playwright’s Python Request API exposes request, response, completion, and failure events. A completed request is not necessarily a successful page fetch in the meaning your extractor needs: HTTP statuses such as 404 and 503 still complete as HTTP responses. Inspect response status explicitly. See the Playwright Request API.
Or skip the browser setup
If the goal is a clean visual capture rather than extracting structured fields, ScreenshotNeo can return a screenshot or PDF from one GET request. It is not a substitute for parsing records from HTML. For screenshot-based checks, it accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be switched off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo and its API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Sign up for 1,000 free screenshots a month, with no card required.
Troubleshooting common failures
403, 429, or a bot-check page
The site may restrict automated access or be applying a request limit. Confirm the published access policy and stop if access is not permitted. Do not try to bypass a CAPTCHA or access restriction. If the site offers an API or permission process, use it.
404 or another HTTP error
raise_for_status() reports unsuccessful HTTP status codes as request errors, so the example does not parse the error page as a record. Check whether the URL changed, a redirect led elsewhere, or the resource is unavailable. Log the final response URL and status when diagnosing the problem.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Missing or empty fields
First inspect the response HTML. If the field is absent there, the site may populate it with browser-side rendering; a direct request-and-parse template cannot extract content it never received. If the field exists, adjust the selector and test it against several representative pages. A page redesign can also invalidate previously working selectors.
Timeouts or connection errors
Check connectivity, the target URL, and whether the host is responding. The example’s 10-second connection and 30-second read timeout values are configuration choices, not guarantees. Increase them only when justified, and keep a finite timeout so a run cannot hang indefinitely.
Unexpected file type or encoding
The template rejects responses whose content type does not include HTML. Check redirects and response headers before changing that guard. If the requested resource is a PDF, image, or API response, use a parser appropriate to that format rather than treating it as a web page.
Robots check cannot be read
The example stops when it cannot retrieve or parse the robots.txt location. Check the exact origin, network access, and whether the site has published an alternate policy. Do not convert an uncertain check into assumed permission; consult the site’s instructions or ask for authorization.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Performance, reliability, and cost
For a small job, a direct HTTP request avoids the extra runtime and maintenance of a browser. Browser automation is appropriate only when rendered interactions or browser network events are necessary. Scrapy becomes useful when repeated crawling needs a framework’s request management and middleware. These are workflow distinctions, not measured performance claims.
Best Value
Every request consumes network and target-site resources, and a browser run has additional operational overhead. Keep the scope narrow, pace requests, cache or checkpoint data when appropriate, and stop on repeated failures. Public availability does not establish permission to collect or reuse content. The official documentation cited above describes technical behavior; it does not decide whether a particular scraping project is lawful.
Frequently Asked Questions
Does robots.txt give me permission to scrape a page?
No. It is crawler guidance, not access control or a grant of permission. Check the site’s terms and access instructions separately.
Can this Python template scrape content that appears only after JavaScript runs?
Not if that content is absent from the HTML response it fetches. Use a browser-based workflow only when permitted and when the rendered interaction is necessary.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →When should I move from a one-page script to Scrapy?
Consider a framework when you have repeated crawl work that benefits from request management, scheduling, and middleware; Scrapy’s robots filtering depends on the relevant middleware and setting being enabled.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

