Collect website data by first defining the records you need, then choosing the least complex reliable access path: an official API or feed, ordinary HTTP plus an HTML parser, a crawler such as Scrapy, or a headless browser only when browser execution is genuinely required. Fetch the relevant pages, extract stable fields, validate every record, and store the result with enough source context to audit it.
1. Define exactly what you are collecting
Write a short collection specification before writing code. It prevents an apparently successful crawler from producing data that cannot answer the real question.
Set the scope
- Sites and URLs: list the domains, starting pages, URL patterns and page types that are in scope.
- Fields: name each field and its expected type. For a product record, that might be
name,price,currency,availabilityandsource_url. - Frequency: decide whether this is a one-time export, a daily job or an event-driven update.
- Output: choose JSON Lines, CSV, XML or a database according to how the next system consumes the data.
- Quality rules: define required fields, acceptable missing values, duplicate handling and how you will record collection time.
Keep the scope narrow enough to explain why each page and field is needed. Scrapy’s tutorial illustrates the pattern: select named fields, follow only the relevant pagination link and export one structured item per record.
2. Use an official API or feed when one fits
Look for a documented first-party API, RSS/Atom feed or downloadable data file before parsing page markup. An API normally gives more stable field names and clearer access conditions than HTML. Check authentication, quotas, permitted uses, update cadence and the fields actually returned. Scrapy can request APIs as well as crawl HTML, so a crawler framework is not a reason to ignore a supported endpoint.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
If the official interface lacks a required field, document that gap before falling back to page extraction. Do not assume that a publicly visible page grants permission to use every field for every purpose.
3. Choose the simplest technical approach
| Approach | Use it when | Trade-offs |
|---|---|---|
| Official API or feed | A documented interface contains the required data. | Fields, quotas, terms and update timing are specific to that site. |
| HTTP client plus parser | A small job needs data already present in the HTML response. | Simple and fast to start, but you must add pagination, retries, scheduling and export handling. |
| Scrapy | A repeatable crawl needs selectors, pagination, exports and request controls. | More framework structure, with built-in feed exports and crawl controls. |
| Headless browser | The data or required output depends on browser execution. | More setup and resource use; inspect the underlying request first. |
| Hosted extraction API | You need managed execution and an exportable result. | Compare coverage, data quality, access terms, cost and availability for your case. |
4. Collect ordinary HTML with Python
When the required fields are in the initial response, an HTTP client and parser are enough. The following example requests a page, selects article cards, normalizes whitespace and writes JSON Lines. Replace selectors with ones verified against the target site’s markup.
import json
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
START = "https://example.com/articles"
headers = {"User-Agent": "data-collector/1.0 (contact: you@example.com)"}
response = requests.get(START, headers=headers, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
with open("articles.jsonl", "w", encoding="utf-8") as out:
for card in soup.select("article.card"):
title_node = card.select_one("h2")
link_node = card.select_one("a")
if not title_node or not link_node:
continue
record = {
"title": " ".join(title_node.get_text(" ", strip=True).split()),
"url": urljoin(START, link_node.get("href", "")),
"collected_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"source_url": START,
}
out.write(json.dumps(record, ensure_ascii=False) + "n")
Use CSS selectors for classes, attributes and relationships. XPath is useful when an element is easier to identify by text or position. Beautiful Soup and lxml are parsers; they do not by themselves provide a scheduler, crawl frontier, retry policy or export pipeline.
5. Scale a repeatable crawl with Scrapy
Scrapy combines selectors, request scheduling, pagination, throttling and feed exports. A minimal spider looks like this:
import scrapy
class QuoteSpider(scrapy.Spider):
name = "quotes"
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
"source_url": response.url,
}
next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
Run it with a feed export such as scrapy crawl quotes -O quotes.jsonl. Scrapy also supports JSON, CSV and XML feeds and item pipelines for validation or storage. Configure a download delay, a sensible per-domain concurrency limit and automatic throttling. Follow only links that can lead to records; never let a generic “follow every link” rule create an unbounded crawl.
6. Find the real source of JavaScript-loaded data
A browser may show a table that is absent from the initial HTML. Treat this first as a source-discovery problem, not as proof that a browser is required.
- Open the browser’s developer tools and select the Network panel.
- Reload the page and filter requests by Fetch/XHR, then inspect responses while the missing field appears.
- Identify the request returning JSON, HTML, embedded JavaScript data or another structured response.
- Reproduce that request with the required method, query parameters, headers, cookies or token, subject to the site’s access conditions.
- Parse and validate the response. Record the endpoint and collection time so a later change is diagnosable.
Scrapy documentation states: “When this happens, the recommended approach is to find the data source and extract the data from it.” Use a headless browser when the request cannot reasonably be reproduced, authentication requires a browser flow, or the deliverable is the browser-rendered result itself. Do not use browser automation to defeat a CAPTCHA, bot check or other access control.
7. Validate, normalize and store records
Normalize at the boundary
- Trim and collapse whitespace; normalize field names and date formats.
- Parse numbers with an explicit locale and preserve the original currency.
- Convert relative links to absolute URLs and retain the source URL.
- Represent missing values consistently instead of silently converting them to empty strings.
Validate before persistence
- Reject or quarantine records missing required identifiers.
- Check types, ranges and allowed values.
- Deduplicate using a stable key such as a first-party ID plus source URL.
- Log response status, parser version, collection timestamp and validation errors.
Store raw responses or a permitted excerpt when auditability requires it, while applying an appropriate retention policy. There is no universally best database: choose based on volume, update patterns, query needs and sensitivity.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
8. Pagination, retries and performance
Pagination
Prefer the site’s explicit next-page link or documented cursor. Stop when the cursor is absent, a page repeats, or a defined maximum is reached. Do not infer that every numeric URL is valid.
Retries and backoff
Retry transient network failures and selected 5xx responses with exponential backoff and a maximum attempt count. Do not aggressively retry authentication failures, validation errors or deliberate access denials. Set connection and total timeouts, and record failures for a later run.
Rank #3
Rate and concurrency
Keep request rates proportionate to the site. Scrapy provides download delays, per-domain concurrency limits and automatic throttling; configure them rather than relying on defaults. Reuse connections, avoid downloading assets you do not parse, and cache responses during development when permitted.
Repeatability
Pin your parser and crawler dependencies, keep selectors in version control, and add fixtures for representative pages. A small change to a class name can otherwise produce an empty file with no obvious exception.
9. Responsible access: robots.txt, terms and privacy
Check the target site’s terms and documented access routes. Review its robots.txt instructions and configure your crawler to honor applicable rules. Keep fields and request rates to what the stated purpose requires, and do not attempt to defeat access controls.
Robots.txt is a crawler-access instruction, not authentication, a security boundary or a complete legal decision. Google notes that a blocked URL may still appear in search if linked elsewhere; password protection or noindex serves different goals. Whether a particular collection is permitted depends on the data, site, permissions, jurisdiction and intended use. The 2024 U.S.-focused social-science legal research describes these as fact-specific legal, ethical, institutional and scientific questions, not a universal yes/no rule.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server when your collection needs a rendered page image or PDF rather than parsed records. One GET request returns PNG, JPEG, WebP or PDF. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
For a screenshot of Stripe, see the ScreenshotNeo API documentation:
Free tools Windows power users keep installed
One-click scans. No signup required.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Options include full-page lazy-image loading, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper settings and page ranges, custom CSS/JavaScript, clicks, waits, request blocking, headers/cookies/user agents, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and annual billing gives two months free. Create a free ScreenshotNeo account to try it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.10. Troubleshooting common failures
The response is 200 but records are empty
Inspect the saved HTML, not just the browser view. If the elements are absent, locate the network request that supplies them. If they are present, update selectors and add a fixture test.
Many requests return 403 or 429
Stop and review terms, robots instructions, authentication and rate settings. Reduce concurrency, add delay and backoff, and use an official interface if available. Never try to bypass a challenge.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePagination loops forever
Log every requested URL or cursor, detect repeats, honor a maximum page count and follow only the site’s explicit next link.
Best Value
Prices or dates are wrong
Check locale, currency, timezone and whether the value is rendered after a request. Preserve the raw text alongside the normalized value and test pages from each relevant locale.
The browser automation job times out
Set a realistic navigation timeout, wait for a specific selector or network-idle condition, block unneeded resources, and capture diagnostics. If a direct data request exists, replace browser automation with that request.
FAQ
Is collecting public website data always legal?
No universal rule applies. Evaluate the site’s terms, permissions, data sensitivity, access controls, intended use and the law applicable to your jurisdiction.
Should I save HTML as well as parsed fields?
Save enough raw context to audit and reproduce a result when your policy and the site’s terms permit it; limit retention of unnecessary or sensitive data.
When is a screenshot useful instead of extracted data?
Use a screenshot or PDF when the deliverable is visual evidence, a rendered layout or an archival view. For searchable records, identify and parse the underlying API or HTML response instead.
Frequently Asked Questions
Is collecting public website data always legal?
No universal rule applies. Evaluate the site’s terms, permissions, data sensitivity, access controls, intended use and the law applicable to your jurisdiction.
Should I save HTML as well as parsed fields?
Save enough raw context to audit and reproduce a result when your policy and the site’s terms permit it; limit retention of unnecessary or sensitive data.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhen is a screenshot useful instead of extracted data?
Use a screenshot or PDF when the deliverable is visual evidence, a rendered layout or an archival view. For searchable records, identify and parse the underlying API or HTML response instead.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

