What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape a website with Python, use the simplest permitted method that can reach the data: Requests to fetch HTML, Beautiful Soup to parse it, and explicit normalization, validation, and storage around the result. Move to Scrapy for a multi-page crawl, reproduce an underlying data request for JavaScript pages when practical, and use a headless browser such as Playwright only when the rendered DOM is genuinely required.

The examples below are illustrative. Run them only against a site you own, have permission to access, or that explicitly supports your intended use.

1. Choose a permitted target and define the output

Start with a site that authorizes your use. Check for an official API or documented feed before writing a scraper, read the site’s terms, inspect robots.txt, and collect only the fields you need. Robots rules communicate crawler preferences; they are not authorization and do not settle the legal position for a particular jurisdiction, dataset, access method, or intended use.

Define fields before selectors

Suppose each record needs a title, author, and detail-page URL. Write that schema first:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • title: a non-empty string
  • author: a string when present, otherwise a flagged missing value
  • url: an absolute HTTP(S) URL on an allowed host

This prevents a selector from quietly producing an attractive but unusable dataset. Decide how to represent missing values, duplicates, pagination, and failed records before collection starts.

2. Install the small static-page stack

Create an isolated environment and install the two libraries used for a one-page or small static collection:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4

Requests performs HTTP fetching; Beautiful Soup parses and searches the returned markup. Pin versions in your project once you have a repeatable build, and review their current documentation when you deploy.

3. Fetch a static page safely

Use a finite timeout and make HTTP failures visible. The timeout bounds how long the client waits; raise_for_status() turns 4xx and 5xx responses into exceptions instead of allowing bad input into the parser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
from bs4 import BeautifulSoup

url = 'https://example.com/page'  # illustrative; replace with an authorized target
response = requests.get(url, timeout=15)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')

print(soup.title.get_text(strip=True) if soup.title else 'No title')

Inspect the returned HTML in your browser’s developer tools or by saving a fixture. Replace the illustrative URL only after confirming that the target permits the request. A successful HTTP response can still contain an error page, a consent wall, or an empty application shell, so inspect the content you actually received.

4. Parse, normalize, validate, and store records

A maintainable scraper has separate stages:

  1. Fetch: request a URL with a timeout and status handling.
  2. Parse: locate elements with stable CSS selectors or XPath.
  3. Normalize: trim text, resolve relative links, and standardize values.
  4. Validate: check required fields, schemes, hosts, and types.
  5. Store: write structured output only after validation.

A complete resilient example

from __future__ import annotations

import csv
from urllib.parse import urljoin, urlparse

import requests
from bs4 import BeautifulSoup

START_URL = 'https://example.com/page'  # illustrative authorized target
ALLOWED_HOSTS = {'example.com'}


def clean_text(node):
    return node.get_text(' ', strip=True) if node else None


def allowed_http_url(value: str, base_url: str) -> str | None:
    absolute = urljoin(base_url, value)
    parsed = urlparse(absolute)
    if parsed.scheme not in {'http', 'https'}:
        return None
    if parsed.hostname not in ALLOWED_HOSTS:
        return None
    return absolute


def extract_records(html: str, page_url: str) -> list[dict[str, str | None]]:
    soup = BeautifulSoup(html, 'html.parser')
    records = []
    for card in soup.select('article'):  # replace after inspecting authorized markup
        title_node = card.select_one('h2, h3, .title')
        author_node = card.select_one('.author, [rel="author"]')
        link_node = card.select_one('a[href]')
        href = link_node.get('href') if link_node else None
        record = {
            'title': clean_text(title_node),
            'author': clean_text(author_node),
            'url': allowed_http_url(href, page_url) if href else None,
        }
        if record['title'] and record['url']:
            records.append(record)
    return records


def main() -> None:
    response = requests.get(
        START_URL,
        timeout=15,
        headers={'User-Agent': 'ExampleResearchBot/1.0 (contact: you@example.com)'},
    )
    response.raise_for_status()
    records = extract_records(response.text, response.url)

    # A simple output check catches a changed selector or an unexpected page.
    if not records:
        raise RuntimeError('No valid records found; inspect the response and selectors')

    with open('records.csv', 'w', newline='', encoding='utf-8') as output:
        writer = csv.DictWriter(output, fieldnames=['title', 'author', 'url'])
        writer.writeheader()
        writer.writerows(records)


if __name__ == '__main__':
    main()

The selectors are deliberately placeholders. Do not index an assumed first match such as select('article')[0]; a missing element should produce a flagged or skipped record, not crash an entire crawl. Resolve relative URLs against the response URL, reject unexpected schemes and hosts, and deduplicate records by a stable key such as the canonical URL.

Regression fixtures

Save a small authorized HTML response as a fixture and test extraction against it. A fixture-based check detects a changed class name before a scheduled job silently exports empty columns. Also log the number of fetched pages, parsed records, rejected records, and validation failures.

5. Add pagination and move to Scrapy when the job grows

For several pages, link following, retries, structured crawl state, and exports, use Scrapy rather than hand-rolling queues and crawl bookkeeping. Its workflow centers on a spider, initial requests, callbacks such as parse(), selectors, and yielded items.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Create a project and spider

python -m pip install scrapy
scrapy startproject quote_crawl
cd quote_crawl
scrapy genspider quotes example.com

The generated domain is only a scaffold. Replace it with an authorized practice target and its real host before running the spider.

import scrapy


class QuotesSpider(scrapy.Spider):
    name = 'quotes'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/quotes']  # illustrative

    def parse(self, response):
        for card in response.css('article'):
            title = card.css('h2::text').get()
            author = card.css('.author::text').get()
            href = card.css('a::attr(href)').get()
            if title and href:
                yield {
                    'title': title.strip(),
                    'author': author.strip() if author else None,
                    'url': response.urljoin(href),
                }

        next_href = response.css('a.next::attr(href)').get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

Run an export after replacing selectors and the target:

scrapy crawl quotes -O records.json

Use the Scrapy shell to inspect a response and refine CSS or XPath selectors:

scrapy shell 'https://example.com/quotes'
response.css('article h2::text').getall()
response.xpath('//article//h2/text()').getall()

Prefer .get() or .getall() and test for empty results instead of assuming a selector always matches. Scrapy’s tutorial makes the same resilience point: extraction should continue producing useful data when some page elements are absent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots and crawl controls in Scrapy

Identify your crawler with a descriptive User-Agent and configure Scrapy’s robots middleware to follow the site’s instructions. Keep concurrency and request volume proportionate to the task, add delays where appropriate, and stop when access is denied or disallowed. Following robots.txt is responsible engineering, but it is not a legal authorization.

6. Handle JavaScript-rendered pages without guessing

First, find the data request

If the initial HTML lacks the records, open the browser’s network panel, reload the page, and identify the request that returns the data. When an accessible JSON or HTML endpoint supplies the needed fields, reproduce that request with Requests or Scrapy, subject to the site’s permission and terms. This is usually simpler, cheaper, and more reliable than rendering every page.

Use a browser only when rendering is required

When the data exists only after browser execution or interaction, a headless browser such as Playwright for Python can render the page. Browser automation is a technical fallback, not a method for defeating bot checks, CAPTCHAs, authentication boundaries, or other restrictions.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto('https://example.com/page', wait_until='domcontentloaded', timeout=30_000)
    page.wait_for_selector('article', timeout=10_000)
    titles = page.locator('article h2').all_text_contents()
    print([title.strip() for title in titles])
    browser.close()

Install the browser binaries according to Playwright’s current Python documentation. Set finite navigation and selector timeouts, and close the browser in a finally block in production code so failed jobs do not leak processes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Be polite, secure, and predictable

  • Identify yourself: send a descriptive User-Agent with a contact address where practical.
  • Honor instructions: follow robots.txt and the site’s published usage limits; stop when access is denied.
  • Limit load: fetch only required pages, avoid unnecessary assets, and use measured concurrency.
  • Bound failures: set connect/read timeouts, handle non-success status codes, and record retries rather than looping forever.
  • Protect secrets: keep API keys, cookies, and authorization headers out of source control and logs.
  • Validate URLs: if URLs come from users or scraped content, allow only expected HTTP(S) schemes and hosts to reduce SSRF risk. Do not expose crawler control endpoints to untrusted networks.

Whether a particular collection is lawful depends on the target, data, jurisdiction, access method, contracts, and intended use. This tutorial is technical guidance, not jurisdiction-specific legal advice.

8. Which Python scraping library should you use?

Need Starting point Reason
One or a few static pages Requests + Beautiful Soup Separate HTTP fetching from HTML parsing with a small setup.
Multi-page crawl and structured workflow Scrapy Spiders, requests, callbacks, selectors, link following, and exports provide crawl state.
Dynamic page with an identifiable data source Reproduce the relevant request Request-level extraction avoids unnecessary browser rendering.
Browser-only behavior or rendered-DOM data Playwright or a Scrapy browser integration Use automation when request-level extraction is not practical.

Choose by page complexity, crawl scale, control over requests, setup time, and operational risk—not by an assumed universal speed ranking. Start small, measure response and parse failures, and promote the job to Scrapy or browser automation only when the requirements justify it.

9. Troubleshooting common failures

403, 429, or an access-denied page

Cause: the site rejected the request, rate-limited it, or requires an approved access path. Fix: stop or slow the job, verify permission and the published API, identify your crawler, and do not try to bypass the control.

200 response but no records

Cause: you received an application shell, consent page, changed markup, or the wrong URL. Fix: save the response, inspect its title and body, check network requests for a data endpoint, and update selectors only after confirming the new structure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Timeouts and intermittent connection errors

Cause: slow servers, overloaded concurrency, or an unbounded wait. Fix: use finite connect/read or navigation timeouts, modest retries with backoff, lower concurrency, and metrics that distinguish fetch failures from parse failures.

Relative or unsafe links

Cause: href values are fragments, protocol-relative URLs, or links to unexpected hosts. Fix: resolve with urljoin, allow only HTTP(S), enforce an allowed-host set, and discard or review anything outside it.

Fields disappear after a redesign

Cause: selectors depended on presentation classes or a fixed DOM position. Fix: prefer semantic attributes and stable containers, keep HTML fixtures, validate row counts and required fields, and alert on sudden changes.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

10. Performance, reliability, and cost planning

Request-level extraction normally uses fewer resources than launching a browser per page. Reuse HTTP sessions where appropriate, avoid downloading assets you do not parse, cache only when the site’s rules permit it, and keep a bounded queue. For Scrapy, tune concurrency conservatively and monitor memory, retries, status codes, and item counts. For browser jobs, reuse a browser process carefully, limit pages, and always close contexts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability comes from observable checkpoints: record the requested URL, final URL, status, elapsed time, parser version, extracted count, and validation errors. Store raw responses or hashes when your retention policy allows it so a failed parse can be reproduced without immediately re-requesting the site. There is no meaningful universal speed or success statistic here; the target, network, markup, and access policy determine results.

Or skip the browser setup

If your goal is a clean screenshot rather than parsed fields, ScreenshotNeo provides a website screenshot API and MCP server. Its request accepts a URL and returns PNG, JPEG, WebP, or PDF. Before capture it can accept the cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Use the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom JavaScript and CSS, pre-capture clicks, hide selectors, selector/delay/network-idle waits, ad/tracker/request/resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, 100-URL bulk capture, usage reporting, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work.

One-call examples

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing gives two months free. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Create a free ScreenshotNeo account to try it without a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

How should I store scraped data for later analysis?

Use a stable schema and an explicit encoding such as UTF-8. CSV is convenient for flat records; JSON preserves nested values. Keep the source URL and collection timestamp with each record so downstream users can trace provenance.

How can I test a scraper without repeatedly contacting a live site?

Save a permitted response as an HTML fixture and run your parser against that file in automated tests. Add cases for missing fields, malformed links, duplicate records, and a changed layout.

When should a scraper become a scheduled job?

Schedule it only after selectors, validation checks, rate limits, error handling, and alerting are in place. A job that exports an empty file successfully is still a failed job.

Can I scrape authenticated pages?

Only with explicit authorization and a permitted access method. Treat session cookies and Authorization headers as secrets, avoid logging them, and respect the account’s terms and data boundaries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.