Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape a web page with Python, fetch its HTML with an HTTP client, check the response, parse the document, select the fields you need, and save the results. For a small, mostly static task, requests plus Beautiful Soup is a straightforward starting point. Use Scrapy when you need a repeatable multi-page crawler; use Playwright only when the required content or interaction genuinely depends on a browser. The example below demonstrates the full retrieval-and-parsing loop, including basic checks and CSV output.

How web scraping works

Scraping is a sequence of distinct steps, not a single Python command:

  1. Request: An HTTP client asks a server for a page.
  2. Response: The server returns a status code, headers and a body. The body may contain HTML, or it may be an error, a redirect or something other than the page you expected.
  3. Parse: An HTML parser turns the body into a tree of elements that code can search.
  4. Select and extract: You identify elements that represent the information you want, then read their text or attributes.
  5. Validate and save: You check that the extracted values make sense and write them to a useful format such as CSV or JSON.

Requests handles HTTP; Beautiful Soup handles parsing and searching. Their jobs are separate, so a successful request does not guarantee that the data is present or that your selectors match the page. See the Requests Quickstart and Beautiful Soup documentation for their APIs.

Before you scrape: choose an appropriate access method

First check whether the site offers an official API or data feed that provides the fields you need. A supported interface is usually less fragile than parsing page markup. If you do scrape pages, review the site’s current terms and instructions, confirm that your use is authorized, and consider applicable privacy, data-protection, copyright and database-rights rules. Publicly viewable content is not, by itself, a universal determination of permission or legality.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Identify your crawler with a descriptive User-Agent and a way for the site operator to contact you.
  • Keep the scope limited to pages and fields you actually need. Store only necessary data.
  • Use conservative request pacing and stop or adjust if the site objects, denies access or returns signs of overload.
  • Do not try to bypass access controls, CAPTCHAs or other restrictions. If permission or terms are unclear, ask the operator or use a supported access route.

A robots.txt file expresses crawling instructions, but it is not legal advice or proof that a particular use is permitted. A basic Requests script does not automatically obey it. Scrapy can filter disallowed paths when its RobotsTxtMiddleware is enabled and ROBOTSTXT_OBEY is set; read the Scrapy downloader middleware documentation for configuration and behavior.

Build a small scraper with Requests and Beautiful Soup

Use a page you are allowed to access and whose returned HTML contains the information you want. This example uses the Scrapy tutorial as a practice page: it fetches the document, checks the response, extracts the page title and links, and writes a CSV. It is deliberately a small demonstration, not a promise that the same selectors fit another site’s markup.

1. Install the packages

In a terminal with Python available, install the two dependencies:

python -m pip install requests beautifulsoup4

2. Fetch, inspect, parse and save

import csv
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

url = "https://doc.scrapy.org/en/master/intro/tutorial.html"
headers = {"User-Agent": "LearningScraper/1.0 (contact: you@example.com)"}

response = requests.get(url, headers=headers, timeout=20)
response.raise_for_status()

content_type = response.headers.get("Content-Type", "")
if "html" not in content_type.lower():
    raise ValueError(f"Expected HTML, got {content_type!r}")

soup = BeautifulSoup(response.text, "html.parser")
page_title = soup.title.get_text(" ", strip=True) if soup.title else ""

if not page_title:
    raise ValueError("No page title found; check the response and page structure")

rows = []
for link in soup.select("a[href]"):
    label = link.get_text(" ", strip=True)
    href = link.get("href", "").strip()
    if label and href:
        rows.append({"text": label, "url": urljoin(response.url, href)})

if not rows:
    raise ValueError("No labeled links found; inspect the HTML and adjust the selector")

with open("scraped_links.csv", "w", newline="", encoding="utf-8") as file:
    writer = csv.DictWriter(file, fieldnames=["text", "url"])
    writer.writeheader()
    writer.writerows(rows)

print(f"Title: {page_title}")
print(f"Saved {len(rows)} links to scraped_links.csv")

Replace the contact text with a real contact method you control when identifying an actual crawler. raise_for_status() raises an exception for unsuccessful HTTP responses rather than letting the script quietly parse an error page. The timeout prevents a request from waiting indefinitely. Checking the content type and requiring a title are simple validation guards; production code should validate the particular fields and formats it depends on.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Select fields narrowly

The example selects all anchors with an href attribute using CSS selector a[href]. For real extraction, inspect the returned HTML and scope a selector to a meaningful container so navigation, footer and unrelated links do not become records. For example, if each record is in an element with class product, select .product containers first, then find a title and price inside each one. Those class names are illustrative; use the ones present in the page you are authorized to process.

for card in soup.select(".product"):
    title_el = card.select_one(".product-title")
    price_el = card.select_one(".price")

    title = title_el.get_text(" ", strip=True) if title_el else None
    price = price_el.get_text(" ", strip=True) if price_el else None

    if title is not None:
        print({"title": title, "price": price})

Text and attributes are different kinds of data. Use get_text(" ", strip=True) to normalize whitespace between text nodes; use element.get("href") or element.get("src") to read an attribute. Resolve relative links against the final response URL with urljoin, as the complete example does.

CSS or XPath?

CSS selectors are concise for matching elements by tag, class, attribute or their relationships. XPath is useful when selection depends on traversal or predicates that are awkward to express in CSS. Scrapy’s selector guide documents both CSS and XPath, and explains that its selectors use Parsel, which uses lxml. It also notes that Beautiful Soup handles imperfect markup reasonably well but has a speed drawback in that comparison; that is not a universal timing result, so test your own workload before choosing on performance alone. See Scrapy Selectors.

Make extraction resilient before collecting more pages

Page markup changes, fields may be absent, and a selector can match the wrong thing without raising an error. Treat extracted data as untrusted until checked.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use select_one() for an optional single element and test for None before reading it.
  • Normalize whitespace, and keep missing values explicit rather than silently substituting plausible-looking text.
  • Check that a record has its required fields and that values have the expected shape before writing it.
  • Print or inspect a few records during development. Compare them with the page and the source HTML, not just with a successful script exit.
  • Prefer stable semantic structure over brittle positional selectors such as “the third paragraph.” Revisit selectors when the site’s layout changes.

For a small dataset, CSV is convenient for tabular records. JSON is often a better fit for nested fields. In either case, preserve a clear distinction between a missing value, an empty string and a value that was actually observed.

Follow pagination without crawling indefinitely

If a page links to a next page, follow that link only while it exists and while the crawl remains within the scope you intended. Relative URLs need to be resolved against the current page. For a simple sequential loop, track visited URLs to avoid cycles and impose a maximum page count as a safety stop:

from urllib.parse import urljoin

current_url = start_url
visited = set()
max_pages = 10

for _ in range(max_pages):
    if not current_url or current_url in visited:
        break
    visited.add(current_url)

    response = requests.get(current_url, headers=headers, timeout=20)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")

    # Extract and validate this page's records here.
    next_link = soup.select_one("a[rel='next']")
    current_url = urljoin(response.url, next_link["href"]) if next_link else None

start_url and the record-extraction block must be supplied for your task. Many sites use a different next-link structure, a button, or an API; inspect the page and choose a permitted, reliable stopping rule rather than assuming rel="next" is universal. The Scrapy tutorial demonstrates link following and extracting items into dictionaries as part of a spider workflow.

Choose the tool that fits the page and workload

Situation Starting choice Why
A few pages with the needed content in the returned HTML Requests plus Beautiful Soup or lxml Simple retrieval and parsing, with a selector style suited to the markup.
Many pages, pagination, recurring runs or structured exports Scrapy A project and spider workflow supports requests, parsing, link following, feed exports and crawl controls.
Content appears only after browser-side JavaScript or interaction Playwright for Python, if permitted Browser automation can observe rendered-page behavior and network activity. First check for an authorized API or data feed.
An official API supplies the records The API, subject to its terms It may be more stable and appropriate than parsing pages.

Choose based on where the data exists, how many pages you need, pagination complexity, interaction requirements, selector maintenance, request controls, export and monitoring needs, and applicable terms and permissions. A browser is heavier than a direct HTTP request, so it should not be the default just because a page looks dynamic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Scrapy is the next step

Scrapy is a good fit when the task becomes a repeatable crawl rather than a short script: define a project and spider, issue requests, parse responses, yield dictionaries or items, follow links and export results. Its tutorial walks through that workflow. Scrapy also documents download delays, per-domain concurrency limits and AutoThrottle as controls in its overview. Concurrency is a technical setting, not permission; choose conservative values appropriate to the site and task.

When Playwright is warranted

Use browser automation only when the needed data is absent from the initial response or access to it requires browser behavior you are authorized to perform. Before launching a browser, inspect whether a supported API or data feed is available. Playwright’s Python Request API documents request and response information, redirects and resource details that can help you understand browser network activity. Not every JavaScript site requires browser-based scraping.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a page image or PDF rather than structured records, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is not a substitute for extracting structured fields from HTML. For a one-call screenshot, create an API key and use this cURL example; API options and response details are in the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://doc.scrapy.org/en/master/intro/tutorial.html -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up free for 1,000 screenshots a month, with no card required.

Troubleshoot common problems

  • HTTP error: raise_for_status() reports a non-success response. Check the status, URL and any redirect or access restriction. Do not try to evade a denial; use an authorized route or stop.
  • Timeout: The server or network did not respond within the configured time. Retry cautiously if appropriate, and do not create a rapid retry loop.
  • Expected elements are missing: Inspect the response body. The page may have changed, returned an error page, or require JavaScript. Confirm whether the data is available through an authorized API before considering browser automation.
  • Selector returns no matches: Check spelling, scope and the actual HTML. A visual label on the page does not guarantee the same text or class is in the response.
  • Duplicate or malformed rows: Scope selection to each record container, normalize text, validate required fields and use a visited-URL set for pagination.
  • CSV looks corrupted: Open the file as UTF-8 and ensure it is written with newline=""; check that each row uses the same field names.

Keep a crawler reliable and considerate

For a one-off script, a timeout, status check, narrow scope and deliberate pacing may be enough. For recurring or multi-domain work, log requested URLs, response outcomes and extraction failures so markup changes do not silently degrade your output. In Scrapy, configure its documented delay and per-domain concurrency controls; enable robots filtering deliberately and understand the middleware’s user-agent matching. The project tutorial recommends setting a descriptive User-Agent so an operator can contact the crawler owner: “Website owners who take issue with your crawler can then ask you to adjust it, rather than block it.” See the Scrapy tutorial.

Respond to objections, access denials and rising error rates by pausing and reassessing. Do not scale a crawl merely because the code can send more requests. The appropriate pace, access method and data handling depend on the site, your authorization, the information collected, your jurisdiction and intended use.

Frequently Asked Questions

Does Python have a built-in web scraper?

No single built-in function performs the whole job. A typical small scraper combines an HTTP client such as Requests with an HTML parser such as Beautiful Soup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I scrape any public website?

There is no universal yes-or-no answer. Permission and obligations depend on the site’s terms, access method, data and use, and applicable law; public visibility alone does not settle them.

What should I learn after a first scraper?

Learn to validate extracted records, handle pagination safely, and use a crawler framework such as Scrapy when the job needs repeatable multi-page scheduling and exports.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.