What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Beautiful Soup parses HTML; it does not download web pages or run JavaScript. A typical scraper uses an HTTP client such as Requests or Python’s urllib to retrieve permitted page markup, then passes that markup to Beautiful Soup to locate and extract the fields you need. This tutorial walks through that workflow, from installation and parser choice to handling missing fields and pages whose content depends on JavaScript.
What Is Web Scraping?
Web scraping is the process of retrieving information from a web page and extracting selected data from its content. For a static HTML page, the workflow has two distinct parts: an HTTP client requests the page, and a parser turns the returned markup into a structure your code can navigate.
Beautiful Soup is the parsing part. The Beautiful Soup project describes it as “a Python library for pulling data out of HTML and XML files.” It creates a navigable representation of the markup; it does not make the network request, decide whether access is permitted, or execute the page’s JavaScript.
What is the difference between requests and BeautifulSoup?
Requests is an HTTP client: it sends a request and gives your program the response, including the page’s HTML. Beautiful Soup parses that HTML into a tree so you can find elements and read their text or attributes. Installing Beautiful Soup does not install Requests, and either library can be used without the other.
#1 Best Overall
Python’s standard library also provides urllib.request for fetching URLs. For content that is already saved in a file or string, you can skip the HTTP step and pass the HTML directly to Beautiful Soup.
Install Beautiful Soup and choose a parser
For new Python code, install the beautifulsoup4 distribution and import the class from bs4. Do not install the similarly named BeautifulSoup package by mistake: that is the older Beautiful Soup 3 line, which the project documentation says is no longer developed or supported. Check the project’s official documentation for current installation guidance and version-specific behavior.
python -m pip install beautifulsoup4 requests lxml
This command installs Requests and lxml as well as Beautiful Soup. If you choose a different parser, install its package where necessary. Then import Beautiful Soup like this:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
from bs4 import BeautifulSoup
Beautiful Soup supports three commonly used parser choices. They can build different trees from malformed HTML, so specify one explicitly when you need consistent results across machines.
| Parser | What to know |
|---|---|
lxml |
The Beautiful Soup manual ranks it first among the named parsers and describes it as significantly faster than the other two. Install the lxml package and select it explicitly. |
html5lib |
Uses HTML5 parsing techniques. Its interpretation of malformed markup may differ from the other parsers. |
html.parser |
Python’s built-in HTML parser; it does not require installing a separate parser package. |
There is no universally correct tree for every invalid document. The parser choice can affect which elements your selectors find. Use the same named parser in development and deployment, and consult the Beautiful Soup manual when exact parser behavior matters.
Fetch a permitted practice page
Before making a request, check the site’s terms and robots.txt, and choose a practice target that permits the access you plan. These are practical safeguards, not a complete legal test; rules can depend on the site, data, and jurisdiction. Do not try to evade access controls, and stop if the site disallows your planned access.
The example below uses a target page only as a placeholder for a permitted practice URL. Replace it with a page you are allowed to retrieve. The code checks the HTTP status before parsing the response body.
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "lxml")
For a reproducible project, keep the selected parser installed in every environment. A timeout also prevents the request from waiting indefinitely. Do not disguise your client to bypass a site’s restrictions; request headers should be used for legitimate request configuration, not access-control circumvention.
Find elements and extract fields
Use tag names, attributes, or CSS selectors to locate content. For example, when a page contains article cards with a headline link, you can collect each card’s title and destination. Adjust the selector to match the actual permitted page’s HTML.
cards = soup.select("article.card")
items = []
for card in cards:
link = card.select_one("h2 a")
if link is None:
continue
items.append({
"title": link.get_text(" ", strip=True),
"url": link.get("href"),
})
select() returns all matching elements, while select_one() returns the first match or None. Checking for a match before reading from it avoids errors when a field is absent. get_text(" ", strip=True) joins text with spaces and trims surrounding whitespace; get("href") retrieves the link attribute and returns None if it is missing.
Beautiful Soup also provides methods such as find() and find_all(). Choose whichever makes the selection clear, then validate the result before extracting values. Collect only the fields you need and handle missing or changed markup as an expected possibility.
Why does my scraper return an empty list?
An empty result means your selection did not match the tree Beautiful Soup received. It does not necessarily mean the page has no visible content. Check the stages in order:
Best Value
- Confirm the response. Check the final URL and HTTP status, and inspect a short portion of
response.text. A redirect, error page, or different response may not contain the expected markup. - Check the selector against the HTML. Inspect the returned HTML for the exact tag, class, and attributes. A site may have changed its markup, or the selector may describe a different page structure.
- Check parser consistency. Make sure you are selecting the same parser in all environments. Different parsers can produce different trees from malformed markup.
- Check whether the content is rendered by JavaScript. If the HTML response does not contain the data, Beautiful Soup cannot extract it from that response. Look for an official API or data export; use a rendering or browser automation tool only if access is permitted and rendered page state is genuinely needed.
A quick diagnostic can reveal whether your selector matches anything:
print(response.status_code)
print(response.url)
print(response.text[:500])
print(len(soup.select("article.card")))
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Save only the data you intend to use
Once the extraction is correct, write the selected fields to a format your next step can consume. For example, this saves the list of dictionaries as JSON:
import json
with open("items.json", "w", encoding="utf-8") as file:
json.dump(items, file, ensure_ascii=False, indent=2)
Before relying on the output, check that required fields are present and that the result contains the expected kinds of values. Page markup can change, so a scraper should fail clearly or flag incomplete records rather than silently treating missing data as valid.
Recommended Free Tools
When static HTML is not enough
Beautiful Soup parses the markup you give it; it does not run JavaScript or reproduce a browser session. A page may show data in a browser that is absent from the HTML returned to your script. In that case, first check whether the site offers an official API or data export. If the data truly depends on rendered DOM state, a browser-rendering or automation approach may be relevant only when the site permits that access.
Martin Breuss’s Real Python tutorial, published December 1, 2024, similarly distinguishes retrieving static HTML with Requests from parsing it with Beautiful Soup, and notes that dynamic pages can require additional tools. The distinction remains important: changing the parser cannot make JavaScript-generated content appear in an HTTP response that never contained it.
Quick Recap
Scrape responsibly
- Review the site’s terms and
robots.txtbefore collecting data, and stop when the planned access is disallowed. - Use training targets or other pages that expressly permit practice.
- Request only the fields needed for your task; avoid collecting personal data or content behind a login.
- Do not work around authentication, rate limits, or other access controls.
- Do not assume that scraping is always lawful or always unlawful. Terms, privacy, copyright, and other obligations can depend on the facts and jurisdiction.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

