For most readable-text jobs, parse the markup with Beautiful Soup and call get_text(" ", strip=True). The separator keeps words from different inline elements from running together, while strip=True removes leading and trailing whitespace. Choose and name the parser explicitly—usually lxml, or the standard-library html.parser when you want no third-party parser dependency.
Choose the extraction approach
Your best method depends on whether you need a quick conversion of a complete document, control over one element, or a dependency-free event stream.
| Approach | Strength | Trade-off | Best fit |
|---|---|---|---|
Beautiful Soup + lxml |
Friendly tree API with a robust parser backend | Requires extra dependencies | General extraction from messy pages |
Beautiful Soup + html5lib |
HTML5-style parsing and browser-like error recovery | Usually slower and adds a dependency | Input where recovery behavior matters |
Beautiful Soup + html.parser |
Simple installation and a familiar API | Different recovery behavior on invalid markup | Small scripts and controlled input |
html.parser.HTMLParser |
Python standard library and callback control | You implement collection and cleanup | Low-level or dependency-light processing |
Beautiful Soup supports all three parser choices. The same malformed HTML can produce different trees with different parsers, so specify the parser in code and pin it in your project requirements when reproducibility matters.
Extract readable text with Beautiful Soup
Install the libraries
python -m pip install beautifulsoup4 lxml
Then pass the HTML string to Beautiful Soup and select the parser by name:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
from bs4 import BeautifulSoup
html = """
<article>
<h1>Release notes</h1>
<p>Version <strong>2.0</strong> is available.</p>
</article>
"""
soup = BeautifulSoup(html, "lxml")
text = soup.get_text(" ", strip=True)
print(text)
# Release notes Version 2.0 is available.
get_text() returns text beneath the document or tag. Its first argument is a separator inserted between text fragments; a single space is a practical default for prose. strip=True trims surrounding whitespace from the result.
Extract only the content element
Calling get_text() on the whole document also collects navigation, footers, cookie notices, comments and other page furniture. If the page has a reliable container, select it first:
main = soup.select_one("main")
if main is None:
raise ValueError("Expected a <main> element was not found")
article_text = main.get_text(" ", strip=True)
You can use any CSS selector supported by Beautiful Soup, such as article, #post or .entry-content. Check for None before calling a method so a changed template produces a useful error instead of an attribute exception.
Process fragments with stripped_strings
Use stripped_strings when you need to inspect, filter or transform each text fragment yourself:
Free tools Windows power users keep installed
One-click scans. No signup required.
parts = [fragment for fragment in soup.stripped_strings]
for part in parts:
print(part)
This is useful when you want to discard a particular heading, normalize selected fragments differently, or preserve a list of pieces for later processing. Join the final list with a separator appropriate to your output format.
Rank #2
Use the standard library when you cannot add a parser
Python’s html.parser module provides an event-driven parser. Callbacks receive text and markup events, so you decide exactly what to retain.
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
def handle_data(self, data):
self.parts.append(data)
html = "<div>Hello <em>Python</em></div>"
extractor = TextExtractor()
extractor.feed(html)
text = " ".join(" ".join(extractor.parts).split())
print(text)
# Hello Python
HTMLParser is a parser, not a downloader: it processes the string or bytes you feed it. The example collects every data fragment, then collapses runs of whitespace. For production code, add callbacks such as handle_starttag and handle_endtag when you need to ignore selected regions or track nesting.
Control whitespace without losing words
HTML often separates a sentence across nested tags. Calling get_text() without a separator can concatenate those fragments. Prefer:
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchestext = soup.get_text(" ", strip=True)
For output that should retain paragraph boundaries, iterate through block elements and join them with newline characters rather than treating the entire document as one stream:
paragraphs = [
p.get_text(" ", strip=True)
for p in soup.select("p")
]
text = "nn".join(p for p in paragraphs if p)
Do not assume every newline in source HTML represents a visible line break. Markup indentation is not the same as rendered layout; choose separators based on the structure you intentionally select.
Exclude non-readable regions deliberately
When a page contains scripts, styles, templates or repeated interface elements, target the content subtree instead of blindly flattening the document. With commonly used Beautiful Soup parsers, script, style and template contents are generally not treated as human-readable text. You should still verify the result against representative pages because menus, cookie banners, comments and duplicated responsive markup may remain.
for unwanted in soup.select("nav, footer, aside, .cookie-banner, .comments"):
unwanted.decompose()
clean_text = soup.get_text(" ", strip=True)
decompose() removes the selected nodes from the parse tree. Use selectors that match your own page templates; a class such as .comments is not universal. If you need to retain the original tree, make a separate soup object before removing nodes.
Parser choice and reproducibility
lxml
Use lxml for a convenient Beautiful Soup tree API backed by a robust parser. It is a strong general-purpose choice when inputs are inconsistent and adding a dependency is acceptable.
html5lib
Choose html5lib when browser-like HTML5 error recovery is important. Its behavior can be useful for badly formed documents, but it is usually slower and adds another dependency.
html.parser
Beautiful Soup can use Python’s built-in parser, which keeps installation simple. Its recovery of malformed markup can differ from lxml or html5lib.
Pin and test the behavior
Parser selection is observable behavior: malformed markup can create different trees. Name the parser in every constructor call, pin dependency versions where your application requires repeatable output, and test fixtures that represent the HTML you actually receive.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Build a reusable extraction function
A small function can make selection, validation and normalization consistent across a codebase:
from bs4 import BeautifulSoup
def extract_text(html: str, selector: str | None = None) -> str:
soup = BeautifulSoup(html, "lxml")
root = soup.select_one(selector) if selector else soup
if root is None:
raise ValueError(f"Selector not found: {selector}")
return root.get_text(" ", strip=True)
html = "<main><h1>Docs</h1><p>Read this first.</p></main>"
print(extract_text(html, "main"))
Keep acquisition separate from parsing. Fetch or read the HTML in one layer, pass the resulting string to this function, and test parsing independently with saved fixtures. That separation makes parser changes and template changes easier to diagnose.
Troubleshoot common extraction failures
Words run together
Symptom: output contains values such as Version2.0. Fix: supply a separator: get_text(" ", strip=True). If you are using HTMLParser, insert separators during your own fragment-joining step.
The result contains menus or cookie text
Cause: you extracted from the entire document. Fix: select the article or main-content element first, or remove known interface nodes with selectors before calling get_text().
Recommended Free Tools
The expected selector returns nothing
Cause: the template changed, the selector is wrong, or the supplied HTML does not contain that element. Check select_one() for None, log the input fixture, and update the selector only after confirming the new structure.
Best Value
Different machines produce different text
Cause: parser choice or version differences changed the parse tree for malformed markup. Fix: name the parser explicitly, pin dependencies, and run the same representative fixtures in each environment.
Whitespace is excessive
Fix: use strip=True with Beautiful Soup, or normalize collected fragments with " ".join(text.split()). Apply normalization after you have made any structural decisions, because early collapsing can erase distinctions you need.
Performance, reliability and cost considerations
Parsing is local once you have the HTML; the main engineering variables are input size, selector complexity and how much cleanup you perform. Avoid repeatedly reparsing the same string when several fields can be selected from one soup object. For large jobs, process documents incrementally where an event-driven parser fits, and release references to completed trees so they can be collected.
Reliability comes from explicit assumptions: validate required selectors, retain failing HTML samples for regression tests, and distinguish an empty page from a page whose template no longer matches. No parser can infer which region is the “main article” from tags alone; that decision requires selectors or a dedicated content-extraction step.
Or skip the browser setup
If your workflow also needs a clean visual capture of a rendered page, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page captures with lazy images loaded, CSS-selector element captures, device and viewport settings, custom CSS and JavaScript, waits, request blocking, cookies, headers, geolocation, signed links, asynchronous jobs, webhooks, bulk capture and a usage API. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for request parameters. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for the free plan to try it.
Practical decision checklist
- Use Beautiful Soup when you want a readable tree API and straightforward selectors.
- Call
get_text(" ", strip=True)for compact prose extraction. - Select the content element before flattening a whole document.
- Use
stripped_stringswhen fragments need individual processing. - Use
HTMLParserwhen the standard library and callback control matter more than convenience. - Name and pin the parser when malformed input or cross-machine consistency matters.
- Test selectors and whitespace rules against real representative fixtures.
Frequently Asked Questions
Does Python’s HTMLParser download a webpage for me?
No. It parses HTML that your program supplies through feed() or related methods; fetching the URL is a separate step.
Can I call get_text() on a single tag instead of the whole document?
Yes. Any Beautiful Soup document or tag can be the target, so selecting a specific element first limits the extracted text to that subtree.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

