Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single best Python HTML parser for every job. Choose Beautiful Soup for readable extraction code, lxml for direct tree work when response time matters, html5lib when browser-aligned HTML5 parsing is the priority, Python’s built-in html.parser when you want to avoid another parser package, and selectolax when CSS selectors and throughput are important enough to benchmark. The key distinction is that Beautiful Soup is a convenient interface over parsing backends; it is not one fixed parsing engine.

How to choose a Python HTML parser

Start with the behavior your program needs, not an assumed speed ranking. HTML found on the web is often malformed or incomplete, and parsers can build different trees from the same input. If exact tree behavior matters, test your real pages with the parser you intend to deploy and pin the selected backend or dependency versions.

  • Readable, straightforward extraction: Beautiful Soup.
  • Direct HTML or XML tree work, especially when response time matters: lxml.
  • Parsing according to WHATWG HTML rules: html5lib.
  • No additional parser package: Python’s html.parser.
  • CSS-selector extraction and a throughput candidate: selectolax, particularly its Lexbor backend.

These libraries parse markup; they do not render a webpage or run its JavaScript. If the data you need appears only after browser-side scripts execute, obtain the rendered HTML with a browser or a screenshot/rendering service first, then parse that output if needed.

At a glance: five Python HTML parsers

Library Best fit Main trade-off
Beautiful Soup Approachable, Python-facing extraction API Results and speed depend on the backend you select.
lxml Direct HTML/XML tree work; a strong choice when speed matters Confirm that its handling of malformed input suits your use case.
html5lib HTML5 parsing behavior designed to conform to the WHATWG specification Standards-oriented parsing can trade speed for fidelity; no universal slowdown figure is established here.
html.parser Standard-library parsing without an extra parser package Its resulting tree can differ from other parsers, especially for malformed markup.
selectolax CSS selectors and a candidate for throughput-sensitive extraction Its published benchmark is project-produced and workload-specific, not a universal comparison.

1. Beautiful Soup: the approachable extraction interface

Beautiful Soup is often the simplest place to start when the goal is to find elements and extract text or attributes without writing low-level tree-handling code. Its interface can use different parsers, including Python’s built-in parser, lxml, and html5lib. Consequently, saying “I use Beautiful Soup” does not fully specify how the markup is parsed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Specify the backend for repeatable results

Beautiful Soup may select the best parser installed in an environment if you do not specify one. Two machines with different installed dependencies can therefore handle the same malformed document differently. When consistent behavior matters, name the backend explicitly, for example BeautifulSoup(markup, "lxml"), and ensure that backend is installed wherever the program runs.

from bs4 import BeautifulSoup

html = "<article><h1>Parser guide</h1><a href='/docs'>Docs</a></article>"
soup = BeautifulSoup(html, "lxml")

print(soup.select_one("h1").get_text(strip=True))
print(soup.select_one("a")["href"])

For debugging, Beautiful Soup’s diagnose() helper can report how available parsers handle a piece of input. That makes it useful when an unexpected element is missing or the tree differs between environments.

2. lxml: direct HTML and XML tree processing

Use lxml directly when you want its tree API rather than Beautiful Soup’s more general extraction interface, particularly if response time is a critical constraint or you also need XML facilities. Beautiful Soup’s documentation advises that it cannot be as fast as the parsers underneath it, recommends lxml as a faster backend than html.parser or html5lib, and suggests working directly with lxml when response time is critical.

from lxml import html

markup = "<main><h1>Parser guide</h1><a href='/docs'>Docs</a></main>"
root = html.fromstring(markup)

print(root.xpath("string(.//h1)"))
print(root.xpath(".//a/@href"))

Do not treat “fast” as proof that lxml will produce the tree your application expects. If input is malformed, inspect the parsed result and compare it with the semantics your data extraction depends on.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

3. html5lib: choose HTML5 parsing behavior

html5lib describes itself as designed to conform to the WHATWG HTML specification as implemented by major web browsers. Choose it when that standards-oriented parsing behavior is more important than throughput. The available evidence does not establish one general numeric speed penalty, so benchmark your own workload if performance is material.

import html5lib

markup = "<main><h1>Parser guide</h1></main>"
document = html5lib.parse(markup)

print(document.tag)

html5lib supports different tree builders, including ElementTree, minidom, and lxml.etree. Select the tree representation that fits the rest of your code rather than assuming every consumer must use the default.

4. Python’s built-in html.parser: no extra parser dependency

The standard library’s html.parser is a practical starting point when installing an additional parser package is undesirable. It is also available as a Beautiful Soup backend. Keep in mind that its treatment of malformed markup can produce a different tree from lxml or html5lib; built-in does not mean behaviorally interchangeable.

from html.parser import HTMLParser

class Headings(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_h1 = False

    def handle_starttag(self, tag, attrs):
        if tag == "h1":
            self.in_h1 = True

    def handle_endtag(self, tag):
        if tag == "h1":
            self.in_h1 = False

    def handle_data(self, data):
        if self.in_h1:
            print(data)

parser = Headings()
parser.feed("<h1>Parser guide</h1>")
parser.close()

This event-driven style makes you responsible for the extraction state you need. For richer document navigation or CSS-selector queries, a higher-level interface may be more convenient.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

5. selectolax: CSS selectors and a throughput candidate

selectolax offers HTML parsing and CSS-selector workflows. Its project recommends the Lexbor backend, and its examples use LexborHTMLParser with methods such as css_first. Consider it when selector-based extraction and speed matter, but validate it on your own documents and extraction logic.

from selectolax.lexbor import LexborHTMLParser

markup = "<main><h1>Parser guide</h1><a href='/docs'>Docs</a></main>"
tree = LexborHTMLParser(markup)

heading = tree.css_first("h1")
link = tree.css_first("a")
print(heading.text() if heading else None)
print(link.attributes.get("href") if link else None)

Why malformed HTML changes the answer

Consider the fragment <a></p>. Beautiful Soup’s documentation shows different outcomes depending on the backend: lxml drops the unmatched closing paragraph and supplies html and body elements; html5lib constructs a paragraph and supplies html, head, and body; html.parser leaves a simpler tree. There is no meaningful universal winner for invalid input until you decide which parsing rules and resulting structure your application requires.

  1. Save representative raw pages, including the malformed cases that have caused trouble.
  2. Parse them with the intended backend and inspect the resulting tree, not just the text your first selector happens to return.
  3. Check edge cases such as missing end tags, nested elements, attributes, and unexpected wrappers.
  4. Make the backend explicit in code and keep the dependency environment consistent across development and production.

What the available performance evidence does—and does not—show

The selectolax project describes a benchmark that extracts titles, links, scripts, and a meta tag from the main pages of 754 domains. It reports the following times for that specific task:

Approach in the project benchmark Reported time
Beautiful Soup with html.parser 61.02 seconds
lxml through Beautiful Soup 9.09 seconds
html5_parser 16.10 seconds
selectolax (Modest) 2.94 seconds
selectolax (Lexbor) 2.39 seconds

These are project-reported results for its selected pages and extraction task. They are not an independently comparable benchmark of every candidate, and they do not establish how your workload will rank. The project material does not state a publication year for these figures. Use them as a reason to benchmark promising candidates, not as a promise of a particular speedup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Performance, reliability, and deployment checklist

  • Measure the whole extraction task: include parsing, selector or XPath work, text normalization, and the mix of pages your application actually sees.
  • Separate parsing from fetching: time spent downloading a page is not parser throughput. A parser cannot fix a slow or failed network request.
  • Pin dependencies where output matters: especially for Beautiful Soup, explicitly choose a backend and deploy the same relevant packages across environments.
  • Test both valid and messy input: production pages may contain unclosed elements, inconsistent nesting, or unexpected markup.
  • Keep expectations realistic: none of these parsers executes page JavaScript or produces browser-rendered content on its own.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common parser problems

My script returns different elements on another machine

If Beautiful Soup selected its backend automatically, the installed parser may differ across machines. Pass the backend explicitly, install it consistently, and inspect the generated trees for the affected input.

A selector finds no element that appears in the source

Check whether the selected parser rearranged malformed markup, whether the content is actually present in the raw HTML you supplied, and whether your selector matches the resulting tree. If the content appears only after JavaScript runs, a parser alone will not obtain it.

The result contains unexpected wrappers or missing tags

That can be a parser’s recovery behavior for invalid markup, rather than a bug in your selector. Compare parser output on a minimal reproduction. Use html5lib if WHATWG-style handling is the requirement; otherwise choose and pin the backend whose tree fits your application.

Parsing is taking too long

Measure a representative batch and identify whether the cost is fetching, parsing, or extraction. Try lxml directly if response time is critical, or benchmark selectolax with Lexbor for your selector workload. A project benchmark is a starting point for evaluation, not a substitute for measuring your own inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

I need text produced by a client-side application

First obtain the rendered page or the data source used by the client-side code. Then pass the resulting HTML or relevant response body to a parser. Do not expect switching parsers to execute scripts.

Need a rendered page before parsing? ScreenshotNeo is an alternative

ScreenshotNeo is a website screenshot API and MCP server, not a Python HTML parser. It is the alternative to try first when the task is capturing a webpage rather than extracting a parsed HTML tree. A screenshot can document what a page renders, but it does not replace a parser when your program needs structured text or attributes.

For example, a single GET request can capture a URL as an image. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up free for 1,000 screenshots a month—no card required.

Frequently Asked Questions

Does Beautiful Soup parse HTML by itself?

Beautiful Soup provides the Python-facing interface and uses a selected backend to parse markup; specifying the backend makes that choice explicit.

Which parser should I use if I need browser-like HTML5 recovery?

Use html5lib when its stated WHATWG HTML parsing target is the behavior you need.

Can a Python HTML parser read content that JavaScript adds to a page?

Not by itself. It parses the markup supplied to it and does not execute JavaScript.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.