Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a conventional, already-rendered HTML table, start with pandas.read_html(). It converts tables into a list of pandas DataFrames, so you can inspect the list and select the intended table. Use Beautiful Soup instead when you need custom table selection, links or attributes, nested elements, or cell-by-cell control. In either case, verify the extracted structure before using the data: HTML irregularities, merged cells, headers, number formats, and JavaScript rendering can all change the result.

Choose the right capture method

Situation Best starting point Why
Ordinary table markup already present in the HTML pandas.read_html() Direct conversion to DataFrames
Several tables on one page read_html() with match or attrs Filters the candidates before selection
Need links, attributes, nested elements, or custom rules Beautiful Soup Manual traversal preserves your extraction logic
Table appears only after JavaScript runs Obtain rendered HTML first, then parse it Neither parser executes page JavaScript

“Read HTML tables into a list of DataFrame objects” is the behavior documented by pandas. A page with one table still produces a list, not a single DataFrame.

Install the Python packages

For the pandas route, install pandas and an HTML parser. lxml is commonly the fast choice; html5lib is more tolerant of malformed markup but slower. Beautiful Soup can use Python’s built-in html.parser, or the separately installed lxml and html5lib backends.

python -m pip install pandas lxml beautifulsoup4 html5lib

You do not need every parser for every script. The explicit parser name in Beautiful Soup makes your choice reproducible. pandas normally tries lxml first and can fall back to Beautiful Soup with html5lib when necessary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Capture a table with pandas

Basic URL example

import pandas as pd

url = "https://example.com/table-page"
tables = pd.read_html(url)

print(f"Found {len(tables)} table(s)")
for index, table in enumerate(tables):
    print(f"nTable {index}: {table.shape}")
    print(table.head())

# Select index 0 only after confirming it is the table you want
df = tables[0]
print(df.to_csv(index=False))

The URL may also be a local file path or a file-like object containing HTML. Do not assume index zero is correct merely because the page currently displays one obvious table; navigation, hidden tables, or layout markup can change the order.

Select one table from many

Use distinctive text with match:

tables = pd.read_html(
    "https://example.com/report",
    match="Quarterly revenue"
)
df = tables[0]

You can filter by valid table attributes such as an ID:

tables = pd.read_html(
    "https://example.com/report",
    attrs={"id": "revenue-table"}
)
df = tables[0]

These filters narrow the returned list; still inspect its length and columns. A text match can select an unintended table if the same phrase appears in more than one.

Control headers, rows, and number formats

df = pd.read_html(
    "report.html",
    attrs={"id": "orders"},
    header=0,
    skiprows=[1],
    thousands=",",
    decimal=".",
    encoding="utf-8"
)[0]

Use header when a particular row contains the column names and skiprows for titles or notes above the data. thousands and decimal help pandas interpret formatted numbers. Converters can apply a function to a specific column:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
def clean_percent(value):
    if pd.isna(value):
        return None
    return float(str(value).rstrip("%")) / 100

df = pd.read_html(
    "report.html",
    converters={"Conversion rate": clean_percent}
)[0]

Link extraction and exact option behavior depend on your pandas version. Check the installed API documentation when you rely on a less common parameter.

Use Beautiful Soup for custom extraction

Find a table and walk its rows

from bs4 import BeautifulSoup

with open("report.html", encoding="utf-8") as file:
    markup = file.read()

soup = BeautifulSoup(markup, "html.parser")
table = soup.find("table", id="orders")
if table is None:
    raise ValueError("The orders table was not found")

rows = []
for tr in table.find_all("tr"):
    cells = tr.find_all(["th", "td"])
    values = [cell.get_text(" ", strip=True) for cell in cells]
    if values:
        rows.append(values)

for row in rows:
    print(row)

This approach lets you select a nested element, retain an anchor’s href, read data attributes, or apply row-specific rules:

records = []
for tr in table.select("tbody tr"):
    name_cell = tr.select_one("td.name")
    link = name_cell.find("a") if name_cell else None
    records.append({
        "name": name_cell.get_text(" ", strip=True) if name_cell else None,
        "url": link.get("href") if link else None,
    })

Parser choice matters

html.parser requires no additional parser package. lxml is generally faster but requires its external dependency. html5lib follows browser-like, lenient parsing but is slower. Invalid markup can produce different trees with different parsers, so run your checks against the actual input rather than assuming equivalent results.

Handle JavaScript-rendered tables

Both pandas and Beautiful Soup parse HTML they receive; they do not run the page’s JavaScript. If the initial response contains an empty table shell and a script later inserts rows, first obtain the rendered DOM with an appropriate browser automation workflow or locate the underlying data request. Save that resulting HTML and pass it to your parser. If an API endpoint supplies JSON, using that endpoint is often more reliable than scraping the visual table, provided you are authorized to access it and follow the site’s terms.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate and clean the extracted data

Extraction is not validation. Add explicit checks before exporting or loading the data:

  • Confirm the selected table’s identity, column names, and row count.
  • Print representative first, middle, and last rows.
  • Check missing values and unexpected duplicate headers.
  • Inspect the effects of rowspan and colspan; merged cells can produce shifted or multi-level columns.
  • Verify decimal marks, thousands separators, currencies, percentages, dates, and encoding.
  • If links or attributes matter, test several rows and confirm their values were retained.

Normalize common text problems

import pandas as pd

# Strip surrounding whitespace from text columns
df = df.map(lambda value: value.strip() if isinstance(value, str) else value)

# Convert a numeric column after removing a currency symbol
# errors="coerce" makes unparseable values visible as NaN
df["Amount"] = pd.to_numeric(
    df["Amount"].astype(str).str.replace("$", "", regex=False).str.replace(",", "", regex=False),
    errors="coerce"
)

print(df.dtypes)
print(df.isna().sum())

Do not silently coerce values you cannot explain. Review the resulting missing values and retain the original text when auditability is important.

Performance, reliability, and repeatability

  • For a single ordinary table, pandas is usually the shortest implementation; Beautiful Soup adds control at the cost of more code.
  • When processing many pages, reuse a session, set network timeouts in your fetching layer, cache downloaded HTML, and log the URL and parser used for each result.
  • Pin compatible package versions in a project environment. Parser upgrades can alter how malformed markup is interpreted.
  • Prefer stable table IDs or distinctive labels over positional selectors.
  • Keep a small fixture of representative HTML, including a merged-cell case and a missing-value case, and run it in tests.

Troubleshooting common failures

“No tables found”

The response may contain no actual <table> element, may be blocked, or may rely on JavaScript. Save and inspect the response, look for an API request, and obtain rendered HTML when necessary.

The wrong table was selected

Print len(tables), each table’s shape, and its columns. Replace positional selection with match or attrs, then verify the result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Headers become unnamed or multi-level

Titles, blank cells, rowspan, and colspan can create missing or hierarchical headers. Set header or flatten columns deliberately after inspecting them; do not discard a level without checking what information it contains.

Numbers remain strings

Specify the correct thousands and decimal characters, remove currency or percent symbols, and use pd.to_numeric with an explicit error policy.

Beautiful Soup returns an unexpected tree

Malformed HTML is interpreted differently by parser backends. Try the parser that matches your deployment, compare the selected rows, and add a fixture test for the page structure.

Encoding or access errors

Read local files with the intended encoding and inspect the HTTP response encoding before parsing. A timeout, consent wall, authentication requirement, or bot check must be handled in the fetching step; changing table selectors will not fix an inaccessible page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If you need a rendered screenshot of a table or page rather than structured cell data, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF, while options include full-page capture with lazy images, CSS-selector element capture, custom CSS and JavaScript, waits, cookies and headers, device presets, dark mode, resizing, blocking rules, caching, signed links, asynchronous jobs, and bulk capture.

Its clean-shot workflow accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. The MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all parameters. The same request in Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

The Free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

FAQ

Does read_html return one DataFrame?

No. It returns a list, even when only one table is present. Select a member after checking which table it represents.

Can Beautiful Soup convert a table directly to a DataFrame?

It extracts the cells; you then construct a DataFrame from the rows, deciding how to handle headers, uneven lengths, and merged cells.

Which parser should production code use?

Choose deliberately based on dependency, speed, and malformed-markup tolerance, then test the selected parser against real pages and saved fixtures.

Why is a screenshot not a substitute for table extraction?

A screenshot preserves visual appearance, not structured values. Use pandas or Beautiful Soup for data; use a screenshot service when you need visual evidence or rendered output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I parse a local HTML string instead of a URL?

Yes. Pass a file path or file-like object to pandas, or give the string to Beautiful Soup. Keep the encoding explicit when reading files.

How do I preserve hyperlinks in a pandas result?

Use Beautiful Soup when link URLs or other cell attributes are part of the required output; it gives direct access to each anchor and attribute.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.