Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →For a conventional, already-rendered HTML table, start with pandas.read_html(). It converts tables into a list of pandas DataFrames, so you can inspect the list and select the intended table. Use Beautiful Soup instead when you need custom table selection, links or attributes, nested elements, or cell-by-cell control. In either case, verify the extracted structure before using the data: HTML irregularities, merged cells, headers, number formats, and JavaScript rendering can all change the result.
Choose the right capture method
| Situation | Best starting point | Why |
|---|---|---|
| Ordinary table markup already present in the HTML | pandas.read_html() |
Direct conversion to DataFrames |
| Several tables on one page | read_html() with match or attrs |
Filters the candidates before selection |
| Need links, attributes, nested elements, or custom rules | Beautiful Soup | Manual traversal preserves your extraction logic |
| Table appears only after JavaScript runs | Obtain rendered HTML first, then parse it | Neither parser executes page JavaScript |
“Read HTML tables into a list of DataFrame objects” is the behavior documented by pandas. A page with one table still produces a list, not a single DataFrame.
Install the Python packages
For the pandas route, install pandas and an HTML parser. lxml is commonly the fast choice; html5lib is more tolerant of malformed markup but slower. Beautiful Soup can use Python’s built-in html.parser, or the separately installed lxml and html5lib backends.
python -m pip install pandas lxml beautifulsoup4 html5lib
You do not need every parser for every script. The explicit parser name in Beautiful Soup makes your choice reproducible. pandas normally tries lxml first and can fall back to Beautiful Soup with html5lib when necessary.
#1 Best Overall
Capture a table with pandas
Basic URL example
import pandas as pd
url = "https://example.com/table-page"
tables = pd.read_html(url)
print(f"Found {len(tables)} table(s)")
for index, table in enumerate(tables):
print(f"nTable {index}: {table.shape}")
print(table.head())
# Select index 0 only after confirming it is the table you want
df = tables[0]
print(df.to_csv(index=False))
The URL may also be a local file path or a file-like object containing HTML. Do not assume index zero is correct merely because the page currently displays one obvious table; navigation, hidden tables, or layout markup can change the order.
Select one table from many
Use distinctive text with match:
tables = pd.read_html(
"https://example.com/report",
match="Quarterly revenue"
)
df = tables[0]
You can filter by valid table attributes such as an ID:
tables = pd.read_html(
"https://example.com/report",
attrs={"id": "revenue-table"}
)
df = tables[0]
These filters narrow the returned list; still inspect its length and columns. A text match can select an unintended table if the same phrase appears in more than one.
Control headers, rows, and number formats
df = pd.read_html(
"report.html",
attrs={"id": "orders"},
header=0,
skiprows=[1],
thousands=",",
decimal=".",
encoding="utf-8"
)[0]
Use header when a particular row contains the column names and skiprows for titles or notes above the data. thousands and decimal help pandas interpret formatted numbers. Converters can apply a function to a specific column:
def clean_percent(value):
if pd.isna(value):
return None
return float(str(value).rstrip("%")) / 100
df = pd.read_html(
"report.html",
converters={"Conversion rate": clean_percent}
)[0]
Link extraction and exact option behavior depend on your pandas version. Check the installed API documentation when you rely on a less common parameter.
Rank #2
Use Beautiful Soup for custom extraction
Find a table and walk its rows
from bs4 import BeautifulSoup
with open("report.html", encoding="utf-8") as file:
markup = file.read()
soup = BeautifulSoup(markup, "html.parser")
table = soup.find("table", id="orders")
if table is None:
raise ValueError("The orders table was not found")
rows = []
for tr in table.find_all("tr"):
cells = tr.find_all(["th", "td"])
values = [cell.get_text(" ", strip=True) for cell in cells]
if values:
rows.append(values)
for row in rows:
print(row)
This approach lets you select a nested element, retain an anchor’s href, read data attributes, or apply row-specific rules:
records = []
for tr in table.select("tbody tr"):
name_cell = tr.select_one("td.name")
link = name_cell.find("a") if name_cell else None
records.append({
"name": name_cell.get_text(" ", strip=True) if name_cell else None,
"url": link.get("href") if link else None,
})
Parser choice matters
html.parser requires no additional parser package. lxml is generally faster but requires its external dependency. html5lib follows browser-like, lenient parsing but is slower. Invalid markup can produce different trees with different parsers, so run your checks against the actual input rather than assuming equivalent results.
Handle JavaScript-rendered tables
Both pandas and Beautiful Soup parse HTML they receive; they do not run the page’s JavaScript. If the initial response contains an empty table shell and a script later inserts rows, first obtain the rendered DOM with an appropriate browser automation workflow or locate the underlying data request. Save that resulting HTML and pass it to your parser. If an API endpoint supplies JSON, using that endpoint is often more reliable than scraping the visual table, provided you are authorized to access it and follow the site’s terms.
Recommended Free Tools
Validate and clean the extracted data
Extraction is not validation. Add explicit checks before exporting or loading the data:
- Confirm the selected table’s identity, column names, and row count.
- Print representative first, middle, and last rows.
- Check missing values and unexpected duplicate headers.
- Inspect the effects of
rowspanandcolspan; merged cells can produce shifted or multi-level columns. - Verify decimal marks, thousands separators, currencies, percentages, dates, and encoding.
- If links or attributes matter, test several rows and confirm their values were retained.
Normalize common text problems
import pandas as pd
# Strip surrounding whitespace from text columns
df = df.map(lambda value: value.strip() if isinstance(value, str) else value)
# Convert a numeric column after removing a currency symbol
# errors="coerce" makes unparseable values visible as NaN
df["Amount"] = pd.to_numeric(
df["Amount"].astype(str).str.replace("$", "", regex=False).str.replace(",", "", regex=False),
errors="coerce"
)
print(df.dtypes)
print(df.isna().sum())
Do not silently coerce values you cannot explain. Review the resulting missing values and retain the original text when auditability is important.
Performance, reliability, and repeatability
- For a single ordinary table, pandas is usually the shortest implementation; Beautiful Soup adds control at the cost of more code.
- When processing many pages, reuse a session, set network timeouts in your fetching layer, cache downloaded HTML, and log the URL and parser used for each result.
- Pin compatible package versions in a project environment. Parser upgrades can alter how malformed markup is interpreted.
- Prefer stable table IDs or distinctive labels over positional selectors.
- Keep a small fixture of representative HTML, including a merged-cell case and a missing-value case, and run it in tests.
Troubleshooting common failures
“No tables found”
The response may contain no actual <table> element, may be blocked, or may rely on JavaScript. Save and inspect the response, look for an API request, and obtain rendered HTML when necessary.
The wrong table was selected
Print len(tables), each table’s shape, and its columns. Replace positional selection with match or attrs, then verify the result.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallHeaders become unnamed or multi-level
Titles, blank cells, rowspan, and colspan can create missing or hierarchical headers. Set header or flatten columns deliberately after inspecting them; do not discard a level without checking what information it contains.
Numbers remain strings
Specify the correct thousands and decimal characters, remove currency or percent symbols, and use pd.to_numeric with an explicit error policy.
Beautiful Soup returns an unexpected tree
Malformed HTML is interpreted differently by parser backends. Try the parser that matches your deployment, compare the selected rows, and add a fixture test for the page structure.
Encoding or access errors
Read local files with the intended encoding and inspect the HTTP response encoding before parsing. A timeout, consent wall, authentication requirement, or bot check must be handled in the fetching step; changing table selectors will not fix an inaccessible page.
Or skip the browser setup
If you need a rendered screenshot of a table or page rather than structured cell data, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF, while options include full-page capture with lazy images, CSS-selector element capture, custom CSS and JavaScript, waits, cookies and headers, device presets, dark mode, resizing, blocking rules, caching, signed links, asynchronous jobs, and bulk capture.
Its clean-shot workflow accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. The MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all parameters. The same request in Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
The Free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
FAQ
Does read_html return one DataFrame?
No. It returns a list, even when only one table is present. Select a member after checking which table it represents.
Best Value
Can Beautiful Soup convert a table directly to a DataFrame?
It extracts the cells; you then construct a DataFrame from the rows, deciding how to handle headers, uneven lengths, and merged cells.
Which parser should production code use?
Choose deliberately based on dependency, speed, and malformed-markup tolerance, then test the selected parser against real pages and saved fixtures.
Why is a screenshot not a substitute for table extraction?
A screenshot preserves visual appearance, not structured values. Use pandas or Beautiful Soup for data; use a screenshot service when you need visual evidence or rendered output.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsFrequently Asked Questions
Can I parse a local HTML string instead of a URL?
Yes. Pass a file path or file-like object to pandas, or give the string to Beautiful Soup. Keep the encoding explicit when reading files.
How do I preserve hyperlinks in a pandas result?
Use Beautiful Soup when link URLs or other cell attributes are part of the required output; it gives direct access to each anchor and attribute.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

