Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use pandas.read_html() to turn a page’s HTML tables into a list of DataFrames, then loop over that list with a standard Python for loop. It returns a list even when the page contains just one table, so the loop pattern is the same in either case.

Loop through every HTML table with pandas

For a page with conventional HTML table markup, this is the simplest approach:

import pandas as pd

source = "https://example.com/page"
tables = pd.read_html(source)

for number, df in enumerate(tables, start=1):
    print(f"Table {number}: {df.shape}")
    print(df.head())

pd.read_html() accepts a URL, path, or file-like object. It searches for table elements and returns a list of DataFrames, even if it finds only one table. The pandas API documentation describes the function as reading HTML tables into a list of DataFrame objects.

Each item in tables is a DataFrame, so you can clean, validate, or transform it inside the loop. enumerate(..., start=1) gives the tables human-friendly numbers; omit it if you only need the DataFrames:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
for df in tables:
    print(df.columns)

Select and shape tables while parsing

If a page has many tables, pass filters and structure options to read_html() instead of parsing everything and filtering afterward.

tables = pd.read_html(
    source,
    match="Revenue",
    attrs={"id": "annual-results"},
    header=0,
    index_col=0,
)

for df in tables:
    print(df.head())
  • match selects tables containing text that matches the supplied expression.
  • attrs filters by HTML attributes, such as a stable table id or class.
  • header and index_col tell pandas how to interpret header and index columns.
  • skiprows skips preamble rows that should not be treated as table data.
  • na_values and converters let you control missing-value and column conversion behavior.

These options are documented in the pandas HTML I/O guide. A filter may return more than one DataFrame if several tables match, so keep the loop when you need to process every result.

Use Beautiful Soup when you need to inspect table tags

When tables are visually similar, nested, or surrounded by ambiguous markup, inspect the HTML first. Beautiful Soup can locate the table elements; each tag can then be passed to pandas for conversion:

from bs4 import BeautifulSoup
import pandas as pd

soup = BeautifulSoup(html, "html.parser")

for table_number, table_tag in enumerate(soup.find_all("table"), start=1):
    frames = pd.read_html(str(table_tag))
    for df in frames:
        print(f"Table {table_number}")
        print(df)

This route gives you explicit control over which table tags to inspect before converting them. Beautiful Soup is a Python library for pulling data from HTML and XML, as described in its documentation. For ordinary pages with clear table markup, calling pd.read_html(source) directly is usually less work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate the parsed data before using it

A successful parse only means pandas produced a DataFrame; it does not prove the columns and values match your intended meaning. The pandas guide cautions that HTML varies and cleanup may be needed. Check headers, missing values, types, row counts, and duplicate headers before combining tables.

for number, df in enumerate(pd.read_html(source), start=1):
    df.columns = [str(column).strip() for column in df.columns]

    required = {"Name", "Value"}
    missing = required.difference(df.columns)
    if missing:
        print(f"Skipping table {number}; missing {missing}")
        continue

    df["Value"] = pd.to_numeric(df["Value"], errors="coerce")
    print(df.head())

Numeric-looking identifiers can lose leading zeros when interpreted as numbers. Preserve them as strings with a converter, using the actual column name from your table:

tables = pd.read_html(source, converters={"code": str})

Also inspect how dates, links, and blank cells were interpreted; adjust parsing or clean the resulting DataFrame when the defaults do not fit your data.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose a parser and diagnose failures

The pandas guide describes parser paths involving lxml, Beautiful Soup, and html5lib. lxml is fast but offers weaker guarantees for invalid markup; html5lib is more lenient and can repair malformed HTML, at a potential speed cost. pandas may fall back between parser options depending on what is installed and which parser succeeds. See the parser discussion in the pandas guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If pandas does not find the table, inspect the response with Beautiful Soup and check whether the data is present in the initial HTML. Some pages render tables with JavaScript after page load; static HTML parsing will not necessarily see content added later. The cited pandas and Beautiful Soup documentation does not establish a universal extraction method for JavaScript-rendered tables, so first determine whether the response actually contains the table markup.

For repeatable data work, record the source URL, table index, parser choice, and any filtering arguments alongside the output. That makes it easier to identify which table produced a result if a page changes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.