Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Python’s pandas.read_html() to parse an HTML table, inspect how it expands rowspan and colspan, then export the result with DataFrame.to_csv(). Because CSV cannot store merged cells, decide whether spanning values should repeat or remain blank, and verify the output against the source table before relying on it.

Why merged cells need a decision before conversion

HTML tables can merge cells across rows with rowspan or across columns with colspan. CSV is a flat grid of fields and has no merged-cell layout, so conversion must map each visual span into ordinary rows and columns.

A parser may repeat the spanning cell’s value in each position it covers. For example, a cell with rowspan="2" containing 1 can become a 1 in both corresponding CSV rows. That can be useful when each row needs its own group label for filtering or joining. In other cases, keeping the value only in its original position and leaving covered fields empty better reflects the source’s visual layout. Choose the rule that preserves the meaning needed by the CSV’s users; there is no universally correct representation.

Convert a table with pandas

Install pandas and its HTML parsing dependencies in your Python environment before running this example. The pandas API says it “attempts to properly handle colspan and rowspan attributes,” while also warning that table-specific cleanup may be necessary. See the pandas read_html API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from io import StringIO
import pandas as pd

html = """<table>
  <tr><th>Region</th><th colspan="2">Sales</th></tr>
  <tr><th></th><th>2025</th><th>2026</th></tr>
  <tr><td>North</td><td>10</td><td>12</td></tr>
</table>"""

tables = pd.read_html(StringIO(html))
df = tables[0]  # Select the intended table after checking the list
print(df)
df.to_csv("table.csv", index=False)

read_html() accepts HTML text, files, or URLs and returns a list of DataFrames, even when the input contains only one table. Don’t assume the first item is the table you need: inspect the returned tables and choose the intended one. The pandas HTML I/O guide describes options such as match= to select tables by text, attrs= to filter by table attributes, header= to choose a header row, and index_col= to designate an index column.

In the example, StringIO wraps an HTML string as a file-like object. If you are parsing a page, pass its URL or HTML content instead. Check the parsed headers and rows first: multi-level headers, blank cells, and merged areas may need table-specific cleanup. Use index=False when the DataFrame index is not part of the table data; otherwise retain it or make it an explicit output column.

Control the CSV output explicitly

If you need to define the output rows and the treatment of merged cells yourself, use Python’s standard-library csv.writer:

import csv

with open("table.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.writer(f)
    writer.writerow(["Region", "Sales 2025", "Sales 2026"])
    writer.writerow(["North", "10", "12"])

Python’s CSV documentation recommends opening a CSV file with newline='' when using the writer. Its default QUOTE_MINIMAL mode quotes fields containing special characters such as delimiters, quotation marks, or line breaks. This is important when table cells contain commas, quotes, or embedded newlines. CSV dialects differ across applications, so specify a dialect or delimiter if the receiving system requires one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Express Schedule Free Employee Scheduling Software [PC/Mac Download]
  • Simple shift planning via an easy drag & drop interface
  • Add time-off, sick leave, break entries and holidays
  • Email schedules directly to your employees

Validate the converted table

Before using the CSV, compare it with the source table around each merged area. Check that the selected table is correct, the headers and row count make sense, and values occupy the intended columns. Also inspect blank fields and confirm that spanning values follow your chosen repeat-or-blank policy.

  • Several tables appear: inspect the list returned by read_html() and select the right DataFrame, or narrow the selection with match= or attrs=.
  • Headers or columns look shifted: inspect the source’s header rows and spans, then adjust the header or index selection and clean up the DataFrame as needed.
  • A span attribute is malformed: test the specific page and parser. A pandas GitHub issue documents a ValueError for colspan='2;' with pandas 2.2.2; that example is version-specific, not evidence that every current installation fails. See the pandas issue report.
  • The table is nested, dynamically rendered, or malformed: inspect the fetched HTML and test the parser’s output. The pandas guide discusses backend-related HTML parsing considerations.
  • CSV cells contain punctuation or line breaks: use a CSV writer rather than assembling comma-separated text manually, then confirm the receiving application reads the quoting correctly.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Alternative: HTML Table Takeout

HTML Table Takeout documents a Python alternative: parse_html(...) returns table objects with rows containing expanded cells, and its example uses .to_csv(). The project says it supports row and column spans, links, and nested tables. These are maintainer-documented capabilities, not independent comparative test results, so try it on the target page and verify the output before adopting it.

Best Value
Spreadsheet Calculator Software Budget Templates Case for iPhone 11
  • The spreadsheet design is for accountants or calculator Lover who love to use a software for their budget or bills or need in business for projects. You love Accounting programs and Funny bookkeeping templates? Then you'll love this too!
  • Addicted To Spreadsheets
  • Two-part protective case made from a premium scratch-resistant polycarbonate shell and shock absorbent TPU liner protects against drops
  • Printed in the USA
  • Easy installation
Rank #4
MobiOffice Lifetime 4-in-1 Productivity Suite for Windows | Lifetime License | Includes Word Processor, Spreadsheet, Presentation, Email + Free PDF Reader
  • Not a Microsoft Product: This is not a Microsoft product and is not available in CD format. MobiOffice is a standalone software suite designed to provide productivity tools tailored to your needs.
  • 4-in-1 Productivity Suite + PDF Reader: Includes intuitive tools for word processing, spreadsheets, presentations, and mail management, plus a built-in PDF reader. Everything you need in one powerful package.
  • Full File Compatibility: Open, edit, and save documents, spreadsheets, presentations, and PDFs. Supports popular formats including DOCX, XLSX, PPTX, CSV, TXT, and PDF for seamless compatibility.
  • Familiar and User-Friendly: Designed with an intuitive interface that feels familiar and easy to navigate, offering both essential and advanced features to support your daily workflow.
  • Lifetime License for One PC: Enjoy a one-time purchase that gives you a lifetime premium license for a Windows PC or laptop. No subscriptions just full access forever.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.