Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To convert an HTML table to JSON, read the intended table’s header cells, pair each data cell with its heading, decide how to handle duplicate or empty headings, and serialize the resulting row objects. This works well for a regular table with one header row. Tables with rowspan, colspan, multi-level headers, nested tables, or dynamically rendered rows need an explicit schema or a converter that understands those structures.

The basic mapping: one row becomes one JSON object

A browser exposes a table through HTMLTableElement. A table may contain a caption, column groups, header, body and footer sections, so selecting the correct element is the first step. The common output shape is an array such as:

[{"Name":"Ada","Score":"98"},{"Name":"Lin","Score":"91"}]

That shape is a design choice, not an automatic property of HTML. Decide whether headings are case-sensitive, how whitespace is normalized, what happens to repeated headings, and whether values remain strings or become numbers, booleans, dates or null.

Browser JavaScript: convert a selected table

This browser-only function expects one ordinary header row. It preserves cell content as strings, which avoids silently turning identifiers such as 0017 into numbers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option
function tableToJson(table, options = {}) {
  const {
    duplicate = "suffix", // "suffix" or "error"
    emptyHeader = "column", // "column" or "error"
    parse = value => value // keep strings by default
  } = options;

  const clean = text => text.replace(/\s+/g, " ").trim();
  const headerCells = Array.from(table.querySelectorAll(":scope > thead > tr:first-child > th"));
  const firstRow = table.querySelector(":scope > tr");
  const cells = headerCells.length ? headerCells : (firstRow ? Array.from(firstRow.cells) : []);
  const headings = [];
  const counts = new Map();

  cells.forEach((cell, index) => {
    let key = clean(cell.textContent || "");
    if (!key) {
      if (emptyHeader === "error") throw new Error(`Empty heading at column ${index + 1}`);
      key = `column_${index + 1}`;
    }
    const seen = counts.get(key) || 0;
    counts.set(key, seen + 1);
    if (seen > 0) {
      if (duplicate === "error") throw new Error(`Duplicate heading: ${key}`);
      key = `${key}_${seen + 1}`;
    }
    headings.push(key);
  });

  const rows = table.tBodies.length
    ? Array.from(table.tBodies).flatMap(tbody => Array.from(tbody.rows))
    : Array.from(table.rows).slice(1);

  return rows.map((row, rowIndex) => {
    const output = {};
    headings.forEach((key, columnIndex) => {
      const cell = row.cells[columnIndex];
      const raw = clean(cell ? cell.textContent || "" : "");
      output[key] = parse(raw, { row: rowIndex + 1, column: columnIndex + 1, key });
    });
    return output;
  });
}

const table = document.querySelector("#sales");
if (!table) throw new Error("#sales was not found");
const data = tableToJson(table, {
  parse(value) {
    return value !== "" && /^-?\d+(\.\d+)?$/.test(value) ? Number(value) : value;
  }
});
console.log(JSON.stringify(data, null, 2));

Use thead and tbody in your markup whenever possible. The selector above intentionally targets one table; document.querySelector("table") can accidentally read a navigation, pricing or nested table instead of the data you need.

Preserve or parse values deliberately

Cell text is not automatically a trustworthy JSON type. A currency value such as $1,234.50, a localized date, an empty cell and a value containing a comma all require a policy. Keep strings when the value is an identifier or display text. Parse only with rules appropriate to the source, and record malformed values rather than quietly producing incorrect data.

Handling spans and complex headers

The simple function assumes one cell per column in every data row. That assumption fails when markup uses:

  • colspan: one heading or data cell covers several visual columns.
  • rowspan: a cell occupies rows below it, so later rows have fewer physical cells.
  • Multi-level headings: a child heading may only be unique when combined with its parent, such as Q1 / Revenue.
  • Blank or repeated labels: “Total” may occur in several groups, and an empty heading has no stable property name.
  • Nested markup: icons, links and buttons can add text that you do not want in the value.

For these tables, first build a rectangular grid by expanding row and column spans. Then derive a header path for each column (for example, Region > Q1 > Revenue) or supply a schema yourself. A schema is safer when the table is presentation-oriented or changes frequently. Also inspect whether a footer row is data, a subtotal, or a note before including it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Node.js: parsing saved HTML

Node.js does not provide a browser DOM by default. Use a DOM implementation such as the one already approved in your project, parse the saved HTML, and pass the selected table to the same mapping logic. The following example uses jsdom; install it with npm install jsdom.

import { readFile } from "node:fs/promises";
import { JSDOM } from "jsdom";

const html = await readFile("page.html", "utf8");
const { document } = new JSDOM(html).window;
const table = document.querySelector("#sales");
if (!table) throw new Error("#sales was not found");

const headers = [...table.querySelectorAll("thead th")].map((cell, i) => {
  const key = cell.textContent.replace(/\s+/g, " ").trim();
  return key || `column_${i + 1}`;
});
const rows = [...table.querySelectorAll("tbody tr")].map(row => {
  const cells = [...row.cells];
  return Object.fromEntries(headers.map((key, i) => [key, cells[i]?.textContent.replace(/\s+/g, " ").trim() ?? ""]));
});
console.log(JSON.stringify(rows, null, 2));

If the page creates rows only after JavaScript runs, downloading its original HTML will not contain those rows. Use a browser automation step to wait for the table, or export the data from the page’s network/API response instead.

Python: convert HTML with Beautiful Soup

Install the parser with pip install beautifulsoup4. This example selects a table by ID and keeps all values as strings.

from bs4 import BeautifulSoup
import json

with open("page.html", encoding="utf-8") as f:
    soup = BeautifulSoup(f, "html.parser")

table = soup.select_one("#sales")
if table is None:
    raise ValueError("#sales was not found")

header_row = table.select_one("thead tr") or table.select_one("tr")
if header_row is None:
    raise ValueError("The table has no header row")

headers = []
for index, cell in enumerate(header_row.select("th, td"), start=1):
    key = " ".join(cell.get_text(" ", strip=True).split()) or f"column_{index}"
    if key in headers:
        key = f"{key}_{headers.count(key) + 1}"
    headers.append(key)

rows = []
for row in table.select("tbody tr"):
    values = [" ".join(cell.get_text(" ", strip=True).split()) for cell in row.select("th, td")]
    rows.append({key: values[i] if i < len(values) else "" for i, key in enumerate(headers)})

print(json.dumps(rows, ensure_ascii=False, indent=2))

Choosing a conversion approach

Approach Best input Output and trade-off
Custom DOM mapping A selected, regular table already in a browser Small and controllable array of objects; you own every edge case.
Python parser Saved HTML or server-side processing Convenient text extraction; it will not execute page JavaScript.
Node.js DOM parser HTML files or fetched markup Reuse JavaScript logic; dynamic content still requires a browser or API.
tabletojson HTML markup or a URL Its documentation covers duplicate headings, spans, complex headers, HTML in cells, ignored columns and row limits. Verify the current package version and behavior against your table.
W3C tabular-data conversion An annotated tabular-data model Standards-oriented metadata and parsing, rather than an arbitrary DOM-to-object recipe.

The W3C document “Generating JSON from Tabular Data on the Web” defines minimal and standard conversion modes and states: “A conformant JSON conversion application MUST produce output conforming to this algorithm according to the chosen mode of conversion: standard or minimal.” That requirement applies to its annotated tabular-data model; it does not make every hand-written DOM mapping a standards conversion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validation and failure modes

Wrong table selected

Use a stable ID, class, caption text or a surrounding container. If several tables are expected, iterate over them and label each result instead of assuming the first match.

Rows have different lengths

Compare the number of logical columns with each row. Missing cells should become an explicit empty value or an error according to your contract; never shift later cells into the wrong property.

Duplicate keys overwrite data

Suffix duplicates, construct hierarchical keys, or reject the table. JavaScript object assignment with the same key otherwise discards the earlier value.

Numbers and dates are wrong

Keep the original string alongside a parsed field, or use a parser that understands the table’s locale. Reject ambiguous dates and thousands separators rather than guessing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Rendered tables are empty

Wait for a selector and for the expected row count, or call the endpoint that supplies the data. A static HTTP fetch cannot see rows generated after load.

Accessibility and semantics are lost

Visual spans do not always provide unambiguous header associations, especially for assistive technology. Prefer explicit scope or headers attributes in source markup and preserve header paths in your JSON when meaning depends on them.

Performance, reliability and cost considerations

  • Limit extraction to the required table and columns; avoid serializing hidden navigation tables.
  • For large tables, process rows incrementally or in batches rather than retaining multiple intermediate HTML copies.
  • Cache the source HTML only when its freshness requirements permit it, and validate a representative sample after every source-layout change.
  • Record the source URL, capture time, heading list and parse errors with each export so downstream users can audit changes.
  • Do not claim speed, accuracy or privacy advantages for a library without testing the exact tables and runtime you will deploy.

Or skip the browser setup

If your immediate need is a clean visual capture of the page or table rather than structured cell data, ScreenshotNeo provides a one-call screenshot API. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options such as full-page capture, CSS selectors, waiting, custom headers and cookies, blocking requests, PDF output and bulk jobs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

FAQ

Can JSON contain duplicate property names?

Objects should have unique property names. Rename duplicates, use nested objects, or represent the row as an array with a separate column definition.

Should an empty table cell become null?

Only if your data contract says an empty cell means missing data. Otherwise preserve it as an empty string and distinguish it from an explicit textual value such as “N/A.”

Can I convert every table on a page?

Yes, select all intended tables and return an object keyed by a stable identifier, caption or index. Validate that decorative and nested tables are excluded.

Is a screenshot a replacement for table-to-JSON conversion?

No. A screenshot preserves appearance, while JSON preserves values and structure. Use a DOM or data-source extraction method when another program must consume the cells.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can JSON contain duplicate property names?

Objects should have unique property names. Rename duplicates, use nested objects, or represent the row as an array with a separate column definition.

Should an empty table cell become null?

Only if your data contract says an empty cell means missing data. Otherwise preserve it as an empty string and distinguish it from an explicit textual value such as “N/A.”

Can I convert every table on a page?

Yes, select all intended tables and return an object keyed by a stable identifier, caption or index. Validate that decorative and nested tables are excluded.

Is a screenshot a replacement for table-to-JSON conversion?

No. A screenshot preserves appearance, while JSON preserves values and structure. Use a DOM or data-source extraction method when another program must consume the cells.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

For a regular table, a carefully selected DOM plus explicit heading, typing and error policies is enough. Spans, multi-level headers and dynamic rendering require a grid-aware converter or a schema you control.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.