Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Nokogiri to capture an HTML table in Ruby: parse the document, select the intended <table>, iterate through its <tr> elements, and read each row’s <th> and <td> cells. That produces the cells present in the DOM. If the table uses rowspan or colspan, add a grid-normalization step when you must preserve visual column positions. This guide shows both approaches, safe parsing defaults, HTML4/HTML5 choices, encoding checks, CSV output, troubleshooting, and a way to obtain a rendered page capture without managing a browser.

Install Nokogiri and choose a parser

Nokogiri is the direct Ruby option for parsing HTML and searching it with CSS or XPath. Add it to your project:

gem install nokogiri

For a Bundler project, put gem "nokogiri" in your Gemfile and run bundle install. Record the Ruby runtime, Nokogiri version, and parser choice in applications where reproducibility matters: parser behavior can differ between CRuby and JRuby.

HTML4-compatible parsing

Nokogiri::HTML is the broadly compatible choice and works for ordinary HTML documents:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
require "nokogiri"

doc = Nokogiri::HTML(File.read("page.html"))

HTML5 parsing

Nokogiri::HTML5 is documented as available since Nokogiri 1.12.0. Its HTML5 functionality is unavailable on JRuby, so confirm your runtime before using it:

require "nokogiri"

doc = Nokogiri::HTML5(File.read("page.html"))

The HTML5 API also exposes options such as parse-error reporting, maximum tree depth, maximum attributes per element, and an encoding parameter. Use those options only after checking the API supported by your installed version.

Capture a table as rows and cell values

Start by scoping the search to one known table. A page can contain several tables, nested markup, repeated headers, or empty cells, so inspect a representative result before building downstream logic.

require "nokogiri"

html = File.read("page.html")
doc = Nokogiri::HTML(html)

table = doc.at_css("table#results")
raise "table not found" unless table

rows = table.css("tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

p rows

For a table such as:

<table id="results">
  <tr><th>Name</th><th>Status</th></tr>
  <tr><td>Build</td><td>Passed</td></tr>
</table>

rows is:

[["Name", "Status"], ["Build", "Passed"]]

CSS selectors versus XPath

CSS is concise when the table has an ID or class, for example table#results or table.data. XPath is useful when selection depends on structure or text:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
table = doc.at_xpath("//table[@id='results']")
rows = table.xpath(".//tr").map do |row|
  row.xpath("./th | ./td").map { |cell| cell.text.strip }
end

The leading dot in .//tr and ./th | ./td keeps the search inside the selected table rather than accidentally collecting rows from another table.

When a simple cell list is not enough

The basic pattern returns the cells that occur in each row. It does not expand rowspan or colspan into a rectangular grid. A visual table can therefore have rows of different lengths even though a reader sees aligned columns.

Detect irregular rows first

widths = rows.map(&:length)
warn "irregular table: #{widths.inspect}" unless widths.uniq.length == 1

If all rows have the same width and the document has no spans, the array is usually sufficient. For a rectangular result that preserves spans, place each cell into the next free column and reserve the covered positions.

Normalize rowspan and colspan

def table_grid(table)
  grid = []

  table.css("tr").each_with_index do |row_node, row_index|
    grid[row_index] ||= []
    column = 0

    row_node.css("th, td").each do |cell|
      column += 1 while grid[row_index][column]
      rowspan = (cell["rowspan"] || "1").to_i
      colspan = (cell["colspan"] || "1").to_i
      value = cell.text.strip

      rowspan.times do |r_offset|
        target_row = row_index + r_offset
        grid[target_row] ||= []
        colspan.times do |c_offset|
          target_column = column + c_offset
          grid[target_row][target_column] ||= value
        end
      end

      column += colspan
    end
  end

  width = grid.map(&:length).max || 0
  grid.map { |row| row.fill(nil, row.length...width) }
end

grid = table_grid(table)
p grid

This implementation repeats a spanning cell’s value in every covered position. If your data model needs span metadata instead, retain the original cell, its row and column, and the two span counts rather than repeating text. Also inspect nested tables: a selector such as table.css("tr") can include rows from a nested table unless you deliberately restrict the traversal.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract headers and preserve meaningful text

Header cells may appear in a separate header row or be repeated in a body section. Capture them explicitly when you need named records:

header = table.css("tr").first&.css("th, td")&.map { |cell| cell.text.strip } || []
data = table.css("tr")[1..]&.map do |row|
  row.css("td").map { |cell| cell.text.strip }
end || []

records = data.map { |values| header.zip(values).to_h }

Do not assume the first row is a header if the page uses a caption, a multi-row header, or body rows containing <th>. Inspect the markup and choose the rows intentionally. cell.text returns UTF-8 text in Nokogiri. Verify non-ASCII names, currency symbols, and accents in your output, especially when the source is read from an IO object with a declared encoding.

Export captured rows to CSV

Extraction and CSV serialization are separate operations. Use Ruby’s CSV library rather than joining values with commas; CSV handles quoting, embedded commas, quotes, and line breaks.

require "csv"

CSV.open("results.csv", "wb", write_headers: false) do |csv|
  rows.each { |row| csv << row }
end

With a header row and records, a table-oriented representation is convenient:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
require "csv"

table_data = CSV::Table.new(rows.map { |row| CSV::Row.new(rows.first, row) })
puts table_data.by_col["Status"]

Construct CSV::Table only after deciding which row supplies the headers and ensuring every data row is aligned. If spans were normalized, use the rectangular grid rather than the original variable-length arrays.

Safe parsing for downloaded or user-supplied HTML

Nokogiri treats input as untrusted by default. Its parser does not load external DTDs or access the network for external resources. Keep those protections enabled when processing scraped or uploaded documents.

  • Do not disable network protections for convenience.
  • Do not enable external entity or DTD behavior for untrusted HTML.
  • Remember that parser safety does not grant permission to fetch a site or bypass its access controls.
  • Limit input size and validate the selected table before storing extracted data.

Complete reusable Ruby script

This script accepts a file and a CSS selector, fails clearly when the table is absent, and writes a CSV file.

#!/usr/bin/env ruby
require "nokogiri"
require "csv"

path = ARGV.fetch(0) { abort "usage: ruby capture_table.rb page.html [selector] [output.csv]" }
selector = ARGV.fetch(1, "table")
out_path = ARGV.fetch(2, "table.csv")

html = File.binread(path)
doc = Nokogiri::HTML(html)
table = doc.at_css(selector) or abort "table not found for selector #{selector.inspect}"

rows = table.css("tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

abort "no rows found" if rows.empty?

CSV.open(out_path, "wb") do |csv|
  rows.each { |row| csv << row }
end

puts "wrote #{rows.length} rows to #{out_path}"

Run it with ruby capture_table.rb page.html 'table#results' results.csv. Change the selector to match the actual page and inspect the generated CSV before automating large batches.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and runtime choices

  • Scope early: selecting one table before iterating reduces accidental work and prevents data from unrelated tables entering the result.
  • Read once: parse the document once, then reuse the selected node for headers, rows, and validation.
  • Prefer deterministic inputs: save the HTML and record Ruby, Nokogiri, parser, and encoding details when a result must be reproducible.
  • Expect malformed markup: HTML parsers repair broken structure; test selectors against real samples rather than relying on source indentation.
  • Validate shape: check row counts, expected headers, and span-induced irregularities before importing data.

Common failures and fixes

“table not found”

The selector may target the wrong ID or class, the table may be inserted by JavaScript after the original HTML loads, or the file may not contain the expected page. Print doc.css("table").length, inspect available attributes, and confirm that you are parsing the HTML response rather than a login or error page.

Rows have different lengths

Check for rowspan, colspan, nested tables, and repeated header rows. Use the grid normalizer above when visual column alignment matters; otherwise preserve the variable-length cell arrays and document that choice.

HTML5 parser fails on JRuby

Nokogiri::HTML5 is not available on JRuby. Use the supported parser API for that runtime, or run the HTML5 parser on CRuby after confirming behavior against your project’s Nokogiri version.

Accented text is corrupted

Check the source encoding and the way the IO is passed to Nokogiri. The parser’s text values are UTF-8, but an incorrect input-encoding assumption can still produce bad characters. Test names and symbols, not just ASCII fixtures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CSV columns shift

Do not concatenate values with commas. Serialize through CSV, and ensure every row is aligned to the same header set after handling spans and missing cells.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is to obtain a clean rendered image or PDF of the page containing the table before processing it, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page and element captures, dark mode, device presets or custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, easing migration.

See the ScreenshotNeo documentation for request options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots each month without a card; paid plans start at $5 for 3,000 shots, with every feature on every plan. Sign up for the free ScreenshotNeo plan.

Frequently Asked Questions

Can Nokogiri execute JavaScript to create a table?

No. Nokogiri parses the HTML it receives; it does not provide a browser runtime. Obtain rendered HTML with an appropriate browser workflow or capture service, then parse the resulting document.

Should I use CSS or XPath for table extraction?

Use whichever expresses the target most clearly. CSS is concise for IDs and classes; XPath is useful for structural or text-based conditions. Both can scope rows and cells to one selected table.

How do I keep formulas, links, or images inside cells?

The examples extract visible text with cell.text. Select attributes such as cell["href"] or inspect child nodes separately when the cell’s markup, rather than its text, is the data you need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.