Use Nokogiri to capture an HTML table in Ruby: parse the document, select the intended <table>, iterate through its <tr> elements, and read each row’s <th> and <td> cells. That produces the cells present in the DOM. If the table uses rowspan or colspan, add a grid-normalization step when you must preserve visual column positions. This guide shows both approaches, safe parsing defaults, HTML4/HTML5 choices, encoding checks, CSV output, troubleshooting, and a way to obtain a rendered page capture without managing a browser.
Install Nokogiri and choose a parser
Nokogiri is the direct Ruby option for parsing HTML and searching it with CSS or XPath. Add it to your project:
gem install nokogiri
For a Bundler project, put gem "nokogiri" in your Gemfile and run bundle install. Record the Ruby runtime, Nokogiri version, and parser choice in applications where reproducibility matters: parser behavior can differ between CRuby and JRuby.
HTML4-compatible parsing
Nokogiri::HTML is the broadly compatible choice and works for ordinary HTML documents:
#1 Best Overall
require "nokogiri"
doc = Nokogiri::HTML(File.read("page.html"))
HTML5 parsing
Nokogiri::HTML5 is documented as available since Nokogiri 1.12.0. Its HTML5 functionality is unavailable on JRuby, so confirm your runtime before using it:
require "nokogiri"
doc = Nokogiri::HTML5(File.read("page.html"))
The HTML5 API also exposes options such as parse-error reporting, maximum tree depth, maximum attributes per element, and an encoding parameter. Use those options only after checking the API supported by your installed version.
Capture a table as rows and cell values
Start by scoping the search to one known table. A page can contain several tables, nested markup, repeated headers, or empty cells, so inspect a representative result before building downstream logic.
require "nokogiri"
html = File.read("page.html")
doc = Nokogiri::HTML(html)
table = doc.at_css("table#results")
raise "table not found" unless table
rows = table.css("tr").map do |row|
row.css("th, td").map { |cell| cell.text.strip }
end
p rows
For a table such as:
<table id="results">
<tr><th>Name</th><th>Status</th></tr>
<tr><td>Build</td><td>Passed</td></tr>
</table>
rows is:
[["Name", "Status"], ["Build", "Passed"]]
CSS selectors versus XPath
CSS is concise when the table has an ID or class, for example table#results or table.data. XPath is useful when selection depends on structure or text:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemstable = doc.at_xpath("//table[@id='results']")
rows = table.xpath(".//tr").map do |row|
row.xpath("./th | ./td").map { |cell| cell.text.strip }
end
The leading dot in .//tr and ./th | ./td keeps the search inside the selected table rather than accidentally collecting rows from another table.
Rank #2
When a simple cell list is not enough
The basic pattern returns the cells that occur in each row. It does not expand rowspan or colspan into a rectangular grid. A visual table can therefore have rows of different lengths even though a reader sees aligned columns.
Detect irregular rows first
widths = rows.map(&:length)
warn "irregular table: #{widths.inspect}" unless widths.uniq.length == 1
If all rows have the same width and the document has no spans, the array is usually sufficient. For a rectangular result that preserves spans, place each cell into the next free column and reserve the covered positions.
Normalize rowspan and colspan
def table_grid(table)
grid = []
table.css("tr").each_with_index do |row_node, row_index|
grid[row_index] ||= []
column = 0
row_node.css("th, td").each do |cell|
column += 1 while grid[row_index][column]
rowspan = (cell["rowspan"] || "1").to_i
colspan = (cell["colspan"] || "1").to_i
value = cell.text.strip
rowspan.times do |r_offset|
target_row = row_index + r_offset
grid[target_row] ||= []
colspan.times do |c_offset|
target_column = column + c_offset
grid[target_row][target_column] ||= value
end
end
column += colspan
end
end
width = grid.map(&:length).max || 0
grid.map { |row| row.fill(nil, row.length...width) }
end
grid = table_grid(table)
p grid
This implementation repeats a spanning cell’s value in every covered position. If your data model needs span metadata instead, retain the original cell, its row and column, and the two span counts rather than repeating text. Also inspect nested tables: a selector such as table.css("tr") can include rows from a nested table unless you deliberately restrict the traversal.
Free tools Windows power users keep installed
One-click scans. No signup required.
Extract headers and preserve meaningful text
Header cells may appear in a separate header row or be repeated in a body section. Capture them explicitly when you need named records:
header = table.css("tr").first&.css("th, td")&.map { |cell| cell.text.strip } || []
data = table.css("tr")[1..]&.map do |row|
row.css("td").map { |cell| cell.text.strip }
end || []
records = data.map { |values| header.zip(values).to_h }
Do not assume the first row is a header if the page uses a caption, a multi-row header, or body rows containing <th>. Inspect the markup and choose the rows intentionally. cell.text returns UTF-8 text in Nokogiri. Verify non-ASCII names, currency symbols, and accents in your output, especially when the source is read from an IO object with a declared encoding.
Rank #3
Export captured rows to CSV
Extraction and CSV serialization are separate operations. Use Ruby’s CSV library rather than joining values with commas; CSV handles quoting, embedded commas, quotes, and line breaks.
require "csv"
CSV.open("results.csv", "wb", write_headers: false) do |csv|
rows.each { |row| csv << row }
end
With a header row and records, a table-oriented representation is convenient:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
require "csv"
table_data = CSV::Table.new(rows.map { |row| CSV::Row.new(rows.first, row) })
puts table_data.by_col["Status"]
Construct CSV::Table only after deciding which row supplies the headers and ensuring every data row is aligned. If spans were normalized, use the rectangular grid rather than the original variable-length arrays.
Safe parsing for downloaded or user-supplied HTML
Nokogiri treats input as untrusted by default. Its parser does not load external DTDs or access the network for external resources. Keep those protections enabled when processing scraped or uploaded documents.
- Do not disable network protections for convenience.
- Do not enable external entity or DTD behavior for untrusted HTML.
- Remember that parser safety does not grant permission to fetch a site or bypass its access controls.
- Limit input size and validate the selected table before storing extracted data.
Complete reusable Ruby script
This script accepts a file and a CSS selector, fails clearly when the table is absent, and writes a CSV file.
Rank #4
#!/usr/bin/env ruby
require "nokogiri"
require "csv"
path = ARGV.fetch(0) { abort "usage: ruby capture_table.rb page.html [selector] [output.csv]" }
selector = ARGV.fetch(1, "table")
out_path = ARGV.fetch(2, "table.csv")
html = File.binread(path)
doc = Nokogiri::HTML(html)
table = doc.at_css(selector) or abort "table not found for selector #{selector.inspect}"
rows = table.css("tr").map do |row|
row.css("th, td").map { |cell| cell.text.strip }
end
abort "no rows found" if rows.empty?
CSV.open(out_path, "wb") do |csv|
rows.each { |row| csv << row }
end
puts "wrote #{rows.length} rows to #{out_path}"
Run it with ruby capture_table.rb page.html 'table#results' results.csv. Change the selector to match the actual page and inspect the generated CSV before automating large batches.
Performance, reliability, and runtime choices
- Scope early: selecting one table before iterating reduces accidental work and prevents data from unrelated tables entering the result.
- Read once: parse the document once, then reuse the selected node for headers, rows, and validation.
- Prefer deterministic inputs: save the HTML and record Ruby, Nokogiri, parser, and encoding details when a result must be reproducible.
- Expect malformed markup: HTML parsers repair broken structure; test selectors against real samples rather than relying on source indentation.
- Validate shape: check row counts, expected headers, and span-induced irregularities before importing data.
Common failures and fixes
“table not found”
The selector may target the wrong ID or class, the table may be inserted by JavaScript after the original HTML loads, or the file may not contain the expected page. Print doc.css("table").length, inspect available attributes, and confirm that you are parsing the HTML response rather than a login or error page.
Rows have different lengths
Check for rowspan, colspan, nested tables, and repeated header rows. Use the grid normalizer above when visual column alignment matters; otherwise preserve the variable-length cell arrays and document that choice.
HTML5 parser fails on JRuby
Nokogiri::HTML5 is not available on JRuby. Use the supported parser API for that runtime, or run the HTML5 parser on CRuby after confirming behavior against your project’s Nokogiri version.
Accented text is corrupted
Check the source encoding and the way the IO is passed to Nokogiri. The parser’s text values are UTF-8, but an incorrect input-encoding assumption can still produce bad characters. Test names and symbols, not just ASCII fixtures.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
CSV columns shift
Do not concatenate values with commas. Serialize through CSV, and ensure every row is aligned to the same header set after handling spans and missing cells.
Or skip the browser setup
If your goal is to obtain a clean rendered image or PDF of the page containing the table before processing it, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page and element captures, dark mode, device presets or custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, easing migration.
See the ScreenshotNeo documentation for request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots each month without a card; paid plans start at $5 for 3,000 shots, with every feature on every plan. Sign up for the free ScreenshotNeo plan.
Frequently Asked Questions
Can Nokogiri execute JavaScript to create a table?
No. Nokogiri parses the HTML it receives; it does not provide a browser runtime. Obtain rendered HTML with an appropriate browser workflow or capture service, then parse the resulting document.
Should I use CSS or XPath for table extraction?
Use whichever expresses the target most clearly. CSS is concise for IDs and classes; XPath is useful for structural or text-based conditions. Both can scope rows and cells to one selected table.
How do I keep formulas, links, or images inside cells?
The examples extract visible text with cell.text. Select attributes such as cell["href"] or inspect child nodes separately when the cell’s markup, rather than its text, is the data you need.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

