What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Importing HTML in Rust is a two-step operation: read the file into memory, then pass the resulting text to an HTML parser. For CSS selectors and text extraction, the scraper crate is usually the simplest choice. Use Html::parse_document for a complete page and Html::parse_fragment for an HTML snippet. If you need to modify a DOM-like tree, use Kuchiki; if you need a lower-level standards-oriented parser, use html5ever.

Read the file, then parse the HTML

Rust’s standard library does not include an HTML parser. The standard-library part ends after loading the file:

let html = std::fs::read_to_string("page.html")?;

read_to_string reads the complete file into a String and therefore requires valid UTF-8. Parsing is a separate step performed by a crate such as scraper, Kuchiki, or html5ever.

Minimal selector-based example with scraper

Add the crate with Cargo:

cargo add scraper

Create src/main.rs:

use scraper::{Html, Selector};
use std::error::Error;
use std::fs;

fn main() -> Result<(), Box<dyn Error>> {
    let html = fs::read_to_string("page.html")?;
    let document = Html::parse_document(&html);
    let title_selector = Selector::parse("title")?;

    if let Some(title) = document.select(&title_selector).next() {
        let title_text = title.text().collect::<String>();
        println!("{title_text}");
    } else {
        println!("The document has no title element");
    }

    Ok(())
}

Run it from the directory that contains page.html with cargo run. The ? operator propagates both file-system errors and selector-parsing errors to main.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose document parsing or fragment parsing

Complete HTML document

Use Html::parse_document when the input represents a page, normally including elements such as <html>, <head>, and <body>. The parser builds a document that you can query with CSS selectors.

let source = std::fs::read_to_string("page.html")?;
let document = scraper::Html::parse_document(&source);

HTML fragment

Use Html::parse_fragment when the file contains only a snippet, such as a list item or a component template:

use scraper::Html;

let snippet = std::fs::read_to_string("item.html")?;
let fragment = Html::parse_fragment(&snippet);

Fragment parsing avoids treating a snippet as an entire page. This is useful when your input is <li>Item</li>, a table row, or markup intended to be inserted into an existing document.

Extract text, attributes, and several elements

Read text without keeping markup

use scraper::{Html, Selector};
use std::error::Error;
use std::fs;

fn main() -> Result<(), Box<dyn Error>> {
    let source = fs::read_to_string("page.html")?;
    let document = Html::parse_document(&source);
    let heading = Selector::parse("h1")?;

    for element in document.select(&heading) {
        let text = element.text().collect::<Vec<_>>().join(" ");
        println!("{}", text.trim());
    }

    Ok(())
}

text() walks the element’s descendant text nodes. Joining a vector with spaces prevents words from adjacent inline nodes from running together.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read an attribute

let links = Selector::parse("a[href]")?;

for link in document.select(&links) {
    if let Some(href) = link.value().attr("href") {
        let label = link.text().collect::<String>();
        println!("{} -> {}", label.trim(), href);
    }
}

Use the selector to limit matches, then call value().attr("name") for an optional attribute. The result is None when the attribute is absent, so do not unwrap it unless the markup is guaranteed to contain it.

Parse selectors once

Selector parsing can fail, while selecting elements is cheap to repeat. In a loop over many files, construct each Selector once outside the loop and reuse it. This keeps selector errors near program startup and avoids needless work for every document.

Handle file paths and input errors explicitly

A relative path is resolved against the process’s current working directory, not necessarily the directory containing your Rust source. Print or log the path you intend to open when diagnosing deployment problems.

use std::fs;

fn load_page(path: &str) -> Result<String, std::io::Error> {
    fs::read_to_string(path)
}

fn main() {
    match load_page("page.html") {
        Ok(html) => println!("loaded {} bytes", html.len()),
        Err(error) => eprintln!("could not read page.html: {error}"),
    }
}
  • Not found: verify the current directory, spelling, and case of the path. Use an absolute path temporarily to distinguish a path problem from a parser problem.
  • Permission denied: grant the running process read access or choose a file it is allowed to open.
  • Invalid UTF-8: use the byte-loading path below instead of silently replacing or discarding bytes.

When the file is not valid UTF-8

std::fs::read returns the complete file as Vec<u8>, so it does not impose a UTF-8 requirement:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
use std::error::Error;
use std::fs;

fn main() -> Result<(), Box<dyn Error>> {
    let bytes = fs::read("page.html")?;
    let html = String::from_utf8(bytes)?;
    println!("loaded {} characters", html.chars().count());
    Ok(())
}

String::from_utf8 reports malformed UTF-8 as an error. If your application has a deliberate replacement policy, use String::from_utf8_lossy instead, but remember that replacement characters change the input before parsing. For files in a known legacy encoding, decode the bytes with an encoding library appropriate to that encoding before passing the resulting string to the HTML parser.

Choose the right Rust HTML crate

Crate Best fit Parsing model Mutation and effort
scraper CSS-selector queries, text, attributes, and serialization Complete documents and fragments High-level API; a practical default for extraction
Kuchiki DOM-like traversal and tree manipulation Document and fragment parsing through html5ever Higher-level mutable tree; useful when you must edit nodes
html5ever Custom parsing or serialization pipelines WHATWG/HTML5 parsing callbacks Lower level; it does not provide a DOM tree by itself

The table describes the API roles rather than performance rankings. Pick the smallest abstraction that matches the work you need to do.

Use Kuchiki for a mutable tree

Kuchiki parses HTML into a DOM-like tree. A typical selector traversal looks like this:

use kuchiki::traits::*;
use std::error::Error;
use std::fs;

fn main() -> Result<(), Box<dyn Error>> {
    let html = fs::read_to_string("page.html")?;
    let document = kuchiki::parse_html().one(html);

    for node in document.select("title")? {
        let text = node.text_contents();
        println!("{}", text.trim());
    }

    Ok(())
}

Choose Kuchiki when you need to retain a tree, alter nodes, or perform repeated DOM-style operations. Its document and fragment parsers are built on html5ever.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use html5ever when you need parser-level control

html5ever can parse and serialize HTML according to the WHATWG HTML5 specifications, but it uses callbacks and does not create a DOM representation on its own. That makes it suitable for a custom sink or streaming-style pipeline, while scraper or Kuchiki is usually less work for ordinary application code.

Build a reusable importer

For an application that imports many files, isolate I/O from parsing so each stage has a clear error boundary:

use scraper::{Html, Selector};
use std::error::Error;
use std::path::Path;

fn page_title(path: impl AsRef<Path>) -> Result<Option<String>, Box<dyn Error>> {
    let source = std::fs::read_to_string(path)?;
    let document = Html::parse_document(&source);
    let selector = Selector::parse("title")?;

    Ok(document
        .select(&selector)
        .next()
        .map(|node| node.text().collect::<String>().trim().to_owned()))
}

fn main() -> Result<(), Box<dyn Error>> {
    match page_title("page.html")? {
        Some(title) => println!("title: {title}"),
        None => println!("no title found"),
    }
    Ok(())
}

This function distinguishes three outcomes: the file could not be read, parsing or selector construction failed, or the file was valid but had no matching title. That distinction is more useful than returning an empty string for every failure.

Performance and reliability considerations

  • Memory: read_to_string and read load the complete file, so peak memory includes the input and the parser’s representation. For very large files, consider whether a lower-level parser pipeline is more appropriate.
  • Malformed markup: HTML parsers generally recover from common authoring errors. Treat recovered output as parsed data, not proof that the source was valid; validate required elements after parsing.
  • Repeated work: reuse selectors, avoid reparsing the same source, and keep the parsed document alive while extracting all required fields.
  • Untrusted files: impose application-level limits on file size and processing time before importing user-provided HTML. Do not execute scripts merely because they appear in the source; parsing HTML is not the same as running it in a browser.
  • Testing: include fixtures for a full document, a fragment, missing elements, malformed nesting, empty files, missing files, and invalid UTF-8 bytes.

Troubleshooting selectors and parser results

“The selector finds nothing”

Check that you used parse_document for a page and parse_fragment for a snippet, then print a small portion of the source to confirm the expected tag, class, or attribute is actually present. CSS selectors are case-sensitive for many values, and a class selector such as .price matches a class, not an element’s visible text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The title text is empty or has extra whitespace”

An element can contain nested tags or whitespace-only text nodes. Collect descendant text and normalize it with trim or a deliberate whitespace policy instead of assuming one text node.

“The program works in the IDE but not from a shell”

The working directory differs between launchers. Use std::env::current_dir() while diagnosing, or construct a path from a configured application directory rather than relying on a relative path.

“Kuchiki code does not compile”

Ensure the crate’s selector and node traits are imported (for example, kuchiki::traits::*) and that your dependency version matches the API examples you are using. Keep the crate choice consistent; scraper and Kuchiki expose different node types and selection APIs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your real goal is to obtain a clean image or PDF of a web page rather than parse a local HTML file, ScreenshotNeo provides a single HTTP request. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for all options, including full-page or CSS-selector captures, device and viewport settings, dark mode, retina scale, PDF paper and margin controls, custom CSS or JavaScript, click and wait actions, request blocking, headers and cookies, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing gives two months free. Sign up for the free ScreenshotNeo plan.

FAQ

Can Rust parse HTML without a third-party crate?

The standard library can read the bytes or text, but it does not provide an HTML parser. Add a crate suited to your extraction or tree-editing requirements.

Should I parse a saved web page as a document or fragment?

Use document parsing when the file is a complete page and fragment parsing when it is a standalone snippet. The choice controls how the parser constructs the root structure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does parsing local HTML run JavaScript?

No. These Rust parsers read and represent markup; they do not provide a browser runtime that executes scripts or fetches subresources.

Frequently Asked Questions

Can Rust parse HTML without a third-party crate?

The standard library can read the bytes or text, but it does not provide an HTML parser. Add a crate suited to your extraction or tree-editing requirements.

Should I parse a saved web page as a document or fragment?

Use document parsing when the file is a complete page and fragment parsing when it is a standalone snippet. The choice controls how the parser constructs the root structure.

Does parsing local HTML run JavaScript?

No. These Rust parsers read and represent markup; they do not provide a browser runtime that executes scripts or fetches subresources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.