Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: choose jsoup when the data is already present in the server-returned HTML, HtmlUnit when JavaScript and browser-like state must run inside Java, and Selenium when you need to automate an actual browser. Those are the three choices supported by current official documentation. A defensible, current ranking of ten distinct Java scraping libraries is not established here, so padding the list with unverified or generic HTTP clients would be misleading.

The decision in one table

Tool Executes page JavaScript? Uses a real graphical browser? Best fit Main trade-off
jsoup No No Static HTML extraction, cleanup and parsing Cannot see content generated only after JavaScript runs
HtmlUnit Yes, through browser simulation No; it is GUI-less JavaScript-dependent pages without a full browser runtime Compatibility can vary by site and JavaScript behavior
Selenium Yes Yes Browser-specific behavior, interactions and end-to-end automation Requires browser and driver/runtime setup

The official HtmlUnit comparison makes the central distinction: jsoup performs static extraction, HtmlUnit simulates a browser with JavaScript, and Selenium automates real browsers. Treat execution model—not an unsupported speed claim—as the primary selection criterion.

1. jsoup: the default for ordinary HTML

jsoup fetches and parses HTML, exposes DOM traversal, CSS selectors and XPath extraction, and follows the WHATWG HTML specification while handling malformed real-world markup. Its site currently displays version 1.23.2 and an MIT license; verify the version and Java runtime requirements before adding it to a new project.

When jsoup is the right choice

  • The fields you need appear in the initial HTTP response.
  • You want a small dependency and straightforward selector-based extraction.
  • You also need to clean, transform or serialize HTML.

Runnable Java example

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;

public class ScrapeWithJsoup {
    public static void main(String[] args) throws Exception {
        Document doc = Jsoup.connect("https://example.com")
                .userAgent("Mozilla/5.0 (compatible; ExampleBot/1.0)")
                .timeout(30_000)
                .get();

        for (Element link : doc.select("a[href]")) {
            System.out.println(link.text() + " -> " + link.absUrl("href"));
        }
    }
}

The Connection API is both an HTTP client and a session object. It supports cookies, headers, redirects and proxy configuration. Sessions retain cookies in memory, so avoid keeping one indefinitely; for concurrent work, create a new request per operation rather than sharing one request object across threads. The documentation describes HTTP/2 use on JVM 11 and above.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. HtmlUnit: JavaScript without a visible browser

HtmlUnit describes WebClient as a browser-like entry point. It retrieves pages, executes JavaScript, manages cookies and redirects, and preserves browser state across navigation. Page objects expose the DOM and support links, forms and extraction. This makes HtmlUnit useful when scripts modify the content you need but a full graphical browser is unnecessary.

Runnable Java example

import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;

public class ScrapeWithHtmlUnit {
    public static void main(String[] args) throws Exception {
        try (WebClient client = new WebClient()) {
            client.getOptions().setJavaScriptEnabled(true);
            client.getOptions().setCssEnabled(false);
            client.getOptions().setThrowExceptionOnScriptError(false);
            client.getOptions().setTimeout(30_000);

            HtmlPage page = client.getPage("https://example.com");
            client.waitForBackgroundJavaScript(5_000);

            page.getByXPath("//a[@href]").forEach(node ->
                System.out.println(node.asNormalizedText())
            );
        }
    }
}

Version and compatibility caution

The HtmlUnit project page reports release 5.5.0 on August 30, 2026. Release numbers and JavaScript compatibility change, so confirm the current project page and test the exact sites you intend to crawl. A simulated browser is not identical to Chrome or Firefox; behavior can differ when a site depends on browser-specific APIs.

3. Selenium: real-browser automation

Selenium automates real browsers. Choose it when the requirement is browser-specific behavior, interactive workflows, or end-to-end automation rather than simply parsing a response. You must provision a browser and the matching driver or Selenium Manager setup, and account for the additional runtime in deployment.

Runnable Java example

import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;

public class ScrapeWithSelenium {
    public static void main(String[] args) {
        WebDriver driver = new ChromeDriver();
        try {
            driver.get("https://example.com");
            for (var link : driver.findElements(By.cssSelector("a[href]"))) {
                System.out.println(link.getText() + " -> " + link.getAttribute("href"));
            }
        } finally {
            driver.quit();
        }
    }
}

Use explicit waits for a condition that matters to your extraction instead of sleeping for an arbitrary period. Keep browser instances isolated between jobs, limit concurrency to what the host can support, and always call quit() in a finally block so orphaned processes do not accumulate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why a literal “10 best” ranking would be unreliable

The available official material does not validate ten distinct, currently maintained Java scraping libraries, nor does it provide reproducible comparative benchmarks for speed, accuracy or adoption. Generic HTTP clients and HTML parsers can be useful components, but labeling them the next seven “best scraping libraries” would imply evidence that is not available. Before adding another dependency, check its release activity, Java version support, license, JavaScript requirements and documentation yourself.

A practical selection checklist

  • Inspect the response first: if the target text is in returned HTML, start with jsoup.
  • Identify post-load rendering: if a script inserts the data, try HtmlUnit when browser simulation is sufficient.
  • Require actual browser behavior: use Selenium for browser APIs, complex interactions or workflows that must match a user browser.
  • Separate fetching from extraction: define selectors, normalization and validation independently so a later tool change does not rewrite your data model.
  • Respect access rules: neither parsing nor browser automation is evidence of bypassing CAPTCHAs, bot checks or other access controls. Follow terms, applicable law and published crawling policies.

Reliable scraping patterns in Java

Make requests bounded and observable

Set connection and read timeouts, record the URL and HTTP outcome, and cap response sizes where your client permits it. Retries should be limited and should not blindly repeat authentication failures or deliberate access denials.

Preserve the right state

Cookies, redirects, headers and authentication can change what a server returns. jsoup sessions retain cookies in memory; HtmlUnit keeps browser state in its WebClient; Selenium keeps state in its browser profile. Scope that state to a job or account and avoid sharing mutable clients across threads.

Validate extracted data

Selectors can match zero, one or many elements after a redesign. Treat missing required fields as a structured extraction error, normalize whitespace and URLs, and retain enough logging to diagnose which selector failed without storing secrets.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common failures and fixes

The HTML contains no data

Cause: the page inserts it with JavaScript. Fix: confirm the network response and switch from jsoup to HtmlUnit or Selenium when rendering is required.

HtmlUnit output differs from the visible site

Cause: site code may depend on browser APIs or JavaScript that the simulator does not reproduce. Fix: disable nonessential scripts, wait for the specific element you need, or move the workflow to Selenium.

Selenium cannot start a session

Cause: the browser, driver or container permissions are mismatched. Fix: install a supported browser, let Selenium Manager resolve the driver where appropriate, run headless only when your environment requires it, and inspect the first startup exception rather than retrying indefinitely.

Selectors suddenly return nothing

Cause: markup or class names changed, or the request received a consent, login or block page. Fix: save a diagnostic response, check the final URL and page title, and add explicit validation for the expected document.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Jobs become slower or exhaust memory

Cause: unbounded sessions, browser processes or large documents. Fix: close clients deterministically, limit parallel browsers, avoid retaining whole documents after extraction, and measure your own workload; no universal speed ranking is established by the available sources.

Or skip the browser setup

ScreenshotNeo is the practical alternative when your deliverable is a page image or PDF rather than extracted fields. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

One GET request

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for output formats and options. The service also supports full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification.

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

FAQ

Can jsoup scrape a single-page application?

Only when the required data is included in the server response. If JavaScript creates the data after load, use HtmlUnit or Selenium instead.

Is HtmlUnit a replacement for Selenium?

It can replace a real browser for some JavaScript-driven pages, but it does not guarantee identical browser behavior. Use Selenium when that fidelity is part of the requirement.

Do these libraries bypass CAPTCHAs?

No. The documented capabilities describe fetching, parsing, simulation or browser automation, not bypassing access controls.

Which tool should run in a small container?

Start with jsoup when rendering is unnecessary. HtmlUnit avoids a graphical browser but still needs JavaScript resources. Selenium requires the browser and its runtime, so size and operate the container accordingly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can jsoup scrape a single-page application?

Only when the required data is included in the server response. If JavaScript creates the data after load, use HtmlUnit or Selenium instead.

Is HtmlUnit a replacement for Selenium?

It can replace a real browser for some JavaScript-driven pages, but it does not guarantee identical browser behavior. Use Selenium when that fidelity is part of the requirement.

Do these libraries bypass CAPTCHAs?

No. The documented capabilities describe fetching, parsing, simulation or browser automation, not bypassing access controls.

Which tool should run in a small container?

Start with jsoup when rendering is unnecessary. HtmlUnit avoids a graphical browser but still needs JavaScript resources. Selenium requires the browser and its runtime, so size and operate the container accordingly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

Use jsoup for HTML that already exists, HtmlUnit for JavaScript-driven pages where a simulated browser is enough, and Selenium when only a real browser will do. Do not treat an unsupported ten-library ranking or universal speed claim as fact.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.