Short answer: choose jsoup when the data is already present in the server-returned HTML, HtmlUnit when JavaScript and browser-like state must run inside Java, and Selenium when you need to automate an actual browser. Those are the three choices supported by current official documentation. A defensible, current ranking of ten distinct Java scraping libraries is not established here, so padding the list with unverified or generic HTTP clients would be misleading.
The decision in one table
| Tool | Executes page JavaScript? | Uses a real graphical browser? | Best fit | Main trade-off |
|---|---|---|---|---|
| jsoup | No | No | Static HTML extraction, cleanup and parsing | Cannot see content generated only after JavaScript runs |
| HtmlUnit | Yes, through browser simulation | No; it is GUI-less | JavaScript-dependent pages without a full browser runtime | Compatibility can vary by site and JavaScript behavior |
| Selenium | Yes | Yes | Browser-specific behavior, interactions and end-to-end automation | Requires browser and driver/runtime setup |
The official HtmlUnit comparison makes the central distinction: jsoup performs static extraction, HtmlUnit simulates a browser with JavaScript, and Selenium automates real browsers. Treat execution model—not an unsupported speed claim—as the primary selection criterion.
1. jsoup: the default for ordinary HTML
jsoup fetches and parses HTML, exposes DOM traversal, CSS selectors and XPath extraction, and follows the WHATWG HTML specification while handling malformed real-world markup. Its site currently displays version 1.23.2 and an MIT license; verify the version and Java runtime requirements before adding it to a new project.
When jsoup is the right choice
- The fields you need appear in the initial HTTP response.
- You want a small dependency and straightforward selector-based extraction.
- You also need to clean, transform or serialize HTML.
Runnable Java example
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
public class ScrapeWithJsoup {
public static void main(String[] args) throws Exception {
Document doc = Jsoup.connect("https://example.com")
.userAgent("Mozilla/5.0 (compatible; ExampleBot/1.0)")
.timeout(30_000)
.get();
for (Element link : doc.select("a[href]")) {
System.out.println(link.text() + " -> " + link.absUrl("href"));
}
}
}
The Connection API is both an HTTP client and a session object. It supports cookies, headers, redirects and proxy configuration. Sessions retain cookies in memory, so avoid keeping one indefinitely; for concurrent work, create a new request per operation rather than sharing one request object across threads. The documentation describes HTTP/2 use on JVM 11 and above.
Recommended Free Tools
2. HtmlUnit: JavaScript without a visible browser
HtmlUnit describes WebClient as a browser-like entry point. It retrieves pages, executes JavaScript, manages cookies and redirects, and preserves browser state across navigation. Page objects expose the DOM and support links, forms and extraction. This makes HtmlUnit useful when scripts modify the content you need but a full graphical browser is unnecessary.
Runnable Java example
import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;
public class ScrapeWithHtmlUnit {
public static void main(String[] args) throws Exception {
try (WebClient client = new WebClient()) {
client.getOptions().setJavaScriptEnabled(true);
client.getOptions().setCssEnabled(false);
client.getOptions().setThrowExceptionOnScriptError(false);
client.getOptions().setTimeout(30_000);
HtmlPage page = client.getPage("https://example.com");
client.waitForBackgroundJavaScript(5_000);
page.getByXPath("//a[@href]").forEach(node ->
System.out.println(node.asNormalizedText())
);
}
}
}
Version and compatibility caution
The HtmlUnit project page reports release 5.5.0 on August 30, 2026. Release numbers and JavaScript compatibility change, so confirm the current project page and test the exact sites you intend to crawl. A simulated browser is not identical to Chrome or Firefox; behavior can differ when a site depends on browser-specific APIs.
3. Selenium: real-browser automation
Selenium automates real browsers. Choose it when the requirement is browser-specific behavior, interactive workflows, or end-to-end automation rather than simply parsing a response. You must provision a browser and the matching driver or Selenium Manager setup, and account for the additional runtime in deployment.
Runnable Java example
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
public class ScrapeWithSelenium {
public static void main(String[] args) {
WebDriver driver = new ChromeDriver();
try {
driver.get("https://example.com");
for (var link : driver.findElements(By.cssSelector("a[href]"))) {
System.out.println(link.getText() + " -> " + link.getAttribute("href"));
}
} finally {
driver.quit();
}
}
}
Use explicit waits for a condition that matters to your extraction instead of sleeping for an arbitrary period. Keep browser instances isolated between jobs, limit concurrency to what the host can support, and always call quit() in a finally block so orphaned processes do not accumulate.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Why a literal “10 best” ranking would be unreliable
The available official material does not validate ten distinct, currently maintained Java scraping libraries, nor does it provide reproducible comparative benchmarks for speed, accuracy or adoption. Generic HTTP clients and HTML parsers can be useful components, but labeling them the next seven “best scraping libraries” would imply evidence that is not available. Before adding another dependency, check its release activity, Java version support, license, JavaScript requirements and documentation yourself.
Rank #2
A practical selection checklist
- Inspect the response first: if the target text is in returned HTML, start with jsoup.
- Identify post-load rendering: if a script inserts the data, try HtmlUnit when browser simulation is sufficient.
- Require actual browser behavior: use Selenium for browser APIs, complex interactions or workflows that must match a user browser.
- Separate fetching from extraction: define selectors, normalization and validation independently so a later tool change does not rewrite your data model.
- Respect access rules: neither parsing nor browser automation is evidence of bypassing CAPTCHAs, bot checks or other access controls. Follow terms, applicable law and published crawling policies.
Reliable scraping patterns in Java
Make requests bounded and observable
Set connection and read timeouts, record the URL and HTTP outcome, and cap response sizes where your client permits it. Retries should be limited and should not blindly repeat authentication failures or deliberate access denials.
Preserve the right state
Cookies, redirects, headers and authentication can change what a server returns. jsoup sessions retain cookies in memory; HtmlUnit keeps browser state in its WebClient; Selenium keeps state in its browser profile. Scope that state to a job or account and avoid sharing mutable clients across threads.
Validate extracted data
Selectors can match zero, one or many elements after a redesign. Treat missing required fields as a structured extraction error, normalize whitespace and URLs, and retain enough logging to diagnose which selector failed without storing secrets.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Common failures and fixes
The HTML contains no data
Cause: the page inserts it with JavaScript. Fix: confirm the network response and switch from jsoup to HtmlUnit or Selenium when rendering is required.
HtmlUnit output differs from the visible site
Cause: site code may depend on browser APIs or JavaScript that the simulator does not reproduce. Fix: disable nonessential scripts, wait for the specific element you need, or move the workflow to Selenium.
Selenium cannot start a session
Cause: the browser, driver or container permissions are mismatched. Fix: install a supported browser, let Selenium Manager resolve the driver where appropriate, run headless only when your environment requires it, and inspect the first startup exception rather than retrying indefinitely.
Selectors suddenly return nothing
Cause: markup or class names changed, or the request received a consent, login or block page. Fix: save a diagnostic response, check the final URL and page title, and add explicit validation for the expected document.
Jobs become slower or exhaust memory
Cause: unbounded sessions, browser processes or large documents. Fix: close clients deterministically, limit parallel browsers, avoid retaining whole documents after extraction, and measure your own workload; no universal speed ranking is established by the available sources.
Or skip the browser setup
ScreenshotNeo is the practical alternative when your deliverable is a page image or PDF rather than extracted fields. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
One GET request
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for output formats and options. The service also supports full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification.
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card. Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #4
FAQ
Can jsoup scrape a single-page application?
Only when the required data is included in the server response. If JavaScript creates the data after load, use HtmlUnit or Selenium instead.
Is HtmlUnit a replacement for Selenium?
It can replace a real browser for some JavaScript-driven pages, but it does not guarantee identical browser behavior. Use Selenium when that fidelity is part of the requirement.
Do these libraries bypass CAPTCHAs?
No. The documented capabilities describe fetching, parsing, simulation or browser automation, not bypassing access controls.
Which tool should run in a small container?
Start with jsoup when rendering is unnecessary. HtmlUnit avoids a graphical browser but still needs JavaScript resources. Selenium requires the browser and its runtime, so size and operate the container accordingly.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Frequently Asked Questions
Can jsoup scrape a single-page application?
Only when the required data is included in the server response. If JavaScript creates the data after load, use HtmlUnit or Selenium instead.
Best Value
Is HtmlUnit a replacement for Selenium?
It can replace a real browser for some JavaScript-driven pages, but it does not guarantee identical browser behavior. Use Selenium when that fidelity is part of the requirement.
Do these libraries bypass CAPTCHAs?
No. The documented capabilities describe fetching, parsing, simulation or browser automation, not bypassing access controls.
Which tool should run in a small container?
Start with jsoup when rendering is unnecessary. HtmlUnit avoids a graphical browser but still needs JavaScript resources. Selenium requires the browser and its runtime, so size and operate the container accordingly.
The Bottom Line
Use jsoup for HTML that already exists, HtmlUnit for JavaScript-driven pages where a simulated browser is enough, and Selenium when only a real browser will do. Do not treat an unsupported ten-library ranking or universal speed claim as fact.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

