Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsShort answer: choose a renderer based on the HTML you actually have. OpenHTMLtoPDF is a practical pure-Java option for well-formed XHTML/XML and supported CSS, but it does not execute JavaScript and is not a browser replacement. If your document relies on modern HTML5, CSS3, or JavaScript, evaluate Flying Saucer’s Chrome-backed PDF module, which delegates rendering to chrome-headless-shell. Apache PDFBox is useful for creating and manipulating PDFs, not for laying out arbitrary HTML and CSS.
The reliable workflow is to normalize the HTML, select the renderer, make assets and fonts available, render representative pages, and inspect pagination and security behavior before shipping.
Choose the rendering route first
| Requirement | Best starting point | Important limitation or dependency |
|---|---|---|
| Deliberately authored XHTML/XML, supported CSS, no JavaScript | OpenHTMLtoPDF | Pure Java; supports a reasonable subset of XHTML/XML and some HTML5. It does not run JavaScript and does not implement many modern features, including flex and grid. |
| XHTML and CSS 2.1 in the Flying Saucer family | Flying Saucer PDF module | Pure-Java XML/XHTML renderer. The required Java version depends on the release line: the project README lists Java 11+ for 9.5.0, Java 17+ for 9.6.0, and Java 21+ for 10.0.0. |
| Modern HTML5/CSS3 or browser-like behavior | Flying Saucer Chrome PDF module | Delegates PDF generation to chrome-headless-shell; deploy and patch that external browser component with your application. |
| PDF editing, forms, extraction, signing, or image operations after rendering | Apache PDFBox | PDFBox is a PDF creation and manipulation library, not an HTML/CSS layout engine. |
Do not pick a library from its name alone. Test the exact templates, CSS, images, fonts, page breaks, and Java runtime you will deploy. The inspected project documentation does not establish an empirical performance winner for a typical workload.
Prepare HTML that a Java renderer can understand
OpenHTMLtoPDF and the pure-Java Flying Saucer path expect well-formed XML/XHTML more than arbitrary browser HTML. Make the document deterministic before conversion:
Free tools Windows power users keep installed
One-click scans. No signup required.
- Include a single root element, correctly closed tags, quoted attributes, and valid nesting.
- Use explicit character encoding, normally UTF-8, and ensure the Java input stream uses the same encoding.
- Resolve images, stylesheets, and fonts with stable absolute URLs or a controlled base URI.
- Replace browser-only layout assumptions with supported CSS. The OpenHTMLtoPDF README specifically cautions that modern HTML5 should be crafted for its engine; flex and grid are not implemented.
- Prefer table-based layout for print sections and avoid floats close to page boundaries, where the project documentation warns that results can be poor.
For untrusted HTML, isolate conversion, restrict network access, validate URLs, and use the selected release’s security guidance. XML parser configuration matters: Flying Saucer’s changelog records hardening of DocumentBuilderFactory usage against XXE in a recent release, but that is not a blanket guarantee for every version or application.
OpenHTMLtoPDF: a pure-Java implementation
The following program converts an XHTML file to a PDF. Resolve the current OpenHTMLtoPDF runtime modules from the project’s integration documentation for your chosen release; a parent POM version is not automatically the module your application should add. The API below is the standard builder pattern used by the project family.
import java.io.File;
import java.io.FileOutputStream;
import java.io.OutputStream;
import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
public final class HtmlToPdf {
public static void main(String[] args) throws Exception {
if (args.length != 2) {
throw new IllegalArgumentException("Usage: HtmlToPdf input.xhtml output.pdf");
}
File input = new File(args[0]).getCanonicalFile();
File output = new File(args[1]).getCanonicalFile();
if (!input.isFile()) {
throw new IllegalArgumentException("Input does not exist: " + input);
}
File parent = output.getParentFile();
if (parent != null) parent.mkdirs();
try (OutputStream out = new FileOutputStream(output)) {
PdfRendererBuilder builder = new PdfRendererBuilder();
builder.useFastMode();
builder.withFile(input);
builder.toStream(out);
builder.run();
}
}
}
Compile and run it with the renderer modules on your class path:
java HtmlToPdf invoice.xhtml invoice.pdf
withFile gives the renderer a base location for relative resources. If you supply HTML as a string instead, also provide a base URI so relative CSS, images, and fonts can be found. Add explicit font registrations when a required typeface is not installed in the runtime container, then inspect the resulting PDF for missing glyphs and fallback fonts.
Resource and page-break example
Keep print rules conservative and explicit:
<style>
@page { size: A4; margin: 18mm 16mm 20mm; }
body { font-family: "Noto Sans", sans-serif; font-size: 10pt; }
.page-break { page-break-before: always; }
table { width: 100%; border-collapse: collapse; }
td, th { border: 0.2pt solid #777; padding: 4pt; }
</style>
Render a long document, a document containing images, and a document containing non-Latin text. Check that tables do not split in unacceptable places, headings remain with their following content, images have the expected resolution, and the final page is not unexpectedly blank.
Rank #2
When Flying Saucer is the better fit
Flying Saucer’s pure-Java PDF module is appropriate when your input is well-formed XML/XHTML and CSS 2.1 is sufficient. Select the Java runtime required by the exact release rather than relying on a generic “Flying Saucer supports Java” statement. The project README lists Java 11+ for 9.5.0, Java 17+ for 9.6.0, and Java 21+ for 10.0.0.
For pages that require modern HTML5/CSS3, the project lists a Chrome PDF module that delegates to chrome-headless-shell. This can provide browser-style capabilities, but it changes operations: the browser binary must be present, compatible with the module, patched, sandboxed appropriately, and monitored. Test startup time, filesystem permissions, fonts, proxy settings, and concurrent job limits in the same container image used in production.
Why PDFBox is not an HTML converter
PDFBox provides APIs for PDF creation and manipulation, extraction, forms, printing, images, and signing. It does not parse HTML and CSS into a paginated layout. A common architecture is therefore: render HTML with OpenHTMLtoPDF or Flying Saucer, then use PDFBox for post-processing such as merging, metadata, stamping, or signing.
Recommended Free Tools
Input, asset, and font handling
Local files
Use a canonical, controlled input directory and a base URI. Do not let user-supplied paths escape that directory. Confirm that the process account can read every stylesheet, image, and font.
Remote assets
Remote URLs make output dependent on DNS, TLS, authentication, and changing content. Prefer downloading approved assets first, checking size and content type, and rendering from a local staging directory. If remote loading is unavoidable, apply strict allowlists, timeouts, response-size limits, and no-redirect or limited-redirect policies.
Fonts and international text
Install or register the fonts required by the document. A PDF can be structurally valid while displaying tofu boxes or substituted glyphs. Verify ligatures, right-to-left scripts, combining marks, and emoji separately; support varies by renderer, font, and PDF viewer.
JavaScript and dynamic data
OpenHTMLtoPDF does not execute JavaScript. A chart or table populated in the browser will be absent unless you render the data into the HTML first. If JavaScript and browser APIs are essential, test the Chrome-backed route against the exact page rather than assuming a successful browser screen capture implies identical PDF output.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchProduction checklist
- Pin the renderer and every transitive dependency; review the exact release notes and license obligations.
- Validate and normalize HTML before handing it to the renderer.
- Set bounded timeouts for asset retrieval and conversion jobs.
- Limit memory, CPU, document size, image dimensions, and page count for untrusted input.
- Run conversion in an isolated worker when input is user-controlled.
- Capture logs containing a document ID, renderer version, elapsed time, page count, and failure category without logging secrets.
- Compare generated PDFs in automated tests for representative templates, fonts, page breaks, and images.
- Open the output with a PDF validator and at least one independent viewer.
Troubleshooting common failures
Output is blank or missing sections
Cause: malformed markup, unsupported CSS, or content created only by JavaScript. Fix: validate XHTML, simplify unsupported layout, pre-render dynamic data, and test the same file with the chosen engine’s supported feature set.
Images or styles are missing
Cause: an incorrect base URI, inaccessible relative path, blocked network request, or unsupported image format. Fix: use canonical local paths or approved absolute URLs, verify process permissions, and inspect the conversion log.
Fonts are substituted or characters disappear
Cause: the font is not installed, not registered, or lacks the required glyphs. Fix: package the font, register it with the renderer where supported, and test the target script and weights.
Rank #4
Layout differs from Chrome
Cause: a pure-Java renderer is not a full browser and may not support flex, grid, JavaScript, or other modern standards. Fix: rewrite the print CSS for the selected engine or evaluate the Flying Saucer Chrome PDF module.
Conversion hangs or consumes excessive memory
Cause: huge images, unbounded remote resources, pathological markup, or too many concurrent jobs. Fix: cap input and resource sizes, add timeouts, isolate workers, and limit concurrency.
XXE or unexpected external access is reported
Cause: unsafe XML parsing or unrestricted resource resolution. Fix: use a maintained release, disable external entities and unrelated external access in the parser and resolver, and enforce an allowlist for resources. Verify the configuration with security tests.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Licensing and release diligence
OpenHTMLtoPDF states that it is LGPL 2.1-or-later. Its PDF/A testing module has a separate GPL exception and is not distributed to Maven Central, so do not assume every module has identical terms. PDFBox is Apache License 2.0. Flying Saucer’s exact artifacts and dependencies require their own review. The Flying Saucer changelog dates version 10.4.0 to July 16, 2026 and records CSS transform work, inline PDF elements, SVG fixes, and XXE-related hardening. Apache’s site lists PDFBox 3.0.8 as released July 11, 2026. Confirm current security advisories, licenses, and Java requirements before deployment.
Or skip the browser setup
If your goal is simply to obtain a PDF or image of a URL, ScreenshotNeo provides a hosted capture API rather than requiring you to package a Java renderer or headless browser. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →For a PDF capture, call the API endpoint shown in the ScreenshotNeo documentation:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://stripe.com
-o page.pdf
The same endpoint can be called from Java through any HTTP client, or from Python and Node.js:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can OpenHTMLtoPDF convert any webpage?
No. It is intended for well-formed, deliberately authored markup and supported CSS, not an arbitrary browser page with JavaScript, flexbox, grid, or other unsupported standards.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Should I use PDFBox instead?
Use PDFBox after rendering when you need PDF manipulation. It is not the HTML/CSS renderer itself.
Which Java version should I install for Flying Saucer?
Use the requirement for the exact release: the project README lists Java 11+ for 9.5.0, Java 17+ for 9.6.0, and Java 21+ for 10.0.0.
Is a Chrome-backed renderer always more accurate?
It is the route to evaluate for modern browser features, but the external browser introduces deployment and operational requirements. Validate your actual templates and environment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

