Use a renderer that matches the page you are converting. For HTML or XHTML you control, iText pdfHTML provides the shortest documented Java path with HtmlConverter.convertToPdf(...). OpenHTMLToPDF is a pure-Java, LGPL-compatible choice for a reasonable XHTML/CSS subset. A JavaScript-heavy, browser-oriented page needs a Chromium-backed approach; Flying Saucer lists a flying-saucer-chrome-pdf artifact that delegates PDF generation to chrome-headless-shell. Apache PDFBox is useful for PDF creation and post-processing, but it is not an HTML browser renderer by itself.
This guide shows a complete iText implementation, a pure-Java OpenHTMLToPDF alternative, how to handle relative assets and JavaScript, what Flying Saucer and PDFBox are actually suited for, and the failure modes that make Java HTML-to-PDF jobs unreliable.
Choose the renderer before you write code
The phrase “Java HTML to PDF” covers two different jobs: rendering a controlled HTML document, or reproducing a live web page. Those jobs have different requirements.
| Option | Best fit | Browser and JavaScript behavior | Runtime and licensing notes |
|---|---|---|---|
| iText pdfHTML | Controlled HTML/CSS, templates, invoices and reports | Use it as a document renderer; test advanced browser-only features rather than assuming full browser behavior. | Commercial and open-source licensing terms can differ by version and deployment model. Check the terms that apply to your project. |
| OpenHTMLToPDF | Well-formed XHTML and a reasonable CSS 2.1 subset | It is not a web browser, does not execute JavaScript, and does not implement many modern standards such as flex and grid. | Pure Java and LGPL-compatible. It uses Apache PDFBox underneath and includes modules for PDF/A, accessible PDF, SVG, MathML and font fallback. |
| Flying Saucer | XHTML/CSS 2.1 rendering, or browser-oriented output through its Chrome artifact | flying-saucer-pdf is the conventional XHTML/CSS route. flying-saucer-chrome-pdf delegates to chrome-headless-shell for browser-oriented rendering. |
The project states that 9.5.0 requires Java 11 or later, 9.6.0 Java 17 or later, and 10.0.0 Java 21 or later. Verify the selected artifact’s exact baseline. |
| Apache PDFBox | Creating, editing, rendering or post-processing PDFs around another renderer | It does not convert an arbitrary web page with browser-grade HTML, CSS and JavaScript on its own. | Apache open-source PDF infrastructure; OpenHTMLToPDF uses it as its PDF library. |
If the source is a static template, start with iText or OpenHTMLToPDF. If the source is a live JavaScript web page to PDF, use a browser-backed path and budget for an external browser process, or use a hosted capture service.
iText pdfHTML: the shortest Java implementation
iText’s documented minimal workflow accepts an HTML string and writes a PDF to a file stream. The API also accepts a File or InputStream, and can write to an output stream, file, PdfWriter or PdfDocument.
Convert an HTML string to a PDF file
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;
public final class HtmlToPdf {
private HtmlToPdf() {
}
public static void createPdf(String html, String destination) throws IOException {
try (FileOutputStream output = new FileOutputStream(destination)) {
HtmlConverter.convertToPdf(html, output);
}
}
public static void main(String[] args) throws IOException {
String html = "<!doctype html><html><body><h1>Invoice</h1><p>Paid</p></body></html>";
createPdf(html, "invoice.pdf");
}
}
The output path is created or replaced by the stream in this example. In a service, write to a controlled temporary location or directly to your application’s output stream rather than accepting an arbitrary client path.
Resolve relative images and stylesheets with a base URI
An HTML string containing css/site.css or images/logo.png has no reliable document location unless you provide one. iText documents ConverterProperties.setBaseUri for this case.
import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;
public final class HtmlWithAssets {
public static void createPdf(String baseUri, String html, String destination)
throws IOException {
ConverterProperties properties = new ConverterProperties();
properties.setBaseUri(baseUri);
try (FileOutputStream output = new FileOutputStream(destination)) {
HtmlConverter.convertToPdf(html, output, properties);
}
}
public static void main(String[] args) throws IOException {
String html = "<html><head><link rel='stylesheet' href='css/site.css'></head>"
+ "<body><img src='images/logo.png'></body></html>";
createPdf("file:///opt/my-app/templates/", html, "report.pdf");
}
}
Use a directory URI ending in / when the document refers to sibling files. For a page downloaded from a URL, use that page’s URL as the base so relative references resolve against the same location. Make sure the renderer process can actually read every referenced file or URL.
Fetch a URL in Java, then render the returned HTML
When you need to convert a URL rather than a string you already have, fetch the response explicitly and pass the response body to iText. This keeps HTTP behavior, status checking and timeouts in your code instead of hiding them in a renderer.
import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
public final class UrlToPdf {
public static void main(String[] args) throws Exception {
String pageUrl = "https://example.com/";
HttpClient client = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(20))
.followRedirects(HttpClient.Redirect.NORMAL)
.build();
HttpRequest request = HttpRequest.newBuilder(URI.create(pageUrl))
.timeout(Duration.ofSeconds(60))
.header("User-Agent", "MyPdfService/1.0")
.GET()
.build();
HttpResponse<String> response = client.send(
request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() < 200 || response.statusCode() >= 300) {
throw new IllegalStateException("HTTP status: " + response.statusCode());
}
ConverterProperties properties = new ConverterProperties();
properties.setBaseUri(pageUrl);
try (FileOutputStream output = new FileOutputStream("page.pdf")) {
HtmlConverter.convertToPdf(response.body(), output, properties);
}
}
}
This fetches the initial HTML only. If the visible content is inserted later by JavaScript, the response body will not contain that content; use the browser-oriented options described below.
Rank #2
OpenHTMLToPDF for a pure-Java pipeline
OpenHTMLToPDF describes itself as a pure-Java library that renders a reasonable subset of well-formed XML/XHTML (and some HTML5) with CSS 2.1 and later standards, outputting PDF or images. Its documentation explicitly says it is not a web browser: it does not run JavaScript and does not implement many modern standards, including flex and grid.
That scope makes it a good fit for server-side templates that you can keep close to XHTML and CSS 2.1. The project also documents PDF/A and accessible-PDF support, SVG and MathML modules, font fallback, and a renderer intended to be faster for very large documents. The speed statement is qualitative; no controlled benchmark figure or universal multiplier is established here.
Free tools Windows power users keep installed
One-click scans. No signup required.
Basic OpenHTMLToPDF conversion
import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
import java.io.FileOutputStream;
public final class OpenHtmlToPdfExample {
public static void main(String[] args) throws Exception {
String html = "<html><body><h1>Report</h1><p>Generated in Java.</p></body></html>";
String baseUri = "file:///opt/my-app/templates/";
try (FileOutputStream output = new FileOutputStream("report.pdf")) {
PdfRendererBuilder builder = new PdfRendererBuilder();
builder.withHtmlContent(html, baseUri);
builder.toStream(output);
builder.run();
}
}
}
Keep the input well formed, provide a base URI for images and stylesheets, and register or package the fonts your layout depends on. Confirm the builder API and module names against the OpenHTMLToPDF version selected for production because artifacts and optional modules can change.
OpenHTMLToPDF vs iText: practical choice
- Choose iText when its HTML/CSS conversion features, output integration and support model fit your project and its licensing terms are acceptable.
- Choose OpenHTMLToPDF when you want a pure-Java, LGPL-compatible renderer for controlled XHTML/CSS and can avoid JavaScript, flex and grid.
- Do not choose either solely because the input is a public URL. A live site can depend on client-side rendering, authenticated requests, responsive breakpoints and browser APIs that a document renderer does not reproduce.
Flying Saucer when browser behavior matters
Flying Saucer documents pure-Java XML/XHTML and CSS 2.1 rendering with PDF and image output. Its repository lists org.xhtmlrenderer:flying-saucer-pdf and org.xhtmlrenderer:flying-saucer-chrome-pdf. The latter delegates PDF generation to chrome-headless-shell, making it the browser-oriented path identified in the project material.
Use the conventional PDF artifact for controlled XHTML/CSS. Use the Chrome-backed artifact when you need browser layout or JavaScript behavior, and treat the browser binary as an operational dependency: pin the version, make it available in every deployment environment, set process timeouts, and test the same page in CI and production.
Java compatibility is version-specific. The project states that versions from 9.5.0 require Java 11 or later, 9.6.0 require Java 17 or later, and 10.0.0 require Java 21 or later. Verify the exact artifact and baseline before choosing your build image.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Where Apache PDFBox fits
Apache describes PDFBox as an open-source Java tool for working with PDF documents. It can create, manipulate, render and post-process PDFs, and OpenHTMLToPDF uses it as its underlying PDF library. PDFBox alone should not be presented as a browser-grade HTML/CSS/JavaScript converter.
A common architecture is therefore HTML renderer first, PDFBox second: render with iText, OpenHTMLToPDF or a browser-backed tool, then use PDFBox for operations such as merging, splitting, extracting or inspecting pages. Keeping those responsibilities separate makes failures easier to diagnose.
Handle the parts that usually break
Relative assets
Missing logos, CSS and web fonts are usually URI-resolution failures. Set an explicit base URI, use stable file or HTTP locations, and test an HTML fixture with one relative stylesheet and one relative image before adding complex templates.
Malformed HTML
OpenHTMLToPDF is designed around well-formed XML/XHTML. Close elements, quote attributes, avoid invalid nesting and make the document encoding explicit. Browser error recovery can hide malformed markup; a PDF renderer may stop, omit content or produce unexpected pagination instead.
Recommended Free Tools
JavaScript and delayed content
A non-browser renderer cannot be expected to wait for a single-page application, execute a chart library or click a consent dialog. If the final content is created in the browser, render it with a Chromium-backed path or capture a fully materialized HTML snapshot before passing it to a document renderer.
Fonts and pagination
PDF line breaks depend on the exact font files and metrics available to the renderer. Package the fonts required by your design, test long words and tables, and compare page breaks after every font or CSS change. A PDF generated by one engine is not guaranteed to paginate identically to a browser screenshot or to another Java renderer.
Rank #4
Authenticated or private resources
Fetching the page HTML successfully does not guarantee that its images, stylesheets or fonts are available to the renderer. Arrange access for every resource, or download and reference the assets from a controlled base directory. Do not put secrets in publicly reachable URLs.
Make the conversion reliable in production
Validate the source before rendering
- Check the HTTP status and content type when you fetch a URL.
- Record the final URL after redirects so the base URI is correct.
- Reject unexpectedly tiny or empty HTML responses before spending time rendering.
- Use a representative fixture containing images, long paragraphs, tables, lists and page breaks.
Control time and memory
Set connection and response timeouts in your HTTP client. For browser-backed rendering, also set a maximum browser-process lifetime. Stream the PDF to an output stream where possible, and measure memory with the largest document you expect rather than relying on a small development sample.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Test output, not just exceptions
A successful method call can still produce a PDF with missing assets or blank pages. In automated tests, open the resulting file, verify that it has pages, check expected text or metadata, and inspect a rendered sample for layout regressions.
Review licensing and support
iText pdfHTML licensing depends on the version and deployment model, so confirm the applicable terms before shipping. OpenHTMLToPDF is described as LGPL-compatible. Flying Saucer and PDFBox have their own project licenses and version requirements; keep a record of the exact artifacts selected for your build.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
ClassNotFoundException for an iText or OpenHTMLToPDF class |
The renderer module is missing or versions are mixed. | Use the complete module set for one library version and rebuild from a clean dependency cache. |
| Images or CSS disappear | No base URI, an incorrect directory URI, or inaccessible resources. | Set ConverterProperties.setBaseUri or withHtmlContent(html, baseUri), then verify each asset from the renderer’s environment. |
| Flexbox, grid or JavaScript content is missing | OpenHTMLToPDF is not a browser and does not implement those features. | Simplify the template to supported XHTML/CSS, pre-render the dynamic content, or use the Chrome-backed Flying Saucer path. |
| Flying Saucer fails during startup with a Java version error | The selected artifact requires a newer Java baseline. | Align the JDK with the artifact line: 9.5.0 needs Java 11+, 9.6.0 Java 17+, and 10.0.0 Java 21+, according to the project documentation. |
| PDFBox code creates a PDF but the web page is not rendered | PDFBox is being used as if it were an HTML renderer. | Render HTML with iText, OpenHTMLToPDF or a browser-backed tool first; reserve PDFBox for PDF operations. |
| Output is blank or has far fewer pages than expected | The fetched response is an error page, content is injected by JavaScript, or malformed markup stopped layout. | Log status and response length, inspect the saved HTML, validate markup, and choose a browser-capable renderer for client-side content. |
| Fonts change line breaks between environments | Different font files or fallback fonts are being used. | Package the same fonts in every environment and include font-focused regression tests. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
Use the documented request shape below; this supplied example writes a WebP file. For a PDF response, use the PDF options documented at ScreenshotNeo’s API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to try it.
Best Value
FAQ
Will two Java renderers produce identical PDFs?
No. CSS support, font metrics, pagination and resource handling differ between engines. Treat a renderer switch as a layout change and run visual and text-based regression tests.
Can I use the same pipeline for a static template and a live single-page app?
Usually not without an extra browser stage. Static templates can stay in a Java renderer; a single-page app must first be rendered in a browser-capable environment or captured after its client-side content is ready.
Is a successful HTTP response enough to prove the PDF is correct?
No. A 200 response can contain an error page or shell HTML with no rendered content. Save and inspect the source, verify page count and expected text, and review representative pages.
Frequently Asked Questions
Will two Java renderers produce identical PDFs?
No. CSS support, font metrics, pagination and resource handling differ between engines. Treat a renderer switch as a layout change and run visual and text-based regression tests.
Can I use the same pipeline for a static template and a live single-page app?
Usually not without an extra browser stage. Static templates can stay in a Java renderer; a single-page app must first be rendered in a browser-capable environment or captured after its client-side content is ready.
Is a successful HTTP response enough to prove the PDF is correct?
No. A 200 response can contain an error page or shell HTML with no rendered content. Save and inspect the source, verify page count and expected text, and review representative pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

