Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The most direct Java implementation uses iText pdfHTML: create a java.net.URL, open its stream, and pass that InputStream to HtmlConverter.convertToPdf. The converter machine must reach the URL, and it may also need to download stylesheets, images, fonts, and other referenced resources. This produces a PDF from the HTML source; it does not guarantee pixel-for-pixel browser rendering, especially for JavaScript-heavy pages.

What you need before converting

  • A Java runtime compatible with the renderer version you select.
  • Network access from the Java process to the target URL and its linked assets.
  • Permission to retrieve and reproduce the page and its content.
  • A decision about the renderer’s HTML/CSS support and license before shipping.

Test with the exact pages you will process. A URL stream contains the fetched HTML bytes, but it does not prove that client-side scripts, authenticated state, lazy content, custom fonts, or every relative resource will be reproduced.

Convert a URL with iText pdfHTML

iText’s documented URL approach is to use URL.openStream() and feed the stream to HtmlConverter.convertToPdf. The following class is a complete command-line example. Add iText Core and pdfHTML according to the vendor’s current installation instructions.

import com.itextpdf.html2pdf.HtmlConverter;
import java.io.InputStream;
import java.io.OutputStream;
import java.net.URI;
import java.net.URL;
import java.nio.file.Files;
import java.nio.file.Path;

public class UrlToPdf {
    public static void main(String[] args) throws Exception {
        if (args.length != 2) {
            System.err.println("Usage: java UrlToPdf <url> <output.pdf>");
            System.exit(2);
        }

        URI uri = URI.create(args[0]);
        if (!"http".equalsIgnoreCase(uri.getScheme()) &&
            !"https".equalsIgnoreCase(uri.getScheme())) {
            throw new IllegalArgumentException("Only HTTP and HTTPS URLs are allowed");
        }

        URL url = uri.toURL();
        try (InputStream html = url.openStream();
             OutputStream pdf = Files.newOutputStream(Path.of(args[1]))) {
            HtmlConverter.convertToPdf(html, pdf);
        }
    }
}

Compile and run it with the iText dependencies on the class path:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
java UrlToPdf https://example.com example.pdf

For production code, prefer an HTTP client with explicit connect and read timeouts, redirect policy, response-size limits, and an allowlist of permitted hosts. The small example uses openStream() because it mirrors the documented recipe, not because it is a complete network-hardening policy.

Relative images, stylesheets, and fonts

Relative URLs need a base URI. iText’s converter accepts ConverterProperties; set its base URI to the source page’s origin (or another controlled location) before conversion when your HTML refers to resources such as images/logo.png. Verify that those resources are reachable from the conversion host. Numerous pictures can materially increase download time.

import com.itextpdf.html2pdf.ConverterProperties;

ConverterProperties properties = new ConverterProperties();
properties.setBaseUri("https://example.com/");
try (InputStream html = URI.create("https://example.com/report").toURL().openStream();
     OutputStream pdf = Files.newOutputStream(Path.of("report.pdf"))) {
    HtmlConverter.convertToPdf(html, pdf, properties);
}

How closely will the PDF match a browser?

That depends on the page and renderer. A URL fetch does not execute a browser session in the same way as Chrome or Firefox. JavaScript-generated markup, interactive widgets, unsupported CSS, web fonts, cross-origin requests, login state, and content that appears only after user actions can be missing or laid out differently.

OpenHTMLtoPDF documents support for a reasonable subset of well-formed XML/XHTML and some HTML5, with CSS 2.1-era layout support. Its maintainers warn that you cannot simply send arbitrary modern HTML5 to the engine and expect a great result without adapting the content. Flying Saucer likewise targets well-formed XML/XHTML and CSS 2.1. Therefore, validate representative pages rather than selecting a library from its name alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to adapt the HTML

  • Serve a print-oriented template with valid, well-formed markup.
  • Replace client-side rendering with server-rendered content before conversion.
  • Use print CSS and explicit dimensions for headers, footers, tables, and page breaks.
  • Make images and fonts available at stable, absolute or correctly based URLs.

When a browser engine may be necessary

If the required output depends on extensive JavaScript, modern layout features, or a logged-in browser session, a browser-based capture workflow may fit better than a pure-Java renderer. The cited material does not establish a universal best renderer for dynamic pages, so compare output from your real targets.

Java renderer options and licensing

Option What is established Best fit to investigate License note
iText pdfHTML Accepts HTML as an InputStream; documented URL example uses URL.openStream(). Projects needing its PDF features and HTML/CSS support after validation. AGPL or commercial terms; iText says commercial use requires a commercial license for iText Core and pdfHTML.
OpenHTMLtoPDF Pure Java; reasonable XHTML and some HTML5; CSS 2.1-oriented support. Content you control and can author or adapt to its supported subset. LGPL 2.1 or later, according to the project.
Flying Saucer Pure Java renderer for well-formed XML/XHTML and CSS 2.1. XHTML and CSS 2.1 documents. Check the version’s license and dependencies before distribution.
Apache PDFBox PDF creation, manipulation, and text extraction. Post-processing or programmatic PDF work. Apache License 2.0.

PDFBox alone should not be presented as a turnkey HTML renderer. License summaries are not legal advice: review the current license text and how your application is distributed or offered as a service.

Alternative Java pattern: fetch, then convert

Separating retrieval from conversion gives you control over timeouts, authentication, caching, validation, and logging. Fetch the HTML with your approved HTTP client, save it or keep it in memory, then pass a ByteArrayInputStream to the converter. Set a base URI so relative assets resolve. Never fetch arbitrary user-supplied URLs without SSRF protections: block loopback, private, link-local, and metadata-network addresses; restrict schemes and redirects; cap response size; and avoid forwarding internal credentials.

Performance and reliability checklist

  • Set connect, read, total-job, and resource-download timeouts.
  • Limit HTML, image, font, and stylesheet sizes.
  • Expect pages with many images to take longer because assets must be downloaded.
  • Cache immutable source pages and assets where your compliance policy permits.
  • Record the source URL, timestamp, renderer version, and conversion outcome for reproducibility.
  • Write to a temporary file and atomically rename it after successful conversion.
  • Retry transient network failures with bounded exponential backoff, but do not retry malformed HTML indefinitely.
  • Compare generated PDFs visually and textually against acceptance samples after dependency upgrades.

Common failures and fixes

Connection refused, timeout, or unknown host

The conversion host cannot reach the URL or an asset. Check DNS, firewall and proxy settings, then test every referenced resource from the same network. Increase timeouts only after confirming the page is trusted and bounded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PDF is blank or missing dynamic content

The meaningful markup may be generated by JavaScript, require a login, or depend on an unsupported feature. Use a server-rendered/print template, provide authenticated content through a controlled fetch, or evaluate a browser-based workflow.

Images or CSS are missing

Relative paths may have no base URI, resources may be blocked, or the server may return an unexpected content type. Set ConverterProperties.setBaseUri, inspect response status and content type, and use absolute URLs where appropriate.

Malformed-document or layout errors

Validate and simplify the HTML. XHTML-style well-formedness, explicit character encoding, supported CSS, and print-specific markup improve predictable output.

License concern

Do not assume that a library being downloadable makes your deployment compliant. Review AGPL obligations, commercial terms, LGPL conditions, and your distribution or hosted-service model with qualified counsel.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a dependable screenshot or PDF of a live URL rather than a Java-rendered document, ScreenshotNeo provides a one-request API and an MCP server for AI agents. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be switched off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

For a screenshot, use the documented API parameters:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for PDF options, full-page capture, CSS selectors, JavaScript and click actions, waiting rules, blocking controls, headers and cookies, device presets, custom viewports, retina scale, transparent backgrounds, resizing, caching TTL, signed links, asynchronous webhooks, bulk capture, usage, and the OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I convert a page that requires a username and password?

Only if your retrieval and rendering design supplies authenticated access safely. The basic URL.openStream example does not demonstrate authentication, cookie management, or session handling.

Should I use HTML-to-PDF or a screenshot API for invoices?

Use an HTML-to-PDF renderer when you need selectable, paginated document output you control. Use a capture service when reproducing a live page state is more important than document semantics, and validate privacy and retention requirements.

Does URL.openStream wait for JavaScript to finish?

It opens a network stream for the URL response; it is not a browser automation wait mechanism. Client-side rendering therefore may not appear in the fetched HTML.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.