How do you convert HTML to PDF? Choose a browser print or automation workflow when the PDF must match a live, JavaScript-rendered page. Choose a document renderer such as WeasyPrint or Prince when predictable paged-media layout, server-side document features, or a non-browser deployment matters more. There is no universally best HTML-to-PDF library: the right choice depends on rendering fidelity, print layout control, runtime, security, and the PDF features you must deliver.
Choose a conversion method first
HTML-to-PDF systems fall into three practical groups. The table below is a decision aid, not a performance ranking; the cited projects document capabilities, but the available material does not establish comparative benchmarks.
| Method | Best fit | Important trade-off |
|---|---|---|
| Browser print flow | A person prints an already rendered page from a browser and wants the normal print preview or Save as PDF experience. | Least automation and repeatability; the result depends on the page, browser, user settings, and print CSS. |
| Puppeteer | JavaScript or Node.js services that need a real browser, JavaScript execution, and PDF bytes or a file. | Browser startup and page rendering add operational overhead. PDF output uses print CSS unless you deliberately select screen media. |
| Playwright | Teams already using Playwright for browser automation and needing a PDF buffer plus documented output options. | It has the same browser-rendering considerations as other automation tools, including print media defaults. |
| WeasyPrint | Python applications that need HTML/CSS-to-PDF rendering and document features without a full WebKit or Gecko browser. | It is its own rendering engine, so browser-specific layout or JavaScript behavior is not the target. CSS support should be checked against the feature matrix. |
| Prince | Publishing and reporting workflows that need detailed paged-media composition, headers, footers, numbering, and controlled page breaks. | It is a commercial renderer; licensing and deployment terms must be evaluated with YesLogic. |
Browser print: the simplest human workflow
When a person is looking at the page, use the browser’s print command and choose the PDF destination. This is appropriate for one-off exports, support instructions, and content that has already been reviewed visually. The page’s print stylesheet, selected paper size, margins, scale, headers, footers, background setting, and the browser’s print dialog all affect the result.
For a repeatable service, a manual print flow is usually the wrong abstraction. Move to Puppeteer or Playwright so navigation, authentication, waiting, media selection, and file handling are explicit and testable.
#1 Best Overall
Puppeteer: generate a PDF from a rendered page
Puppeteer’s documented sequence is to launch a browser, open a page, navigate to the content, call Page.pdf(), and close the browser. The API generates with print CSS by default. If the design is intended for the screen, call page.emulateMediaType('screen') before creating the PDF. The guide also notes that font loading is awaited by default.
See the Puppeteer PDF generation guide and the Page.pdf() API reference for the current option names and version-specific behavior.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/invoice/123', {
waitUntil: 'networkidle0'
});
// Omit this line when print CSS is the intended design.
await page.emulateMediaType('screen');
await page.pdf({
path: 'invoice.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
} finally {
await browser.close();
}
path writes a file; without it, the API can return PDF data for further processing. Keep the media decision intentional: print CSS often hides navigation and changes colors, while screen media may preserve a dashboard’s on-screen arrangement but produce poor page breaks. Print color treatment can also change what appears in the output, so verify background and color settings on representative pages.
When Puppeteer is a good fit
- The source depends on client-side JavaScript, web fonts, or browser APIs.
- You need the same Chromium rendering model used by an existing end-to-end test suite.
- You can operate browsers safely in your deployment environment and accept their startup and memory costs.
Playwright: browser automation with PDF options
Playwright’s page.pdf() returns a PDF buffer and uses print CSS by default. Its documented options include an output path and control over whether CSS page size is preferred. Use the Playwright Page API for the complete, version-specific option set.
Free tools Windows power users keep installed
One-click scans. No signup required.
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/report', {
waitUntil: 'networkidle'
});
// Playwright also defaults to print media. Select screen only when required.
await page.emulateMedia({ media: 'print' });
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
await writeFile('report.pdf', pdf);
} finally {
await browser.close();
}
Use Playwright when it is already part of your automation stack or when its browser-management model fits your deployment. The conversion questions remain the same as with Puppeteer: wait for the content that matters, decide between print and screen media, provide credentials safely, and make external resources deterministic.
WeasyPrint: a Python document renderer
WeasyPrint 70.0 is documented as a Python 3.10+ HTML/CSS rendering engine under the BSD license, not a wrapper around a full WebKit or Gecko browser. In the project’s words, “From a technical point of view, WeasyPrint is a visual rendering engine for HTML and CSS that can export to PDF.” It is a strong option for server-side documents whose layout is expressed in HTML and print-oriented CSS rather than browser JavaScript.
Rank #2
The API accepts HTML supplied as a string, file, file object, or URL. When relative images, stylesheets, or fonts are referenced from an HTML string, set a suitable base_url; otherwise those resources cannot be resolved reliably.
from weasyprint import HTML
html = '''
Invoice 123
Content loaded from the application.
'''
HTML(string=html, base_url='https://example.com/').write_pdf('invoice.pdf')
WeasyPrint uses print media by default. Page dimensions, margins, running content, and page breaks belong in CSS @page and related paged-media rules. Its documented document features include hyperlinks, bookmarks, attachments, and forms. PDF/A and PDF/UA generation is described as supported but not guaranteed valid; if conformance matters, validate the produced file with an appropriate checker rather than assuming the label is sufficient.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →WeasyPrint security boundary
Rendering user-modifiable HTML or CSS can expose your application to security problems. Treat markup, CSS, images, fonts, and every URL fetched during rendering as untrusted input. Restrict network access, validate or sanitize content, control the URL-fetching configuration, and isolate the renderer where the threat model requires it. Do not allow arbitrary user HTML to reach internal services through resource URLs.
Prince for advanced paged-media publishing
Prince is a commercial HTML/XML-to-PDF engine whose official guide covers HTML, Markdown, and XML input. Its styling documentation describes controls for page dimensions, headers, footers, page numbering, and page breaks. That makes it relevant for books, invoices, regulatory reports, and other publication workflows where page composition is a first-class requirement.
The Prince documentation also discusses server-side integration and the need for reliable, secure configuration. Confirm current licensing, supported platforms, and deployment terms with YesLogic before selecting it for a production system. The reviewed documentation does not establish a price or an affiliate program.
How to choose an HTML-to-PDF library
1. Decide whether browser fidelity is required
If the page is a single-page application, relies on JavaScript, or must look like a particular browser viewport, start with Puppeteer or Playwright. If the input is primarily semantic HTML and CSS for reports, WeasyPrint or Prince may be easier to operate and more predictable.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems2. Define the paper and pagination contract
Write down the paper size, margins, orientation, repeating headers and footers, numbering, acceptable widows and orphans, image scaling, and whether a table row may split. Browser APIs expose PDF options, but complex publication rules are often clearer in paged-media CSS or a dedicated renderer.
3. Check runtime and integration constraints
- Node.js teams commonly choose Puppeteer or Playwright because the browser and application code share a language.
- Python services can call WeasyPrint directly and keep conversion inside the application process.
- Prince may fit organizations that need a specialized publishing engine and can manage a commercial dependency.
4. Plan resource loading and authentication
Relative URLs, private images, CSS, fonts, and API data are frequent causes of incomplete PDFs. In a browser, establish the required session before navigation and wait for the page state that proves the data is ready. In WeasyPrint, provide base_url and configure URL fetching deliberately. For every renderer, decide whether external requests are allowed at all.
5. Treat security as part of conversion
HTML-to-PDF is a document-processing boundary. Sanitize untrusted markup, restrict outbound requests, avoid passing secrets into page content, and isolate browser or renderer processes. Log which input produced each PDF without logging credentials or sensitive document contents.
Reliability, performance, and output checks
- Wait for the real readiness signal. Network-idle alone may not mean that a chart, font, or client-side request has finished. Add an application-level marker or wait for a required element.
- Make assets deterministic. Pin fonts and stylesheets, use stable URLs, and avoid time-dependent data when PDFs are compared or archived.
- Reuse infrastructure carefully. Keeping a browser process alive can reduce launch overhead, but isolate pages and clear state between jobs. Set timeouts and terminate stuck work.
- Validate the artifact. Check that the file opens, has the expected page count, contains selectable text where required, includes links or bookmarks when promised, and does not contain clipped content or missing fonts.
- Measure your own workload. Page complexity, images, fonts, JavaScript, and concurrency dominate resource use. The cited project documentation describes features, not cross-tool speed or quality benchmarks.
Troubleshooting common failures
The PDF shows an old or empty page
The capture probably ran before client-side rendering completed. Wait for a specific application element or readiness signal, and verify that authentication cookies and API requests are present. A successful navigation event alone is not proof that the page is complete.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Screen styling appears instead of the print layout, or vice versa
Puppeteer and Playwright default to print CSS. Remove an unnecessary screen-media override when print rules are intended; call the documented media-emulation method before page.pdf() when screen styling is the requirement.
Images, CSS, or fonts are missing in WeasyPrint
Set base_url when constructing HTML from a string, use resolvable absolute URLs where appropriate, and inspect URL-fetching permissions. A restricted network policy can also block resources by design.
Rank #4
- Programming Language Lover Code Apparel. App or Web Design and Development Expert Funny Dress. Best Valentines Idea For Coding Lover. HTML Code or Meaning Costume
- Funny I Know HTML - How To Meet Ladies Computer Programmer Quotes
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
Pages break in the wrong places
Move pagination rules into print CSS or @page, avoid forcing large unbreakable blocks, and test long headings, tables, and images. Browser PDF options cannot compensate for a layout that has no viable break points.
The result fails a PDF/A or PDF/UA check
Do not infer conformance from successful PDF generation. WeasyPrint documents these variants as supported but not guaranteed valid. Run the output through the validator required by your compliance process and fix the reported metadata, tagging, font, or structural issues.
Recommended Free Tools
Conversion is slow or consumes too much memory
Reduce unnecessary images and JavaScript, limit concurrency, reuse browser infrastructure where safe, and apply job timeouts. Profile representative documents rather than relying on a generic benchmark.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is a website screenshot API that can return PNG, JPEG, WebP, or PDF from one GET request. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Use the documented request shape below; see the ScreenshotNeo documentation for output and option details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThe Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan, and yearly billing gives two months free. Sign up free for ScreenshotNeo and try the 1,000 monthly screenshots without a card.
Best Value
FAQ
Should I keep the source HTML after generating a PDF?
Usually yes. Store the template or content version, renderer version, CSS inputs, and relevant data identifiers so a document can be reproduced or audited later. Retain the PDF according to your document-retention policy.
Can one service use more than one renderer?
Yes. Many teams use a browser renderer for interactive pages and a document renderer for formal reports. Define a routing rule based on JavaScript dependence, pagination requirements, and conformance needs, then test each class of document separately.
What should an automated conversion test assert?
Assert file validity and page count first, then check text, links, required headings, and key visual regions. Include long-content, missing-resource, authenticated, and non-Latin-font cases so regressions are caught before release.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Frequently Asked Questions
Should I keep the source HTML after generating a PDF?
Usually yes. Store the template or content version, renderer version, CSS inputs, and relevant data identifiers so a document can be reproduced or audited later. Retain the PDF according to your document-retention policy.
Can one service use more than one renderer?
Yes. Many teams use a browser renderer for interactive pages and a document renderer for formal reports. Define a routing rule based on JavaScript dependence, pagination requirements, and conformance needs, then test each class of document separately.
What should an automated conversion test assert?
Assert file validity and page count first, then check text, links, required headings, and key visual regions. Include long-content, missing-resource, authenticated, and non-Latin-font cases so regressions are caught before release.

