The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Use Puppeteer to render each webpage as a PDF, then use pdf-lib to copy those PDFs’ pages into one document. Puppeteer’s page.pdf() prints the current page; it does not merge documents. The example below processes URLs sequentially so the final page order matches the URL list.
What the workflow does
The job has two distinct stages: a browser renders each URL, and a PDF library assembles the rendered page sets. Puppeteer documents Page.pdf() as generating a PDF with the print CSS media type (Puppeteer Page.pdf() API). pdf-lib provides the page-copying and document-creation operations used for the merge (pdf-lib documentation).
- Launch Puppeteer and navigate to each URL.
- Check the navigation response and wait for the content your page needs.
- Render the page to PDF bytes.
- Load each rendered PDF with pdf-lib and append its pages, in URL order, to a destination document.
- Save the combined bytes and close the browser.
This produces a single PDF whose pages are the print-rendered output of each input webpage. It does not combine webpages into one continuous, reflowed page: each source PDF contributes its own pages.
Install the libraries
Start a Node.js project and install Puppeteer and pdf-lib:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Convert your PDF files into Word, Excel & Co. the easy way
- Convert scanned documents thanks to our new 2022 OCR technology
- Adjustable conversion settings
- No subscription! Lifetime license!
- Compatible with Windows 11, 10, 8.1, 7 - Internet connection required
npm init -y
npm install puppeteer pdf-lib
Use ES modules by adding "type": "module" to the top-level of package.json, or save the script below with an .mjs extension. Puppeteer’s package generally installs a compatible browser as part of its installation; in restricted or preconfigured environments, consult the installation guidance for your selected Puppeteer version and browser setup.
Runnable example: render and merge URLs in order
Save this as combine-webpages.mjs. It accepts one or more URLs as command-line arguments and writes combined.pdf to the current directory.
import puppeteer from 'puppeteer';
import { PDFDocument } from 'pdf-lib';
import { writeFile } from 'node:fs/promises';
const urls = process.argv.slice(2);
if (urls.length === 0) {
console.error('Usage: node combine-webpages.mjs <url1> <url2> ...');
process.exit(1);
}
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
const renderedPdfs = [];
for (const url of urls) {
console.log(`Rendering ${url}`);
const response = await page.goto(url, { waitUntil: 'networkidle2' });
if (!response) {
throw new Error(`No main-resource response received for ${url}`);
}
if (!response.ok()) {
throw new Error(`HTTP ${response.status()} ${response.statusText()} for ${url}`);
}
const pdfBytes = await page.pdf({
format: 'A4',
printBackground: true,
});
renderedPdfs.push({ url, bytes: pdfBytes });
}
const combined = await PDFDocument.create();
for (const { url, bytes } of renderedPdfs) {
const source = await PDFDocument.load(bytes);
const copiedPages = await combined.copyPages(
source,
source.getPageIndices(),
);
for (const pdfPage of copiedPages) {
combined.addPage(pdfPage);
}
console.log(`Added ${copiedPages.length} page(s) from ${url}`);
}
const outputBytes = await combined.save();
await writeFile('combined.pdf', outputBytes);
console.log(`Wrote combined.pdf (${outputBytes.length} bytes)`);
} finally {
await browser.close();
}
Run it with URLs in the order you want their pages to appear:
node combine-webpages.mjs https://example.com/page-one https://example.com/page-two
For each URL, every page in that webpage’s generated PDF is appended before the next URL’s pages. The resulting document therefore follows the argument order, with each source’s internal page order preserved. The script checks the main-resource HTTP response before printing; it treats a missing response or non-success status as an error instead of silently making the next PDF.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
- Convert over 50 document file formats.
- Preview your files from Doxillion before converting them.
- Use batch conversion to convert thousands of files at once.
- Enjoy an easy-to-use, intuitive interface with a Drag and Drop file option.
- Burn your converted or original files directly to disc.
Choose how each webpage is rendered
Print styling or screen styling
Puppeteer generates PDFs using print CSS by default. That is appropriate when the site provides a print stylesheet, but print rules may remove navigation, change colors, or rearrange content. To render with screen CSS instead, call await page.emulateMediaType('screen') after navigation and before page.pdf(). This is a deliberate styling choice: screen output may be more visually faithful to the browser viewport, while print output follows the website’s print layout. See the Puppeteer PDF generation guide.
Paper size, orientation, margins and backgrounds
The example chooses A4 paper and enables background printing. Puppeteer’s PDF options also cover width and height, landscape orientation, margins, scale, page ranges, timeout, and preference for CSS-defined page size (PDFOptions API). Select these to suit the pages and intended use. If you rely on CSS @page sizing, consider preferCSSPageSize: true; otherwise the selected format or dimensions govern the paper size. Current documented defaults include letter paper, no margins, background printing off, and a 30,000 ms PDF timeout; verify defaults against the version installed in your project rather than depending on an implicit setting.
Fonts and page readiness
Puppeteer’s PDF options document waitForFonts as true by default, so PDF generation waits for fonts to load. That does not mean every page has finished rendering its content. The official PDF guide uses waitUntil: 'networkidle2' in its navigation example, but network quiet is not a universal readiness signal: sites may poll continuously or insert content later. The navigation method returns the main-resource response, which lets you check HTTP status as shown above (Page.goto() API).
For a page with a known content marker, wait for that marker after navigation before printing:
Rank #3
- EDIT text, images & designs in PDF documents. ORGANIZE PDFs. Convert PDFs to Word, Excel & ePub.
- READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.
- CREATE, COMBINE, SCAN and COMPRESS PDFs
- FILL forms & Digitally Sign PDFs. PROTECT and Encrypt PDFs
- LIFETIME License for 1 Windows PC or Laptop. 5GB MobiDrive Cloud Storage Included.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#article-content', { timeout: 15000 });
const pdfBytes = await page.pdf({ format: 'A4', printBackground: true });
Replace #article-content with a selector that means the content you need is ready. If the page relies on an application-specific event or delayed data, wait on that condition instead. A fixed delay can be useful for a known animation or brief delayed render, but it is less reliable than waiting for a meaningful selector or condition.
Control page order and placement
The simplest ordering rule is to process the URLs sequentially and call addPage() for each copied page as it comes. If you need pages inserted at a particular position instead of appended, pdf-lib also documents insertPage() (PDFDocument API). Keep an explicit mapping between each URL and its source PDF if your workflow later adds sorting, retries, or parallel rendering.
Sequential rendering or concurrency?
The example reuses one Puppeteer page sequentially. This makes it straightforward to associate each navigation with its output and preserve requested order, while avoiding the extra simultaneous browser pages and resource use that parallel work entails. Puppeteer supports multiple pages in a browser (Page class), but the documentation does not establish a universal performance winner. If you introduce concurrency, limit the number of active pages to fit your environment, retain each result’s original URL index, and merge in that index order rather than whichever task finishes first.
Practical constraints and output checks
- Access: The browser must be able to reach the page. Authentication, paywalls, bot protections, or network restrictions require application-specific access handling; Puppeteer and pdf-lib do not guarantee arbitrary sites can be fetched or printed.
- Page completeness: Review pages with delayed, lazy-loaded, or interactive content. A successful navigation response only confirms the main resource response; it does not prove that every desired element appeared.
- Print layout: Inspect page breaks, clipped content, backgrounds, and elements hidden by print CSS. Test a representative page before processing a large batch.
- PDF fidelity: This pattern is for ordinary rendered-page PDFs. The cited pdf-lib documentation establishes page copying and merging, but does not establish preservation of every advanced PDF feature such as forms, outlines, digital signatures, or tagged accessibility metadata. Verify current library behavior against representative files if those features matter.
- Memory and size: The example retains every rendered PDF byte array before assembling the destination. Large pages or many URLs can consume substantial memory. For larger workloads, process in bounded batches or write intermediate PDFs to storage and adapt the merge flow, while keeping track of ordering and cleanup.
Troubleshooting
The script reports no response or an HTTP error
page.goto() did not provide a main-resource response, or the response status was not successful. Check that the URL is correct and reachable from the machine running Chromium, and determine whether the site redirects, requires authentication, blocks automation, or returns an error page. Do not remove the status check unless your workflow intentionally wants to include error pages in the output.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #4
- Edit PDFs with Ease. Modify text, images, and layouts directly within your PDF documents.
- Convert & Organize. Export PDFs to Word, Excel, or ePub, and organize files with ease.
- Read & Annotate. Enjoy intuitive reading modes and powerful tools to comment, highlight, and mark up PDFs.
- Create & Manage PDFs. Create new PDFs, combine multiple files, scan documents, and compress for easy sharing.
- Fill & Sign Forms. Complete forms and digitally sign documents with secure e-signature tools.
The PDF is blank or missing late-loading content
The page may have navigated before its application rendered the needed content. Replace a broad network-idle wait with a meaningful selector or app-specific readiness condition, then print. Confirm the selector exists on every target page and increase its timeout only when the site’s expected behavior warrants it.
The output looks different from the browser
PDF generation uses print media unless you explicitly emulate screen media. Check the site’s print CSS and test page.emulateMediaType('screen') if the desired output is the on-screen layout. Also check paper size, margins, background printing, and page-break rules.
Fonts or images are absent
Wait for the relevant content and assets before calling page.pdf(). Font waiting is enabled by default in the documented options, but external font requests can still fail or be blocked. For lazy images, scroll or otherwise trigger their loading before printing, then wait for the images your output requires.
The browser fails to launch
Check that the Puppeteer package and its expected browser are installed for the environment, and that the runtime permits Chromium to start. Containers and locked-down hosts may need environment-specific browser dependencies or configuration; use the installation documentation corresponding to your Puppeteer version.
Best Value
- ALL-IN-ONE SOLUTION – read, edit, convert, merge and protect your PDF files
- MAXIMUM FUNCIONALITY – create interactive forms, compare PDFs, bates numbering, find and replace text or colors, convert documents, OCR engine, comment, highlight, fill out and print forms, document protection and others
- EASY TO INSTALL AND USE – well-structured user-interface, in-program instructions, free tech support whenever you need it
- GREAT VALUE FOR MONEY - why spend a fortune if you can have maximum functionality at a reasonable price - this also fits the requirements of companies very well
The merged PDF is out of order
In the sequential example, output follows command-line URL order. If you add concurrent rendering or retries, save each result with its original index and sort by that index before calling copyPages() and addPage(). Completion order is not necessarily input order.
Or skip the browser setup
If you want a screenshot of a page rather than a multi-URL merged PDF, ScreenshotNeo offers a one-request screenshot API and an MCP server. Its API returns an image or PDF for one URL; the Puppeteer-plus-pdf-lib workflow above remains the way to assemble PDFs from several URLs into one document.
Example request (see the ScreenshotNeo API documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
- Cookie banners and consent prompts, newsletter popups, and chat widgets are handled before the shot; those cleanup steps can be turned off.
- Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; responses identify the page verdict and billing status in headers.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
Frequently Asked Questions
Does Puppeteer merge PDFs by itself?
No. Puppeteer renders a page to PDF; use a PDF library such as pdf-lib to combine the generated page sets.
Can I combine URLs into one continuous webpage before printing?
The example creates one PDF by appending each webpage’s rendered PDF pages. It does not create a single reflowed webpage from the source URLs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

