Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract an embedded PDF with Puppeteer, first find the PDF resource URL in the page’s frames or embedded-element markup. If the site loads the document dynamically, monitor network requests instead. Then retrieve the resource itself and check that the response contains a PDF. page.pdf() is different: it generates a PDF of the current page using print CSS; it does not download a PDF already embedded in that page.

What “extract an embedded PDF” means

An embedded PDF is a document resource displayed inside a web page, commonly through an <iframe>, <embed>, or <object> element, or through a viewer that loads the document separately. Extraction means identifying and retrieving that original resource. The page’s HTML may expose its URL directly, or scripts may load it only after navigation or an interaction.

By contrast, Puppeteer’s page.pdf() prints the currently rendered page to a new PDF, using print CSS by default. That can be useful when you want a printable copy of the web page, but it is not a way to obtain the original file embedded in the page.

Start with the page’s frames and embedded elements

Inspect the top-level page and its attached frames, then look for URL-bearing attributes on likely embed elements. A candidate URL might be the PDF itself, or it might point to an HTML viewer. Treat it as a lead until you confirm what it returns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Epson Workforce ES-50 Compact & Lightweight Mobile Document Scanner
  • PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
  • QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
  • VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
  • INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
  • EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0

Runnable Node.js discovery script

Install Puppeteer in a Node.js project with npm install puppeteer. Save this as inspect-pdf.js, replace the example page URL, and run node inspect-pdf.js. It prints each frame URL and any iframe, embed, or object attributes found in that frame.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com/page-with-pdf', {
      waitUntil: 'domcontentloaded',
      timeout: 60000,
    });

    for (const frame of page.frames()) {
      console.log('nFrame:', frame.url());
      try {
        const elements = await frame.evaluate(() =>
          Array.from(document.querySelectorAll('iframe, embed, object')).map(el => ({
            tag: el.tagName.toLowerCase(),
            src: el.getAttribute('src'),
            data: el.getAttribute('data'),
            type: el.getAttribute('type'),
          }))
        );
        for (const element of elements) console.log(element);
      } catch (error) {
        console.log('Could not inspect this frame:', error.message);
      }
    }
  } finally {
    await browser.close();
  }
})();

The attributes are deliberately printed as found. A relative URL needs to be resolved against the frame or page that contains it, and an empty or absent attribute does not prove that no PDF is loaded. A cross-origin frame may also be inaccessible to page JavaScript; continue with request observation rather than assuming the document is absent.

How to interpret the results

  • An iframe src ending in a PDF path is a useful candidate, but the extension alone does not confirm the response is a PDF.
  • A URL that points to a viewer page may require another inspection step: the viewer’s own frame or its network activity may reveal the document URL.
  • For an object, check data as well as type; for an embed, check src.
  • If scripts replace or populate the element after initial navigation, inspect after the relevant page activity or user action.

Monitor requests when markup does not reveal the file

Attach request listeners before navigation so you do not miss early activity. Puppeteer exposes request, requestfinished, and requestfailed events. A finished request has completed downloading its response body, but that does not mean the HTTP status was successful: a 404 or 503 can still finish at the transport level.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    page.on('request', request => {
      console.log('REQUEST', request.method(), request.url());
    });
    page.on('requestfinished', async request => {
      const response = request.response();
      console.log(
        'FINISHED',
        response ? response.status() : 'no response',
        request.url()
      );
    });
    page.on('requestfailed', request => {
      console.log('FAILED', request.url(), request.failure()?.errorText);
    });

    await page.goto('https://example.com/page-with-pdf', {
      waitUntil: 'domcontentloaded',
      timeout: 60000,
    });
    // If the viewer loads the PDF only after a click, perform that site's
    // required interaction here, after the listeners have been attached.
    await page.waitForTimeout(5000);
  } finally {
    await browser.close();
  }
})();

Look for requests initiated by the viewer around the time it displays the document. A likely file request can use a URL without a .pdf suffix, so consider response status and content type as clues rather than relying only on the URL name. Puppeteer’s request lifecycle helps identify candidate resources; it does not guarantee a universal way to recognize every viewer’s internal format.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Brother DS-640 Compact Mobile Document Scanner, (Model: DS640)
  • FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
  • ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
  • READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
  • WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
  • OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)

Retrieve the original resource and verify it

Once you have a candidate resource URL, retrieve that URL rather than calling page.pdf(). The site may require the same cookies, authorization, or other session context as the browser page. There is no single authenticated-download recipe that works for every site: how to carry over session state depends on the target site and your access rights.

  1. Record the candidate URL and the frame or request that exposed it.
  2. Check the HTTP response status. Do not treat a completed request as a successful document response.
  3. Save the response body as bytes, not as decoded page text.
  4. Verify that the saved response is actually a PDF before passing it to another tool. A viewer HTML page, an access-denied response, or an error document can be returned from a URL that looked promising.

Do not assume that a filename ending in .pdf, a viewer label, or a successful browser request proves the downloaded bytes are a valid PDF. The applicable verification method can depend on the target’s response and the downstream PDF software; the available Puppeteer guidance does not establish one validation algorithm for all embedded viewers.

Choose the discovery path that fits the page

Method Best first use What it can reveal Important limit
Frame and DOM inspection Start here when the page has an ordinary embed or iframe. Element attributes and frame URLs that may identify the resource or viewer. The URL may be hidden from initial markup or inaccessible to page script.
Request monitoring Use when the viewer loads content dynamically or after interaction. Requests made while scripts run or the user operates the viewer. A finished request can still have an HTTP error status; inspect the response.

Neither method is guaranteed for every website. Some pages expose the document directly; others use a viewer, session-dependent access, or dynamic loading. Use DOM inspection as the simpler first check and request observation as the fallback.

Browser-mode caveat for direct PDF navigation

Puppeteer’s page.goto() documentation warns that headless shell mode does not support navigation to a PDF document. This warning is specific to headless shell; it should not be generalized to every Puppeteer mode or browser configuration. If navigating directly to a PDF fails, first check which browser mode you launched. Discovering the PDF URL from an HTML host page and navigating directly to that URL are separate steps.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Epson Workforce ES-400 II High-Speed Color Duplex Desktop Document Scanner
  • FAST DOCUMENT SCANNING — Document scanner with feeder allows you to speed through stacks with a 50-sheet Auto Document Feeder (ADF); Efficient office scanner to help you scan more productively
  • INTUITIVE, HIGH-SPEED SOFTWARE — Quickly scan with this desktop document scanner; Epson ScanSmart Software lets you easily preview scans, email files, upload to the cloud, and more; Plus, automatic file naming saves even more time
  • SEAMLESS INTEGRATION — Easily incorporate your data into most document management software with the included TWAIN driver; Office document scanner integrates seamlessly with business workflows
  • EASY SHARING — Duplex scanner allows you to scan straight to email or popular cloud storage2 services like Dropbox, Evernote, Google Drive, and OneDrive for simple storage and sharing
  • SIMPLE FILE MANAGEMENT — Scanner allows the creation of searchable PDFs with Optical Character Recognition (OCR) and convert scans to editable Word or Excel files effortlessly; Designed for home and office document scanning

Troubleshooting

No PDF URL appears in the DOM

The viewer may insert its resource after scripts run, or the PDF may be loaded in another frame. Print page.frames() URLs and inspect each accessible frame. If that does not expose a candidate, attach request listeners before navigation and repeat any interaction that causes the viewer to load.

The candidate URL opens a viewer instead of a file

You may have found the viewer page, not the PDF resource. Inspect the viewer’s frames and watch its requests while it loads the document. Do not save the viewer’s HTML response and call it the extracted PDF.

The request finished, but the saved file is unusable

Check the response status and inspect whether the body is a PDF rather than an error or access page. Request completion reports that downloading ended; it does not certify that the server returned a valid document.

The browser can display the PDF, but a separate download fails

The resource may depend on the session used to load the page. A separate retrieval must preserve whatever session context the site requires. The exact transfer method is site-specific; the documented Puppeteer inspection methods do not prescribe one universal authenticated-download procedure.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
  • Scanner type: Document
  • Connectivity technology: USB
  • With Auto Scan Mode, the scanner automatically detects what you're scanning
  • Digitize documents and images

page.goto() fails on the PDF URL

Check whether the browser is running in headless shell mode, which has a documented limitation for navigating to PDF documents. Do not assume that the same limitation applies to all Puppeteer modes.

The downloaded filename looks right, but its contents are not a PDF

Use the HTTP status and actual response content to assess the result. A URL path or filename is not proof of the returned format; viewer pages and error responses can be misleading.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost considerations

Request monitoring can produce substantial log output on pages with many assets. Filter or record only relevant requests once you know what the target viewer loads. A fixed delay can miss slow or interaction-triggered requests, so use it only as a simple demonstration; adapt the wait to the page’s actual behavior. Timeouts, site access controls, and session requirements can all affect whether you can obtain the resource. No general success rate or extraction time applies across websites.

Use this technique only for documents you are permitted to access and retrieve. Puppeteer can help inspect what the page loads, but it does not bypass a site’s access requirements or make a protected resource publicly downloadable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
ScanSnap iX2500 Wireless or USB High-Speed Document Scanner, Black
  • OUR MOST ADVANCED SCANSNAP. Large touchscreen, fast 45ppm double-sided scanning, 100-sheet document feeder, Wi-Fi and USB connectivity, automatic optimizations, and support for cloud services. Upgraded replacement for the discontinued iX1600
  • CUSTOMIZABLE. SHARABLE. Select personalized profiles from the touchscreen. Send to PC, Mac, mobile devices, and clouds. QUICK MENU lets you quickly scan-drag-drop to your favorite computer apps
  • STABLE WIRELESS OR USB CONNECTION. Built-in Wi-Fi 6 for the fastest and most secure scanning. Connect to smart devices or cloud services without a computer. USB-C connection also available
  • PHOTO AND DOCUMENT ORGANIZATION MADE EFFORTLESS. Easily manage, edit, and use scanned data from documents, receipts, photos, and business cards. Automatically optimize, name, and sort files
  • AVOIDS PAPER JAMS AND DAMAGE. Features a brake roller system to feed paper smoothly, a multi-feed sensor that detects pages stuck together, and skew detection to prevent paper damage and data loss

Or skip the browser setup

If your goal is to create a clean screenshot or a new PDF rendition of a web page—not retrieve the original PDF embedded in it—ScreenshotNeo offers a one-request capture API. This is a different job from extracting the original document bytes. The following cURL request captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page-with-pdf -o shot.webp

See the ScreenshotNeo documentation for request options, including PDF capture. Before capture, it can accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are not billed, and the response identifies the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, with no card required.

Frequently Asked Questions

Does Puppeteer’s page.pdf() download the PDF embedded in a page?

No. It creates a PDF of the current page using print CSS by default; extraction requires finding and retrieving the embedded resource.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Will this method work for every PDF viewer?

No universal method is established. Some pages expose a direct resource URL, while others load content dynamically or require site-specific session context.

Quick Recap

Bestseller No. 4
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Canon Canoscan Lide 300 Scanner (PDF, AUTOSCAN, Copy, Send)
Scanner type: Document; Connectivity technology: USB; With Auto Scan Mode, the scanner automatically detects what you're scanning
$75.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.