Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If a Puppeteer PDF looks correct but copied text is reversed, scrambled, missing characters, or pasted as gibberish, do not assume a single “encoding bug.” PDF display and text extraction use different data: the page can draw the right glyphs while its Unicode mapping or reading order is unusable. Isolate the failure by comparing the source string, loaded font, print CSS, Puppeteer/Chromium version, and the PDF reader or extractor, then test a minimal document with a known font.

First determine what is actually broken

Save a short sample containing ordinary Latin text, punctuation, accented letters, and the language that fails. Compare three things: the HTML string, the text selected and pasted from the PDF into a plain-text editor, and text extracted by a second PDF reader or extractor. Record whether the problem is visual or limited to selection.

  • Wrong glyphs: the PDF itself displays incorrect symbols. Investigate source characters, encoding, font files, and fallback fonts.
  • Missing characters or replacement boxes: check that the selected font contains the glyphs and that the browser really loaded the intended font.
  • Reversed or shuffled words: rendering may be fine while the PDF’s character mapping or reading order is not.
  • Broken spaces: compare extraction in another reader and inspect font shaping, ligatures, and layout changes.

Keep the same PDF viewer and extractor while changing one variable at a time. A symptom that occurs only in one application may be an extraction-order limitation rather than a malformed PDF.

Why a PDF can look right and copy wrong

PDF 32000-1:2008, published by Adobe Systems Incorporated, separates drawing glyphs from making text intelligible to other software. Display uses font character codes to paint glyphs. Copy, search, text-to-speech, and export need Unicode mappings and a meaningful reading order. The specification says that “Tagged PDF defines a set of rules for representing text in the page content so that characters, words, and text order can be determined reliably.” It also states that Tagged PDF requires each character code to be mappable to Unicode.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
PDF Reader, PDF Viewer, PDF Editor- file document
  • Fast PDF reader with read aloud, night mode, reading mode, search and bookmarks
  • Highlight, underline, draw, add notes and text on any PDF
  • Fill PDF forms, sign documents with your finger and protect PDFs with a password
  • Convert PDF to Word or JPG; merge, extract and reorder pages; scan with your camera
  • Works on Fire TV: send PDFs from your phone over Wi-Fi and read them on the big screen

That is why changing a character set declaration alone may not repair a PDF: the browser can receive the correct Unicode string, yet the embedded font or generated PDF can still expose an ambiguous mapping. Conversely, a bad HTML source string can produce wrong text before PDF generation begins.

Build a minimal reproduction before changing versions

Use an explicit UTF-8 document

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <style>
    body { font-family: Arial, sans-serif; }
  </style>
</head>
<body>
  <p>Copy test: café — Привет — 東京 — مرحباً</p>
</body>
</html>

Start with a system or broadly supported font rather than your custom webfont. If this file copies correctly, add your application’s CSS, then the custom font, then the real content until the failure returns. This binary-search approach tells you which input introduces the defect.

Make the PDF and inspect the same sample

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch();
  try {
    const page = await browser.newPage();
    await page.setContent(`<!doctype html>
      <html lang="en">
      <head><meta charset="utf-8">
      <style>body{font-family:Arial,sans-serif}</style></head>
      <body>Copy test: café — Привет — 東京 — مرحباً</body>
      </html>`, { waitUntil: 'networkidle0' });

    await page.pdf({
      path: 'test.pdf',
      format: 'A4',
      printBackground: true
    });
  } finally {
    await browser.close();
  }
})();

Compare the visible page, selected text, and extracted text. Do not call a version change a fix until the same reproduction passes after the change.

Verify fonts, not just font declarations

Wait for the intended font to load

The current Puppeteer PDF guide (shown as version 25.12.0 when consulted) says Page.pdf() waits for fonts by default. That is a documented default, not proof that every webfont completed its network request, selected the expected face, embedded correctly, or received a usable extraction map.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.goto('https://example.com/document', { waitUntil: 'networkidle0' });
await page.evaluate(async () => {
  await document.fonts.ready;
});
console.log(await page.evaluate(() => ({
  status: document.fonts.status,
  loaded: [...document.fonts].map(f => ({ family: f.family, status: f.status }))
})));
await page.pdf({ path: 'document.pdf' });

Check the browser’s network log for failed font requests, blocked cross-origin resources, incorrect MIME types, and redirects. Inspect the computed font-family and use the browser’s font panel, where available, to see the face actually used. Test the same content with a simple known font. If the custom face causes the copied text to reverse or lose characters, compare a different font format or a subsetted versus full font file.

Rank #2
XTEINK X3 3.7" Pocket E-Ink eBook Reader,58g,Magnetic, Mini Ereader Devices
  • 3.7" Pocket eBook Reader, Only Approx. 58g: Take your library anywhere with the XTEINK X3, a compact 3.7-inch lightweight eReader designed for everyday portability. Weighing approximately 58g and measuring just 5.1mm thin, it easily slips into your pocket or bag, making it ideal for reading during commutes, while traveling, or during quick breaks.
  • Paper-feel E-Ink Reading, Made for Focus: Enjoy a clean, paper-feel E-Ink reading experience that feels gentle on the eyes and helps you stay focused. No constant notifications, no social media distractions—just a simple mini eReader built for books, manga, notes, and quiet reading time.
  • Gyroscope Page-Turn + Physical Buttons: Read comfortably with one hand using gyroscope page-turn control and responsive physical buttons. Whether you are standing, commuting, or relaxing, XTEINK X3 makes page turning smoother, easier, and more intuitive than traditional touch-only reading devices.
  • Personalized Features & Long-Lasting Battery:Switch between reading, photos, clock, and more for a customizable experience beyond traditional eReaders. Designed for everyday portability, XTEINK X3 delivers up to 10 hours of reading time, supporting about a week of casual reading on a single charge. For safe charging, use a locally certified charger and keep conductive objects away from the charging pin contacts during charging to help prevent short circuits.
  • Magnetic-Ready Design with Pogo-Pin Charging: XTEINK X3 includes an Adhesive Metal Ring to enable magnetic attachment on compatible non-magnetic phone cases or surfaces, expanding compatibility for everyday use. The magnetic pogo-pin charging design maintains a clean, minimalist appearance while supporting convenient daily charging.

Check shaping and fallback

Mixed scripts, ligatures, right-to-left text, and fallback fonts can expose mapping problems that ordinary English does not. Include representative characters in the minimal test. A visually similar fallback does not guarantee identical Unicode extraction behavior.

Control print and screen CSS deliberately

Page.pdf() generates with the print CSS media type by default, as documented in the Page.pdf() API documentation and Puppeteer’s PDF generation guide. Print rules can replace fonts, hide elements, alter direction, or change layout. To test screen styling instead, set the media type before generating:

await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-styles.pdf', printBackground: true });

This changes which CSS rules apply; it is not a general repair for Unicode mappings. Generate one PDF with print media and one with screen media, then compare both visual output and copied text. If only one fails, inspect the corresponding @media print rules and font declarations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare Puppeteer and Chromium versions carefully

Record the exact Puppeteer package, bundled or executable Chromium version, Node.js version, operating system, and PDF consumer. A historical report, PDF Reverse Words, opened March 16, 2018, described a PDF that looked normal while copied words were reversed; later comments associated a similar symptom with a custom font and said it also appeared when printing through Chrome. It is evidence of a symptom pattern, not proof of one universal defect or remedy.

A December 28, 2024 report, PDF Reverse Words Copy, described reverse-order copying with embedded Noto Sans font data on Puppeteer 23.8.0. The reporter said versions 23.0.0 through 23.7.1 worked in that test and later releases through 23.11.1 did not; the issue was closed as not planned. Treat those observations as case-specific. Do not blindly downgrade production: reproduce with your HTML, font, and Chromium build, then pin a version only after verifying the result and assessing security and maintenance implications.

Rank #3
PDF Extra Ultimate | Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Yearly License | 1 Windows PC & 2 Mobile Devices | 1 User
  • EDIT text, images & designs in PDF documents. ORGANIZE PDFs. Convert PDFs to Word, Excel & ePub.
  • READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.
  • CREATE, COMBINE, SCAN and COMPRESS PDFs
  • FILL forms & Digitally Sign PDFs. PROTECT and Encrypt PDFs
  • 1 Year License for 1 Windows & 2 Mobile (Android and/or iOS) devices.

A separate May 16, 2024 report, Issue with Text Encoding in PDF Generation Using Puppeteer, used Puppeteer 22.8.2 and Node 18.16 on Windows but was labeled not reproducible and unconfirmed. It does not establish a general Puppeteer encoding bug.

Use a controlled comparison matrix

Variable Test A Test B What it tells you
Symptom Visual rendering Selection/extraction Separates glyph drawing from text data
Media Print (default) Screen via emulateMediaType Identifies print-CSS effects
Font Known system font Custom webfont Tests font loading, embedding, and mapping
Runtime Current pinned build One documented alternate build Checks version-specific behavior
Consumer Your normal reader Another reader or extractor Finds viewer-specific order problems

Keep the HTML, URL, viewport, locale, and PDF options constant for each pair. Capture the copied output and the PDF files as artifacts so another developer can reproduce the comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common failure modes and fixes

The page is visibly wrong

  • Inspect the original JavaScript string and database values for accidental byte decoding or replacement characters.
  • Confirm <meta charset="utf-8"> and the HTTP Content-Type charset where applicable.
  • Check font coverage and failed font requests; temporarily use a known font.

Only copied words are reversed

  • Run the known-font minimal reproduction.
  • Compare print and screen media.
  • Test another reader or extractor to distinguish order interpretation from malformed mapping.
  • Record the exact Puppeteer/Chromium build and test a narrowly chosen alternate version.

Accents or non-Latin characters disappear

  • Verify the selected font contains those code points and that fallback is not unexpectedly used.
  • Wait for document.fonts.ready after navigation and inspect font requests.
  • Test a full, un-subsetted font file and a minimal string in the affected script.

Adding emulateMediaType('screen') changes nothing

That method controls CSS media selection only. If copied text remains wrong in both modes, focus on font mapping, extraction order, and the runtime rather than treating media emulation as a text-repair switch.

A proposed downgrade “works” once and then fails in deployment

Ensure production uses the same Chromium executable, OS, font installation, locale, and launch flags as the test. Pin both package and browser inputs, retain the minimal PDF as a regression fixture, and review security updates before keeping an older build.

Make the PDF more extraction-friendly

  • Use real text nodes instead of drawing letters to a canvas or converting them to images when copy/search matters.
  • Keep logical reading order in the DOM, especially for columns, sidebars, and right-to-left content.
  • Avoid unnecessary absolute positioning and overlapping text layers.
  • Use stable, properly licensed font files with complete Unicode coverage for the languages you publish.
  • When accessibility and reliable extraction are requirements, evaluate whether your PDF pipeline can produce tagged structure; visual correctness alone is not a sufficient acceptance test.

PDF 32000-1:2008 describes Unicode mappings and tagged structure as the mechanisms supporting copy-and-paste, searching, speech, and export. Puppeteer’s PDF call does not automatically guarantee that every generated document meets those semantic goals.

Rank #4
PDF Extra 2024| Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Lifetime License | 1 Windows PC | 1 User [PC Online code]
  • EDIT text, images & designs in PDF documents. ORGANIZE PDFs. Convert PDFs to Word, Excel & ePub.
  • READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.
  • CREATE, COMBINE, SCAN and COMPRESS PDFs
  • FILL forms & Digitally Sign PDFs. PROTECT and Encrypt PDFs
  • LIFETIME License for 1 Windows PC or Laptop. 5GB MobiDrive Cloud Storage Included.

Performance, reliability, and cost considerations

Font downloads, JavaScript-heavy pages, and networkidle0 waits can lengthen generation. Cache trusted font assets, avoid waiting for permanently open connections, and use a targeted readiness condition when appropriate. Do not remove font or content waits merely to reduce latency; an early capture can embed fallback fonts or incomplete text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For reproducible builds, record the HTML input hash, Puppeteer and Chromium versions, font file checksums, launch options, media type, locale, and PDF options. Keep a fixture containing the characters that previously failed. Validate visual output, selection, search, and extraction order in continuous integration when PDFs are user-facing.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean screenshot or PDF of a URL rather than debugging a local PDF pipeline, ScreenshotNeo provides a single website screenshot API request and an MCP server for AI clients. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also offers PDF output, custom CSS and JavaScript, font and network controls, wait conditions, device presets, signed links, asynchronous jobs, bulk capture, and usage reporting.

Use the API examples in the ScreenshotNeo documentation as the starting point:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to include in a bug report

  • Whether the PDF looks wrong, copies wrong, searches wrong, or only fails in one consumer.
  • A short pasted sample showing the expected and actual strings, including affected languages.
  • Minimal HTML, CSS, and a generated PDF that reproduces the issue.
  • Puppeteer version, bundled Chromium revision or executable version, Node.js version, operating system, and PDF reader/extractor.
  • Font family, file format, source URL, loading status, and whether the known-font test passes.
  • Media type, viewport, locale, wait strategy, and relevant PDF options.

FAQ

Is this always a UTF-8 problem?

No. UTF-8 source errors are one possibility; font embedding, Unicode maps, reading order, print CSS, extraction software, and version-specific behavior can produce similar symptoms.

Best Value
PDF Extra Lifetime - Professional PDF Editor - Best Adobe Acrobat Pro Alternative - Lifetime License for Windows PC
  • Perfect Adobe Acrobat Pro alternative – lifetime license for Windows 10 and 11.
  • EDIT text, images, pages, hyperlinks, designs in PDF documents. ORGANIZE PDFs.
  • READ and Comment on PDFs – Intuitive reading modes & document commenting and mark up tools!
  • CREATE, COMBINE, SCAN and COMPRESS PDFs.
  • FILL forms & Digitally Sign PDFs. Work with Digital certificates

Should I immediately downgrade Puppeteer?

No. Version reports are tied to particular fonts, Chromium builds, and inputs. Reproduce first, compare a controlled alternate, and consider security and maintenance before pinning an older release.

Can a different PDF reader repair the file?

A different reader can extract the same content in a different order and help identify a consumer-specific problem, but it does not rewrite an invalid mapping in the original PDF.

Does waiting for fonts guarantee correct copy-and-paste?

No. It reduces the chance of capturing before a font loads. The generated font embedding and Unicode mapping still need to be tested with the actual font and content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What is the fastest first test?

Generate the same minimal HTML twice: once with a known system font and once with the custom font, then compare selected text and extraction in two readers.

Which details matter most when asking for help?

Provide the minimal HTML/PDF pair, affected characters, actual loaded font, Puppeteer and Chromium versions, operating system, media type, and the reader or extractor that shows the failure.

Quick Recap

Bestseller No. 1
PDF Reader, PDF Viewer, PDF Editor- file document
PDF Reader, PDF Viewer, PDF Editor- file document
Fast PDF reader with read aloud, night mode, reading mode, search and bookmarks; Highlight, underline, draw, add notes and text on any PDF
$6.85
Bestseller No. 3
PDF Extra Ultimate | Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Yearly License | 1 Windows PC & 2 Mobile Devices | 1 User
PDF Extra Ultimate | Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Yearly License | 1 Windows PC & 2 Mobile Devices | 1 User
READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.; CREATE, COMBINE, SCAN and COMPRESS PDFs
$83.88
Bestseller No. 4
PDF Extra 2024| Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Lifetime License | 1 Windows PC | 1 User [PC Online code]
PDF Extra 2024| Complete PDF Reader and Editor | Create, Edit, Convert, Combine, Comment, Fill & Sign PDFs | Lifetime License | 1 Windows PC | 1 User [PC Online code]
READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.; CREATE, COMBINE, SCAN and COMPRESS PDFs
$99.99
Bestseller No. 5
PDF Extra Lifetime - Professional PDF Editor - Best Adobe Acrobat Pro Alternative - Lifetime License for Windows PC
PDF Extra Lifetime - Professional PDF Editor - Best Adobe Acrobat Pro Alternative - Lifetime License for Windows PC
Perfect Adobe Acrobat Pro alternative – lifetime license for Windows 10 and 11.; EDIT text, images, pages, hyperlinks, designs in PDF documents. ORGANIZE PDFs.
$99.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.