Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Base64 string is only an encoding of PDF bytes. Decode it first, then give the resulting binary data to a PDF parser and map the parser’s page text into the JSON shape your application needs. In Node.js, use Buffer.from(value, 'base64') and a buffer-capable extractor. In a browser, convert the string to a Uint8Array and pass it to PDF.js. Neither route performs OCR on an image-only scan.

This guide shows both implementations, a practical page-oriented JSON schema, handling for data-URI prefixes and memory limits, and fixes for common failures.

What the conversion actually involves

There are two separate operations:

  1. Decode: Base64 text becomes the original PDF bytes.
  2. Extract: A PDF library reads those bytes and exposes text items, page metadata, coordinates, or other document structures.

Returning JSON is your application’s design decision. A useful default is one object per page, with a page number and a single normalized text string. You can later add coordinates, raw text items, or extracted rows without changing the decoding step.

Node.js: decode a Base64 PDF and return page JSON

Install the extractor

The pdf.js-extract documentation provides extractBuffer(buffer, options, callback) and exposes page content items whose text is in item.str.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
npm install pdf.js-extract

Complete example

The following function accepts either plain Base64 or a data:application/pdf;base64, value, decodes it once, and serializes one result object per page.

import { PDFExtract } from 'pdf.js-extract';

function removeDataUriPrefix(value) {
  const comma = value.indexOf(',');
  return value.startsWith('data:') && comma !== -1
    ? value.slice(comma + 1)
    : value;
}

export function extractPdfJson(base64Pdf) {
  const payload = removeDataUriPrefix(base64Pdf).replace(/s+/g, '');
  const pdfBuffer = Buffer.from(payload, 'base64');
  const extractor = new PDFExtract();

  return new Promise((resolve, reject) => {
    extractor.extractBuffer(pdfBuffer, {}, (error, data) => {
      if (error) {
        reject(error);
        return;
      }

      const pages = data.pages.map((page) => ({
        page: page.info.num,
        text: page.content.map((item) => item.str).join(' ')
      }));

      resolve({ pages });
    });
  });
}

const result = await extractPdfJson(base64Pdf);
console.log(JSON.stringify(result, null, 2));

Node’s Buffer documentation states that the Base64 decoder accepts the URL-safe alphabet and ignores whitespace. Stripping whitespace yourself makes input handling explicit, while Buffer.from performs the actual conversion. Validate that the decoded value is non-empty before parsing so an empty or truncated upload does not become an opaque parser error.

Typical Node output

{
  "pages": [
    { "page": 1, "text": "Invoice 1042 Acme Corporation ..." },
    { "page": 2, "text": "Payment terms Net 30 ..." }
  ]
}

The property names are not imposed by PDF.js or by the package. Rename page to pageNumber, return a combined fullText, or retain each item’s coordinates when your consumer needs layout information.

Browser: turn Base64 into a Uint8Array for PDF.js

PDF.js accepts binary document data and documents Uint8Array as the preferred representation for memory use. Its API documentation and examples show loading a document, retrieving each page, and calling getTextContent(). The project’s FAQ advises decoding Base64 before supplying the bytes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Browser extraction function

This example assumes your application has loaded PDF.js and made its API available as pdfjsLib. A bundler setup must also configure PDF.js’s worker as required by the version you install.

function base64ToUint8Array(input) {
  const comma = input.indexOf(',');
  const payload = input.startsWith('data:') && comma !== -1
    ? input.slice(comma + 1)
    : input;
  const binary = atob(payload.replace(/s+/g, ''));
  const bytes = new Uint8Array(binary.length);

  for (let i = 0; i < binary.length; i += 1) {
    bytes[i] = binary.charCodeAt(i);
  }
  return bytes;
}

async function extractPdfJson(base64Pdf) {
  const data = base64ToUint8Array(base64Pdf);
  const loadingTask = pdfjsLib.getDocument({ data });
  const pdf = await loadingTask.promise;
  const pages = [];

  for (let pageNumber = 1; pageNumber <= pdf.numPages; pageNumber += 1) {
    const page = await pdf.getPage(pageNumber);
    const textContent = await page.getTextContent();
    pages.push({
      page: pageNumber,
      text: textContent.items.map((item) => item.str).join(' ')
    });
  }

  return { pages };
}

const result = await extractPdfJson(base64Pdf);
console.log(JSON.stringify(result));

atob() produces a binary string, so the loop copies each byte into a typed array. Do not pass the original Base64 characters as the data value; PDF.js expects decoded binary data.

When the input is a data URL

A caller may send data:application/pdf;base64,JVBERi0x... instead of the payload alone. The prefix is metadata, not part of the encoded bytes. Remove everything through the first comma before calling atob. If your API contract accepts only raw Base64, reject a data URL with a clear validation message rather than silently producing incorrect bytes.

Choose a JSON shape deliberately

There is no universal “PDF text JSON” standard. Select a shape based on what downstream code must do.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Use case Suggested fields Trade-off
Search or indexing pages[].page, pages[].text Small and easy to consume, but layout is lost.
Auditing or highlighting Page number plus each item’s str, transform, and bounding data Preserves locations; consumers must reconstruct lines.
Simple export A single fullText string Convenient for prompts and logs, but page boundaries disappear.
Table-oriented processing Page text items grouped into lines or rows Useful for known templates; grouping is not guaranteed semantic table recognition.

pdf.js-extract exposes coordinates and row-grouping utilities, but inspect representative documents before treating a grouped row as a real table record. Columns, headers, and reading order vary between PDFs.

Handling difficult input

Scanned or image-only PDFs

Text extraction reads a PDF’s text layer. The pdf.js-extract package explicitly states “NO OCR!” A scan made from page images can therefore produce empty strings even though a person can see words on the page. Add a separate OCR stage, retain the page images, and mark OCR-derived text differently if accuracy matters. Do not report an empty extraction as proof that the document contains no text until you have checked whether it is image-only.

Password-protected files

PDF.js’s API includes a password-loading parameter. Your application must obtain the password and provide it through the parser’s supported mechanism. Encryption compatibility and error wording depend on the library and document, so surface a specific “password required” state instead of retrying indefinitely.

Malformed or unsupported PDFs

Catch parser errors around the extraction call. Preserve an input identifier and byte length, return a stable application error such as PDF_PARSE_FAILED, and avoid returning partially assembled JSON as if it were complete. Test the exact classes of PDFs your service receives; no parser guarantees success for every malformed or feature-rich file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Tables, columns, and reading order

Text items commonly include positions, but concatenating them in array order can interleave columns or separate labels from values. For reliable layout-sensitive processing, keep the item coordinates, group by page and line, and add document-specific sorting rules. A visual table is not automatically a semantic table.

Memory and performance considerations

Base64 expands binary data and decoding can temporarily create more than one copy in memory. The PDF.js FAQ recommends supplying raw binary data as a typed array when possible. If an upstream system already gives you Base64, decode it once, release the original string when safe, and avoid converting the bytes back to Base64 between parsing steps.

  • Set an input-size limit before decoding to protect browser tabs and server processes.
  • Process pages incrementally when your parser permits it instead of retaining every intermediate object.
  • Return page results progressively or store them in a bounded queue for very large documents.
  • Normalize whitespace only after extraction; premature normalization can hide missing text or alter meaningful spacing.

No benchmark is implied by these practices. Actual throughput depends on page count, embedded fonts and images, parser version, available memory, and whether OCR is added.

Browser PDF.js, Node.js, or a viewer SDK?

Option Runtime Input path Best fit OCR included?
PDF.js Browser Decode to typed-array binary and call getDocument Client-side viewing and page text Not established as included by the cited API material
Node.js plus pdf.js-extract Node.js Buffer.from, then extractBuffer Server-side extraction, coordinates, and custom JSON No; the package says “NO OCR!”
PDF.js Express Browser viewer SDK Vendor documents Base64-to-Blob loading Applications needing broader viewer/document operations Not established by the cited Base64 page

PDF.js is an open-source project. PDF.js Express is described in its Base64 documentation as a commercial SDK; it is not required for ordinary text extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting checklist

Symptom Likely cause Fix
“Invalid PDF” immediately after decoding The data-URL prefix, whitespace, or a JSON wrapper was passed as payload. Extract the value, remove the prefix, strip whitespace, and verify the decoded bytes begin with a valid PDF header.
atob throws a decoding error The browser received URL-safe characters, corrupted text, or non-Base64 input. Validate the producer’s encoding contract and reject malformed input before calling atob.
Every page has empty text The PDF is scanned or has no usable text layer. Inspect the file visually and add OCR for image-only pages.
Words appear in the wrong order Multiple columns or positioned text items were joined without layout rules. Retain coordinates and implement page-specific line and column grouping.
Password prompt or encryption error The file is protected. Collect the password and pass it through the parser’s documented password option.
Browser tab or process runs out of memory Large Base64 strings and duplicate byte arrays remain live. Limit input size, decode once, release references, and process pages incrementally.
Node callback reports an extraction error The PDF is malformed, unsupported, or the installed package API differs. Log the library version, test the file independently, and consult the installed package documentation before changing the callback or options.

Or skip the browser setup

If your real input is a live webpage and you need a clean PDF or screenshot before a downstream text pipeline, ScreenshotNeo provides a website screenshot API and MCP server. It is not a replacement for decoding an existing Base64 PDF buffer; it creates a capture from a URL. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets, with each cleanup step configurable. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

For a one-call image capture, see the ScreenshotNeo API documentation:

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request from Python:

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

And from Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then((fs) => fs.writeFile('shot.webp', data));

ScreenshotNeo offers 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account if generating a clean capture is easier than maintaining browser automation.

Operational practices for production services

  • Define whether your endpoint accepts raw Base64, a data URL, or both, and document the maximum decoded byte size.
  • Return a stable schema with a parser status, page count, and per-page results so callers can distinguish an empty page from a failed extraction.
  • Keep parser errors separate from OCR errors; they require different remediation and monitoring.
  • Use fixture PDFs representing text, columns, tables, scans, encryption, and malformed input. Compare normalized expected JSON in automated tests.
  • Apply access controls to uploaded PDFs and avoid logging raw Base64, which can contain confidential document content.

FAQ

Can I change PDF libraries without changing my API?

Yes. Keep your public page-oriented schema in an adapter layer. Each parser maps its own page and text-item structures into that contract, allowing library upgrades or runtime changes without forcing every client to rewrite its JSON handling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should regression tests compare?

Store representative PDFs with expected page counts and normalized page text. Add separate assertions for coordinates or row grouping when layout matters, because those details can change independently of the decoded bytes.

Frequently Asked Questions

Can I change PDF libraries without changing my API?

Yes. Keep your public page-oriented schema in an adapter layer. Each parser maps its own page and text-item structures into that contract, allowing library upgrades or runtime changes without forcing every client to rewrite its JSON handling.

What should regression tests compare?

Store representative PDFs with expected page counts and normalized page text. Add separate assertions for coordinates or row grouping when layout matters, because those details can change independently of the decoded bytes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.