The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A missing /ToUnicode map can explain why some PDF text extracts as wrong characters, but it cannot explain every copy-and-paste or extraction failure. It addresses how a font’s character codes map to Unicode—not whether a page contains selectable text, whether OCR recognized a scan correctly, or whether extracted text will follow the visual reading order.
What a ToUnicode map does—and what its absence tells you
A PDF font uses character codes to select glyphs for display. A font’s encoding and other information may help convey what those characters mean; a /ToUnicode CMap can provide an explicit mapping from the PDF’s character codes to Unicode values used for text extraction. Adobe’s PDF Reference, Second Edition describes the entry as optional and explains why it can matter: “In the absence of a /ToUnicode entry, there would be no information available about what the characters mean.” That statement describes a case where no other character-meaning information is available; absence of the entry alone does not prove every font or PDF will fail to extract correctly.
A presence-only test has a limit, too. Finding a CMap does not prove that its mappings are appropriate for the visible glyphs. If the PDF has selectable text but extraction returns the wrong characters, inspect the font encoding and mapping—including whether the CMap is absent, malformed, or mismatched—rather than treating the test as a verdict. The PDF Association’s text errata for PDF 32000-2:2020, clause 9 provides standards context for font and text behavior.
Why the text can still be wrong when the map is present
The page may be an image, not text
A scanned page can consist of a raster image with no underlying selectable text. A PDF parser that extracts text does not automatically recognize letters in that image. Optical character recognition (OCR) can add a machine-readable text layer, but OCR can also misread characters. The pypdf extraction guide and its current extraction documentation distinguish text extraction from OCR and recommend OCR for image-only pages.
#1 Best Overall
Correct characters can come out in the wrong order
PDFs are designed to display content at specified positions. Their text-drawing commands do not necessarily encode the order a person would read the page. An extractor may return recognizable words but place them in a surprising sequence, especially across columns or separately positioned text. pypdf notes that extraction behavior can vary with the PDF generator; its PageObject documentation describes layout-oriented extraction options, which attempt to account for placement rather than simply returning text in command order.
Visual layout is not necessarily semantic structure
A heading that looks like a heading, a paragraph, or a table may be encoded only as positioned text—not as a semantic heading, paragraph, or table object. Extracting Unicode characters does not by itself restore those relationships. A table’s cells may be laid out with positioning and spacing, leaving an extraction tool to infer which values belong together. pypdf’s guide to text extraction discusses this gap between visual presentation and semantic structure.
Rank #2
Diagnose the symptom before choosing a fix
Start with the rendered page and compare it with the extracted result. The visible page, selectable text, character accuracy, reading sequence, and structural relationships are different things to check.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches| What you observe | What to investigate |
|---|---|
| No selectable text on a page that visibly contains words | Whether the page is image-only and needs OCR. |
| Selectable text, but extracted characters differ from the visible glyphs | The font encoding and character mapping, including the /ToUnicode CMap. |
| Characters are right, but words, lines, or columns are jumbled | Text sequence and layout reconstruction. |
| Characters are legible, but table cells or heading relationships are lost | Whether the document contains semantic structure and whether the extraction method can preserve or infer it. |
| Extracted text from a scanned PDF contains recognition mistakes | The OCR text layer, checked against the page image. |
Work through the PDF in layers
- Check whether text is selectable. Compare the rendered page with what your PDF viewer lets you select and copy. If the words are visible but no text can be selected, investigate an image-only page and OCR before focusing on
/ToUnicode. - Compare characters with the rendered glyphs. If selectable text extracts as symbols or incorrect letters, inspect the font’s encoding and mapping. A missing CMap may be relevant, but also check whether an existing one is malformed or unsuitable for the displayed text.
- Check order separately from character accuracy. If individual characters are right but sentences or columns are scrambled, compare the extraction with the page’s visual order. Where the tool offers layout-oriented extraction, try it and verify the result against the rendered page. pypdf documents extraction options and cautions that results depend on how the PDF was generated in its current guide and PageObject reference.
- For a scan with selectable text, check OCR quality. A hidden OCR layer can make scanned words selectable without making them accurate. Compare the extracted text with the page image to see whether recognition errors, rather than font mapping, account for the mismatch.
- For a conformance question, use a validator for that purpose. veraPDF’s validation documentation describes validation for PDF/A and PDF/UA. A conformance result is not, by itself, proof that a particular reader or extractor will produce the desired wording, reading order, or layout.
Choose the remedy that matches the failure
- Wrong characters: investigate font encoding and Unicode mapping.
- No selectable text: determine whether the page is an image and whether OCR is needed.
- Right characters, wrong sequence: investigate content order and layout-aware extraction.
- Missing table or heading relationships: assess semantic structure; a character-map change alone will not recreate it.
There is no established prevalence figure in the cited PDF specification and tool documentation for the share of extraction failures caused by missing /ToUnicode maps. Treat the map as one diagnostic clue, then judge the output against the rendered page and the result you actually need: plain text, visually sensible layout, or structured content.
Quick Recap
Rank #4
Rank #3
- hole punched
- high quality card stock
- 4 pages
- made in USA
- keyboard shortcuts
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

