Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a five-PDF benchmark, CleanMD handled some structures well but did not win every comparison. MarkItDown and pymupdf4llm each had document-specific strengths, and CleanMD’s first run missed 11 RFC sections before a font-handling fix. The results are useful for spotting failure modes—not for declaring a universal best converter.

What the benchmark tested

Giacomo, CleanMD’s developer, compared three converters using their default settings: MarkItDown 0.1.7 from Microsoft, pymupdf4llm 1.28.2 from Artifex, and CleanMD 0.93.0, run in Node. These versions and setup are those reported in the article published September 24, 2026; they are not a statement about current releases. The author excluded Pandoc because it cannot read PDFs and returned code 21 for all five files.

The five public PDFs covered different document types. The author says they were not selected after viewing converter results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
PDF Length What it represents
RFC 9110, HTTP Semantics 194 pages Technical specification with sections, running footers, code and collected ABNF grammar
BERT, arXiv 1810.04805 16 pages Research paper with headings and tables
Think Python, 2nd edition 292 pages Technical book with a deep section hierarchy and many code examples
Loper Bright v. Raimondo 114 pages Court opinion with majority, concurrence and dissent sections
NIST CSF 2.0 32 pages Framework document with sections and tables

All five PDFs had text layers. The author compared outputs with ground truth based on document structures, such as official text or tables of contents. The court opinion has no table of contents, so its 31 heading markers were initially defined geometrically. The checks covered heading recognition and depth, fenced code, table counts and “prose cells” (sentences incorrectly represented as table cells), along with document-specific items such as whether collected ABNF appeared as one block. Timing was measured only for RFC 9110, on the same laptop.

#1 Best Overall
MixPad Free Multitrack Recording Studio and Music Mixing Software [Download]
  • Create a mix using audio, music and voice tracks and recordings.
  • Customize your tracks with amazing effects and helpful editing tools.
  • Use tools like the Beat Maker and Midi Creator.
  • Work efficiently by using Bookmarks and tools like Effect Chain, which allow you to apply multiple effects at a time
  • Use one of the many other NCH multimedia applications that are integrated with MixPad.

The author links a kit containing run.sh, ground truth, results and raw CleanMD Markdown, and describes a process that downloads the PDFs, installs the two open-source tools in a virtual environment, and scores the outputs. The kit makes the comparison inspectable; the results below are the author’s reported measurements, not an independently reproduced run. Read the benchmark and access the kit.

Where CleanMD did well

RFC code, grammar and recurring footers

In RFC 9110, CleanMD produced 158 fenced blocks, while MarkItDown and pymupdf4llm produced none. The collected ABNF appeared as one block only in CleanMD’s output. CleanMD also left no counted RFC footer lines, compared with 187 for MarkItDown and 194 for pymupdf4llm.

On the author’s initial RFC run, CleanMD recognized 280 of 291 sections; pymupdf4llm recognized 290 of 291. The developer traced CleanMD’s 11 missed headings to hyphen glyphs that had been split across font IDs, fixed the issue, and reran the benchmark at 291 of 291. The article retains the initial 280/291 result rather than replacing it with the corrected score. As Giacomo put it, “I am keeping the first number in the article and on the page, because a benchmark that only shows the after is marketing.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Think Python’s heading hierarchy

CleanMD placed 122 of 126 Think Python headings at the correct depth. MarkItDown got none of the 126 at the correct depth, and pymupdf4llm also got none: it recognized 113 sections but placed them all at H4. This is a hierarchy result, not simply a count of whether any headings appeared.

NIST section recognition

CleanMD and pymupdf4llm each recognized eight NIST CSF 2.0 sections; MarkItDown recognized none. The NIST table result, however, favors a different tool, as the next section shows.

Where CleanMD lost or produced trade-offs

RFC section recall on the first run

Before the fix, CleanMD’s 280/291 RFC section result trailed pymupdf4llm’s 290/291. The corrected 291/291 shows the defect was addressed in a later run, but it does not erase the original failure from this comparison.

Rank #4
DeskFX Free Audio Effects & Audio Enhancer Software [PC Download]
  • Transform audio playing via your speakers and headphones
  • Improve sound quality by adjusting it with effects
  • Take control over the sound playing through audio hardware

Think Python: more fences did not mean more useful fences

CleanMD emitted 569 Think Python fenced blocks, 52 of them single-line. pymupdf4llm emitted 657, but 328 were single-line fences around inline code. MarkItDown emitted none. A raw fence count therefore does not tell you whether code examples are grouped usefully: the contents and boundaries matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

BERT and NIST tables

For BERT, the article reports 11 tables from CleanMD, 134 from MarkItDown and 9 from pymupdf4llm. All three had zero counted prose cells in that document. The large difference in table counts makes table boundaries worth inspecting rather than assuming that a higher count means better fidelity.

Best Value
WavePad Audio Editing Software - Professional Audio and Music Editor for Anyone [Download]
  • Full-featured professional audio and music editor that lets you record and edit music, voice and other audio recordings
  • Add effects like echo, amplification, noise reduction, normalize, equalizer, envelope, reverb, echo, reverse and more
  • Supports all popular audio formats including, wav, mp3, vox, gsm, wma, real audio, au, aif, flac, ogg and more
  • Sound editing functions include cut, copy, paste, delete, insert, silence, auto-trim and more
  • Integrated VST plugin support gives professionals access to thousands of additional tools and effects

On NIST CSF 2.0, MarkItDown had zero counted prose cells, while CleanMD had 23. MarkItDown also emitted zero NIST headings, so its clean prose-cell count does not make it the better fit for a document where section structure matters.

Court-opinion headings have a ground-truth caveat

For Loper Bright v. Raimondo, the benchmark counted part markers identified as headings: CleanMD found 28 of 31, pymupdf4llm 18, and MarkItDown none. Treat this row cautiously. The original geometric rule used to define the 31 markers was close to CleanMD’s own heading heuristic, which could favor that converter. In a comment, Giacomo acknowledged the concern and said the majority-opinion markers were checked against Cornell LII HTML (15 of 15); the concurrence and dissent had not yet been checked at that point. The author summarized the distinction this way: “The rule generated the list, but the list can be checked without the rule.”

RFC runtime is one-file evidence

On the same laptop, RFC 9110 conversion took 1.3 seconds for CleanMD, 5.5 seconds for MarkItDown and 10.9 seconds for pymupdf4llm. Those are the author’s measurements for that file and setup only; they do not establish a general speed ranking.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What this comparison cannot tell you

The test covers five text-layer PDFs and selected structural checks. It does not measure scanned-document OCR, dense three-column reading order, mathematical notation, or prose quality. A converter could score well on the reported headings and fences yet still struggle with the specific documents you need to convert.

Giacomo describes the benchmark as one that “shows failure modes, not a universal ranking.” That is the right way to use it: as evidence about these files and as a guide to what to inspect in your own.

Quick Recap

Bestseller No. 1
MixPad Free Multitrack Recording Studio and Music Mixing Software [Download]
MixPad Free Multitrack Recording Studio and Music Mixing Software [Download]
Create a mix using audio, music and voice tracks and recordings.; Customize your tracks with amazing effects and helpful editing tools.
Bestseller No. 4
DeskFX Free Audio Effects & Audio Enhancer Software [PC Download]
DeskFX Free Audio Effects & Audio Enhancer Software [PC Download]
Transform audio playing via your speakers and headphones; Improve sound quality by adjusting it with effects

How to choose a converter for your PDFs

  1. Start with your document type. If you need specifications, check section depth, code blocks, grammar and running headers or footers. For books, inspect whether code examples remain coherent instead of being fragmented into single-line fences. For research papers and frameworks, compare table boundaries and check whether prose has been misclassified as a cell.
  2. Inspect reading order and OCR separately. This benchmark does not establish performance on scanned pages or dense multi-column layouts. Test representative pages from those files and verify that the extracted text follows the intended reading order.
  3. Use a known answer where possible. Compare headings, tables and code with the source PDF or an authoritative text version. A benchmark score is only as meaningful as its ground truth; pay particular attention when the expected structure was defined by a geometric rule similar to a converter’s own.
  4. Time your own representative files. The RFC timings do not predict which tool will be fastest on your PDFs, machines or configurations.
  5. Choose by the errors you can tolerate. Missing a heading, splitting a code block, inventing a table or retaining repeated footer text can have different costs depending on whether you plan to search, edit, publish or feed the Markdown into another tool.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.