Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In a five-PDF benchmark, CleanMD handled some structures well but did not win every comparison. MarkItDown and pymupdf4llm each had document-specific strengths, and CleanMD’s first run missed 11 RFC sections before a font-handling fix. The results are useful for spotting failure modes—not for declaring a universal best converter.
What the benchmark tested
Giacomo, CleanMD’s developer, compared three converters using their default settings: MarkItDown 0.1.7 from Microsoft, pymupdf4llm 1.28.2 from Artifex, and CleanMD 0.93.0, run in Node. These versions and setup are those reported in the article published September 24, 2026; they are not a statement about current releases. The author excluded Pandoc because it cannot read PDFs and returned code 21 for all five files.
The five public PDFs covered different document types. The author says they were not selected after viewing converter results.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall| Length | What it represents | |
|---|---|---|
| RFC 9110, HTTP Semantics | 194 pages | Technical specification with sections, running footers, code and collected ABNF grammar |
| BERT, arXiv 1810.04805 | 16 pages | Research paper with headings and tables |
| Think Python, 2nd edition | 292 pages | Technical book with a deep section hierarchy and many code examples |
| Loper Bright v. Raimondo | 114 pages | Court opinion with majority, concurrence and dissent sections |
| NIST CSF 2.0 | 32 pages | Framework document with sections and tables |
All five PDFs had text layers. The author compared outputs with ground truth based on document structures, such as official text or tables of contents. The court opinion has no table of contents, so its 31 heading markers were initially defined geometrically. The checks covered heading recognition and depth, fenced code, table counts and “prose cells” (sentences incorrectly represented as table cells), along with document-specific items such as whether collected ABNF appeared as one block. Timing was measured only for RFC 9110, on the same laptop.
#1 Best Overall
- Create a mix using audio, music and voice tracks and recordings.
- Customize your tracks with amazing effects and helpful editing tools.
- Use tools like the Beat Maker and Midi Creator.
- Work efficiently by using Bookmarks and tools like Effect Chain, which allow you to apply multiple effects at a time
- Use one of the many other NCH multimedia applications that are integrated with MixPad.
The author links a kit containing run.sh, ground truth, results and raw CleanMD Markdown, and describes a process that downloads the PDFs, installs the two open-source tools in a virtual environment, and scores the outputs. The kit makes the comparison inspectable; the results below are the author’s reported measurements, not an independently reproduced run. Read the benchmark and access the kit.
Where CleanMD did well
RFC code, grammar and recurring footers
In RFC 9110, CleanMD produced 158 fenced blocks, while MarkItDown and pymupdf4llm produced none. The collected ABNF appeared as one block only in CleanMD’s output. CleanMD also left no counted RFC footer lines, compared with 187 for MarkItDown and 194 for pymupdf4llm.
Rank #2
On the author’s initial RFC run, CleanMD recognized 280 of 291 sections; pymupdf4llm recognized 290 of 291. The developer traced CleanMD’s 11 missed headings to hyphen glyphs that had been split across font IDs, fixed the issue, and reran the benchmark at 291 of 291. The article retains the initial 280/291 result rather than replacing it with the corrected score. As Giacomo put it, “I am keeping the first number in the article and on the page, because a benchmark that only shows the after is marketing.”
Recommended Free Tools
Think Python’s heading hierarchy
CleanMD placed 122 of 126 Think Python headings at the correct depth. MarkItDown got none of the 126 at the correct depth, and pymupdf4llm also got none: it recognized 113 sections but placed them all at H4. This is a hierarchy result, not simply a count of whether any headings appeared.
NIST section recognition
CleanMD and pymupdf4llm each recognized eight NIST CSF 2.0 sections; MarkItDown recognized none. The NIST table result, however, favors a different tool, as the next section shows.
Where CleanMD lost or produced trade-offs
RFC section recall on the first run
Before the fix, CleanMD’s 280/291 RFC section result trailed pymupdf4llm’s 290/291. The corrected 291/291 shows the defect was addressed in a later run, but it does not erase the original failure from this comparison.
Rank #4
- Transform audio playing via your speakers and headphones
- Improve sound quality by adjusting it with effects
- Take control over the sound playing through audio hardware
Think Python: more fences did not mean more useful fences
CleanMD emitted 569 Think Python fenced blocks, 52 of them single-line. pymupdf4llm emitted 657, but 328 were single-line fences around inline code. MarkItDown emitted none. A raw fence count therefore does not tell you whether code examples are grouped usefully: the contents and boundaries matter.
BERT and NIST tables
For BERT, the article reports 11 tables from CleanMD, 134 from MarkItDown and 9 from pymupdf4llm. All three had zero counted prose cells in that document. The large difference in table counts makes table boundaries worth inspecting rather than assuming that a higher count means better fidelity.
Best Value
- Full-featured professional audio and music editor that lets you record and edit music, voice and other audio recordings
- Add effects like echo, amplification, noise reduction, normalize, equalizer, envelope, reverb, echo, reverse and more
- Supports all popular audio formats including, wav, mp3, vox, gsm, wma, real audio, au, aif, flac, ogg and more
- Sound editing functions include cut, copy, paste, delete, insert, silence, auto-trim and more
- Integrated VST plugin support gives professionals access to thousands of additional tools and effects
On NIST CSF 2.0, MarkItDown had zero counted prose cells, while CleanMD had 23. MarkItDown also emitted zero NIST headings, so its clean prose-cell count does not make it the better fit for a document where section structure matters.
Court-opinion headings have a ground-truth caveat
For Loper Bright v. Raimondo, the benchmark counted part markers identified as headings: CleanMD found 28 of 31, pymupdf4llm 18, and MarkItDown none. Treat this row cautiously. The original geometric rule used to define the 31 markers was close to CleanMD’s own heading heuristic, which could favor that converter. In a comment, Giacomo acknowledged the concern and said the majority-opinion markers were checked against Cornell LII HTML (15 of 15); the concurrence and dissent had not yet been checked at that point. The author summarized the distinction this way: “The rule generated the list, but the list can be checked without the rule.”
RFC runtime is one-file evidence
On the same laptop, RFC 9110 conversion took 1.3 seconds for CleanMD, 5.5 seconds for MarkItDown and 10.9 seconds for pymupdf4llm. Those are the author’s measurements for that file and setup only; they do not establish a general speed ranking.
What this comparison cannot tell you
The test covers five text-layer PDFs and selected structural checks. It does not measure scanned-document OCR, dense three-column reading order, mathematical notation, or prose quality. A converter could score well on the reported headings and fences yet still struggle with the specific documents you need to convert.
Giacomo describes the benchmark as one that “shows failure modes, not a universal ranking.” That is the right way to use it: as evidence about these files and as a guide to what to inspect in your own.
Quick Recap
How to choose a converter for your PDFs
- Start with your document type. If you need specifications, check section depth, code blocks, grammar and running headers or footers. For books, inspect whether code examples remain coherent instead of being fragmented into single-line fences. For research papers and frameworks, compare table boundaries and check whether prose has been misclassified as a cell.
- Inspect reading order and OCR separately. This benchmark does not establish performance on scanned pages or dense multi-column layouts. Test representative pages from those files and verify that the extracted text follows the intended reading order.
- Use a known answer where possible. Compare headings, tables and code with the source PDF or an authoritative text version. A benchmark score is only as meaningful as its ground truth; pay particular attention when the expected structure was defined by a geometric rule similar to a converter’s own.
- Time your own representative files. The RFC timings do not predict which tool will be fastest on your PDFs, machines or configurations.
- Choose by the errors you can tolerate. Missing a heading, splitting a code block, inventing a table or retaining repeated footer text can have different costs depending on whether you plan to search, edit, publish or feed the Markdown into another tool.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

