Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

MarkItDown converts files such as PDFs, DOCX documents, and XLSX workbooks into Markdown through a Python API or command-line tool. It is useful for text analysis and indexing, but its output is not a faithful visual copy—and a successful conversion does not guarantee that every piece of a PDF was extracted. A reported edge case describes text disappearing after a specially encoded inline image.

What MarkItDown does—and what it does not

The MarkItDown project describes the package as a lightweight Python utility for converting files to Markdown for use with LLMs and related text-analysis pipelines. Its official README lists formats including PDF, PowerPoint, Word, Excel, images, audio, HTML, CSV, JSON, XML, ZIP contents, YouTube URLs, and EPUB.

Markdown is a text representation, not a reconstruction of the source file. Tables, page layout, images, and embedded text may be represented differently or omitted depending on the input and converter path. Treat the result as extracted content to inspect, not as proof that the original document’s appearance or meaning has been preserved.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install support for the formats you need

The project README specifies Python 3.10 through 3.14 and recommends using a virtual environment. Format converters rely on optional dependencies, so a base installation may not support every format. Install all optional format groups for broad coverage:

python -m venv .venv
source .venv/bin/activate
pip install 'markitdown[all]'

For a targeted setup with PDF, DOCX, and XLSX support, install those extras instead:

pip install 'markitdown[pdf,docx,xlsx]'

According to the project’s package metadata, the PDF extra uses pdfminer.six and pdfplumber; DOCX uses Mammoth and lxml; and XLSX uses pandas and openpyxl. Optional dependencies explain why installing only the base package may not provide the converter you expect.

Convert a file with Python or the command line

Python API

Call convert() with a file path, then read the Markdown from the returned result’s markdown attribute:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from markitdown import MarkItDown

md = MarkItDown()
result = md.convert("report.pdf")
print(result.markdown)

Change the path to a DOCX, XLSX, or another supported input for the corresponding converter.

Command-line interface

The documented CLI pattern writes the conversion output to a Markdown file:

markitdown report.pdf > report.md

Use the appropriate installed extra for the input format before relying on the command.

What the reported silent PDF failure means

MarkItDown issue #1870, opened May 9, 2026, reports a specific case in which text after an inline image was missing from extracted PDF text. The reported content stream encoded the inline image with BI ... ID ... EI, ASCII85 and Flate filters, and a bare ~ terminator. The reporter said both the pdfplumber and pdfminer extraction paths returned text before the image but did not surface the text after it, making the conversion appear successful despite partial output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The issue describes a synthetic reproduction and a real-world invoice. Its reported environment was MarkItDown commit 4b65609 (May 7, 2026), pdfplumber 0.11.9, pdfminer.six 20251230, PyMuPDF 1.27.2.3, macOS 15.6, and Python 3.13. This is an individual, version-specific issue report—not evidence that PDFs generally fail or that every current installation has the problem. The issue’s existence alone does not establish whether a fix has since shipped.

The practical implication is narrow but important: nonempty Markdown is not a completeness check. For important PDFs, check that expected headings, values, and text after embedded images appear in the output; where useful, compare page counts or other known markers with the source. These are validation practices, not a built-in MarkItDown completeness feature.

OCR is separate, and it can also omit content

The optional markitdown-ocr plugin documents LLM vision OCR for images embedded in PDF, DOCX, PPTX, and XLSX files. Its Python example enables plugins and supplies an LLM client and model; installing or enabling the plugin alone is not enough to ensure OCR runs.

The plugin documentation says OCR is silently skipped if no llm_client is supplied. If an LLM call fails, conversion continues without that image’s text. For scanned PDFs, the README describes automatic detection and full-page rendering at 300 DPI for pages with no extractable text, plus PyMuPDF rendering as a recovery path for malformed PDFs. These are documented plugin behaviors, not a guarantee that every image or PDF edge case will be handled. The inline-image issue above concerns a different core PDF extraction path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Inspect the output before using it downstream

For routine text analysis, a quick read-through may be enough. For ingestion into search, automation, or a high-stakes workflow, validate the fields that matter rather than relying on whether conversion returned text.

  • Check for expected headings, page markers, totals, identifiers, or other known values.
  • For PDFs containing inline images, compare text before and after those images against the original.
  • For scans or image-heavy files, confirm OCR is configured with the required client and model, then verify that the expected image text appears.
  • For spreadsheets, inspect table headers and representative rows to catch empty or malformed output.

A separate report, issue #2136, opened June 16, 2026, describes a CSV with a blank first line becoming a Markdown table with empty cells and no warning. The reported setup was MarkItDown 0.1.6 and Python 3.12; the reporter said the blank first row was treated as the header, leaving zero columns for subsequent rows. This is a distinct CSV report, not the PDF defect, and it likewise illustrates why output should be checked against the input.

Issue reports capture particular observations and can change status. For a version-specific decision, check the relevant issue and current project release information rather than assuming that a reported behavior remains unresolved.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.