Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Your converter tests can pass while a downstream parser still turns your Markdown into flat text because the tests may verify only the converter’s string—not how the production parser interprets that string. The failure is at the boundary between conversion and parsing, and its cause cannot be identified without the exact HTML, Markdown, converter settings, and parser dialect.

Why conversion tests can miss the failure

A converter test can establish that particular HTML produces an expected Markdown string. It does not establish that a separate parser will read that string as the intended headings, lists, tables, or line breaks. Even when every word survives, structural relationships may not.

Markdown parsing has both block-level structure and inline interpretation. CommonMark specifies that block structure is determined before inline parsing, so indentation, blank lines, and HTML boundaries can change how later text is grouped. The available reference is CommonMark Spec 0.26; check the dialect and version actually used in your application before relying on version-specific behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What can make Markdown look flat

Whitespace normalization and post-processing

Spaces, tabs, and newlines can carry structural meaning. The html-to-markdown Python API documentation, for example, describes a whitespace_mode option: its Normalized mode collapses consecutive whitespace, while Strict preserves source whitespace. The documentation says Normalized produces cleaner output for most documents and recommends Strict when deliberate whitespace outside <pre> matters. This is one library’s behavior, not evidence that it was used for the failure in this title.

That API also documents strip_newlines, which can produce a single-line result, and optional wrapping at word boundaries. Check those settings as well as any custom trim, whitespace collapse, serialization, or transport step between conversion and parsing.

Tabs and indentation

CommonMark defines tabs as advancing to four-column tab stops in structural contexts. Indented lines can therefore affect code blocks or list nesting; tabs inside text can behave differently. Inspect the exact characters and indentation in the Markdown passed to the parser rather than judging spacing from a rendered preview.

Soft and hard line breaks

A Markdown soft break may render as either a line ending or a space under CommonMark. A line that appears separate in source is not necessarily a hard break in the rendered result. If the content requires a visible line break, test for that behavior explicitly in the target dialect and renderer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Raw HTML mixed with Markdown

CommonMark has explicit rules for HTML blocks, and those rules differ from the original Markdown description. A block tag such as <table> or <div> can affect how nearby Markdown is parsed; spacing and indentation around the tag may matter. Do not assume pasted or converted HTML blocks will interact consistently with Markdown in every parser.

Dialect and extension differences

“Markdown” can mean CommonMark, GitHub Flavored Markdown, or a library-specific dialect. Extensions such as tables may be available only when enabled. A test against one parser or option set does not prove equivalent behavior in production. Match the downstream parser’s version, dialect, and extensions in the diagnostic test.

Trace the exact string through the production path

  1. Freeze the input. Save the original HTML fixture byte-for-byte. Record the converter name, version, settings, and intended output format.
  2. Capture the parser input. Save the exact Markdown string passed to the downstream parser. Compare it with the converter’s output to find changes caused by trimming, whitespace collapse, newline removal, wrapping, serialization, or transport.
  3. Reproduce with production settings. Feed that exact string to the exact parser version and dialect used in production. Record its syntax tree (AST) or rendered HTML—not merely whether parsing returned without an error.
  4. Minimize the example. Reduce the HTML to the smallest fixture that still flattens. If present in the real input, isolate deliberate whitespace, tabs, nested lists, line breaks in table cells, and raw block HTML rather than changing them all at once.
  5. Check parser conformance separately. The CommonMark project README says the specification contains over 500 embedded examples that serve as conformance tests. Running the applicable suite can check parser conformance; it does not prove that your converter preserves the semantics of a particular HTML document.
  6. Add a seam-level regression fixture. Keep the source HTML, expected Markdown structure, and expected parsed result together. Test conversion and downstream parsing in the same test so the handoff is covered.
  7. Change one layer at a time. Adjust converter configuration, custom post-processing, parser options, or fixture expectations individually. Do not hide a structural problem by deleting whitespace that the content needs.

Make the test assert meaning, not just text

A string-presence assertion can pass even if a heading becomes a paragraph or a nested list loses its nesting. For each fixture, assert the semantic result that matters: the expected heading level, list hierarchy, table structure, code block, or required line break in the parser’s AST or rendered output.

Keep the converter’s expected Markdown assertion too: it helps isolate whether the change happened during conversion or later. The paired checks answer different questions—what the converter emitted and what the production parser understood.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What can and cannot be concluded

Without the converter, parser, versions, settings, original HTML, emitted Markdown, and observed output, there is no basis for naming a root cause or estimating how often this failure occurs. The cited specification is CommonMark 0.26, and the converter options above come from one library’s documentation; neither identifies the components in the reported case. The practical next step is to capture and test the exact Markdown at the parser boundary.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.