Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable method is a three-stage pipeline: fetch or render the page, extract the content you actually need, then convert that HTML or DOM to Markdown. A normal HTTP request is sufficient when the response already contains the text. For an SPA shell or a client-rendered route, run the page in a real browser (such as Playwright), wait for the relevant content, select its container, and pass that HTML to an HTML-to-Markdown converter such as Turndown.

Why a successful HTTP request can still produce empty Markdown

An HTTP client receives the server’s initial response. Browsers then parse HTML, build a DOM and CSSOM, run JavaScript, and update the DOM as the application starts. In a single-page application, the initial document may contain little more than a root element and script tags. The article, table, or product data you see in a browser can therefore be absent from the response you downloaded.

HTML-to-Markdown libraries solve a different problem. Turndown accepts an HTML string or DOM node and serializes elements such as headings, links, lists, tables, and emphasis as Markdown; it does not execute application JavaScript or classify the page’s main content. Rendering and conversion must be treated as separate stages.

Choose the pipeline that matches the page

Approach Use it when Trade-off
Static fetch plus converter The response already contains the text you need Fast and simple, but an SPA shell can yield empty or incomplete Markdown.
Browser render, extraction, then converter The route depends on JavaScript, interaction, authentication, or deferred loading More setup and a page-specific readiness strategy, but it exposes the post-JavaScript DOM.
Hosted rendering service You want one service to run a browser and return processed output Less infrastructure to operate; coverage, quality, limits, and pricing must be checked for your workload.

Static-first with browser fallback is a practical pattern: inspect the response before paying the cost of launching Chromium. A hosted service such as Firecrawl describes browser rendering and Markdown output, while Microlink documents SPA rendering and readiness controls. Those are vendor descriptions, not independent quality or performance comparisons.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stage 1: Try a static fetch and prove whether content is present

Do not infer completeness from a 200 OK. Search the returned body for a distinctive heading, paragraph, or selector that should exist in the final page.

import requests

url = "https://example.com/article"
r = requests.get(url, timeout=30, headers={"User-Agent": "markdown-extractor/1.0"})
r.raise_for_status()
html = r.text

marker = "Expected article heading"
if marker in html:
    print("Useful content is in the initial response")
else:
    print("Likely client-rendered; use a browser fallback")

If static content is present, select the article region and convert it. A converter applied to the entire document commonly preserves navigation, cookie notices, footers, and repeated interface text.

Static conversion with Node.js

npm install turndown cheerio
import fs from "node:fs/promises";
import TurndownService from "turndown";
import * as cheerio from "cheerio";

const html = await (await fetch("https://example.com/article")).text();
const $ = cheerio.load(html);
const region = $("article").first().html() || $("main").first().html();
if (!region) throw new Error("No article or main element in initial HTML");

const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" });
await fs.writeFile("article.md", turndown.turndown(region), "utf8");

Stage 2: Render an SPA with Playwright

Install Playwright and its browser once in the environment that will run the extractor.

npm install playwright turndown
npx playwright install chromium

The following script navigates to a URL, waits for a content-specific condition, extracts the smallest useful region, and converts it. Replace the URL, selector, and readiness test with values for the target site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import fs from "node:fs/promises";
import { chromium } from "playwright";
import TurndownService from "turndown";

const url = "https://example.com/app/article/42";
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({
    viewport: { width: 1440, height: 1000 },
    userAgent: "markdown-extractor/1.0"
  });

  await page.goto(url, { waitUntil: "domcontentloaded", timeout: 60000 });
  await page.locator("article").waitFor({ state: "visible", timeout: 30000 });

  // Optional: trigger lazy content only when this page requires it.
  await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
  await page.waitForTimeout(500);

  const html = await page.locator("article").first().innerHTML();
  const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" });
  await fs.writeFile("article.md", turndown.turndown(html), "utf8");
} finally {
  await browser.close();
}

domcontentloaded only means the document was parsed. It is not proof that an SPA has finished fetching data. Waiting for the expected article selector, a known heading, or another page-specific condition is safer. There is no universal selector or timeout that works for every application.

When the page needs interaction

  • Click a “Load more” control before reading innerHTML.
  • Choose a tab or expand an accordion if that state is part of the material you need.
  • Supply authentication through a stored browser context, cookies, or a login flow when permitted by the site.
  • Scroll only when lazy loading is known to depend on viewport exposure. Scrolling is not a guarantee that every deferred resource has loaded.

Keep the interaction steps narrowly scoped. Clicking arbitrary controls can change the content or submit forms, and automated access must comply with the site's terms and access controls.

Extract the right DOM region before converting

Prefer a semantic container such as article or a site-specific content selector. If none exists, identify a stable class or data attribute and remove obvious chrome before conversion.

const content = await page.locator("[data-testid='article-body']").first().innerHTML();

Extraction improves signal, but it can also remove context. Check that the selected region still contains the title, headings, links, lists, tables, images, and code blocks you need. Hosted products advertise automatic content cleanup, but cleanup quality varies by page and has not been established by an independent comparison here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preserve structure and handle difficult elements

Headings, links, lists, and code

Turndown generally maps heading levels, anchor URLs, ordered and unordered lists, block quotes, and fenced code. Inspect the result for nested lists and relative links. If links must work outside the source site, resolve relative URLs against the page URL before conversion or post-process the Markdown.

Tables

Markdown tables cannot represent every HTML table feature. Verify merged cells, nested markup, and very wide tables. For complex data, retain an HTML table alongside the Markdown or export the underlying data separately.

Images and lazy resources

Markdown usually retains an image URL, not the image bytes. Make sure the final src is the loaded URL rather than a placeholder such as data-src. If the application swaps attributes after intersection with the viewport, scroll or trigger the component before extraction.

Shadow DOM and iframes

Content inside an iframe belongs to another document; locate the frame and extract from its frame context. Shadow-root content may require querying inside the component rather than reading the light DOM. A converter cannot recover content you never extracted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A production-ready static-first fallback

  1. Request the page with a reasonable timeout and identify a marker that proves the required content exists.
  2. If the marker is present, parse and convert the selected region without launching a browser.
  3. If it is absent, launch Playwright and navigate to the same URL.
  4. Wait for a page-specific selector or text condition, then perform only the interactions needed for deferred content.
  5. Extract the content region and convert it with the same Turndown configuration.
  6. Save the rendered HTML when debugging so you can distinguish a rendering failure from a conversion failure.

This design keeps ordinary pages inexpensive while still handling client-rendered routes. It also makes failures observable: you can compare the initial response, the post-render DOM, and the final Markdown.

Troubleshooting empty or incorrect output

Symptom Likely cause Fix
Markdown contains only a root element or script tags You converted the SPA shell Render with Playwright, then extract after the app populates the DOM.
Navigation and cookie text dominate the file You converted the whole document Select the article or main-content container first.
Selector timeout The selector is wrong, the route failed, or content is behind a login/state change Capture a screenshot and rendered HTML, inspect the DOM, and replace the selector or authentication flow.
Content is visible manually but absent in automation Different viewport, user agent, cookies, geolocation, or bot challenge Reproduce the required context, and do not attempt to bypass access controls.
Some paragraphs or images are missing Lazy loading or an interaction has not occurred Scroll or click the specific control, wait for the resulting content, and re-extract.
Links or tables are malformed Converter limitations or unusual HTML Inspect the selected HTML, configure Turndown rules, or preserve problematic fragments as HTML.
Navigation succeeds but data never appears Background API call failed or readiness was assumed too early Wait for a content condition, inspect network/application errors, and increase timeout only after diagnosing the cause.

Performance, reliability, and cost decisions

  • Use static fetches for static pages. They avoid browser startup and are easier to scale.
  • Reuse a browser process. For batches, keep Chromium running and create isolated contexts or pages per job.
  • Bound every wait. Set navigation and selector timeouts, record the URL and failure stage, and retry only transient failures.
  • Cache where appropriate. A rendered DOM can become stale; choose a cache policy that matches how often the source changes.
  • Control concurrency. Too many pages can exhaust CPU, memory, or the target site's acceptable request rate.
  • Retain diagnostics. Store response status, final URL, rendered HTML, console errors, and a screenshot for failed jobs.

Hosted services trade browser operations for a service call. Evaluate JavaScript support, readiness controls, extraction quality, interaction and authentication handling, raw-HTML access, deployment, limits, and total cost for your workload. Vendor feature pages are not substitutes for a workload-specific quality test.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is useful when you need a rendered visual capture rather than Markdown text. It runs the browser-side capture for a URL and can return PNG, JPEG, WebP, or PDF; it does not replace the extraction-and-Turndown stage for a Markdown pipeline.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for parameters. Equivalent Python and Node.js calls are:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

FAQ

Can Turndown render a React or Vue page?

No. It converts HTML that already exists. Render the route first, then pass the resulting DOM or HTML to Turndown.

Is waiting for networkidle always correct?

No. Applications with analytics, websockets, or polling may never become idle. A condition tied to the content you need is usually more meaningful.

Should I convert the entire page or only body?

Usually neither. Extract the narrowest stable content region that contains the material you intend to publish or analyze.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does the browser show different content from my script?

Compare viewport, cookies, user agent, authentication, locale, geolocation, and bot-check behavior. Any of these can change the rendered DOM.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.