The reliable method is a three-stage pipeline: fetch or render the page, extract the content you actually need, then convert that HTML or DOM to Markdown. A normal HTTP request is sufficient when the response already contains the text. For an SPA shell or a client-rendered route, run the page in a real browser (such as Playwright), wait for the relevant content, select its container, and pass that HTML to an HTML-to-Markdown converter such as Turndown.
Why a successful HTTP request can still produce empty Markdown
An HTTP client receives the server’s initial response. Browsers then parse HTML, build a DOM and CSSOM, run JavaScript, and update the DOM as the application starts. In a single-page application, the initial document may contain little more than a root element and script tags. The article, table, or product data you see in a browser can therefore be absent from the response you downloaded.
HTML-to-Markdown libraries solve a different problem. Turndown accepts an HTML string or DOM node and serializes elements such as headings, links, lists, tables, and emphasis as Markdown; it does not execute application JavaScript or classify the page’s main content. Rendering and conversion must be treated as separate stages.
Choose the pipeline that matches the page
| Approach | Use it when | Trade-off |
|---|---|---|
| Static fetch plus converter | The response already contains the text you need | Fast and simple, but an SPA shell can yield empty or incomplete Markdown. |
| Browser render, extraction, then converter | The route depends on JavaScript, interaction, authentication, or deferred loading | More setup and a page-specific readiness strategy, but it exposes the post-JavaScript DOM. |
| Hosted rendering service | You want one service to run a browser and return processed output | Less infrastructure to operate; coverage, quality, limits, and pricing must be checked for your workload. |
Static-first with browser fallback is a practical pattern: inspect the response before paying the cost of launching Chromium. A hosted service such as Firecrawl describes browser rendering and Markdown output, while Microlink documents SPA rendering and readiness controls. Those are vendor descriptions, not independent quality or performance comparisons.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Stage 1: Try a static fetch and prove whether content is present
Do not infer completeness from a 200 OK. Search the returned body for a distinctive heading, paragraph, or selector that should exist in the final page.
import requests
url = "https://example.com/article"
r = requests.get(url, timeout=30, headers={"User-Agent": "markdown-extractor/1.0"})
r.raise_for_status()
html = r.text
marker = "Expected article heading"
if marker in html:
print("Useful content is in the initial response")
else:
print("Likely client-rendered; use a browser fallback")
If static content is present, select the article region and convert it. A converter applied to the entire document commonly preserves navigation, cookie notices, footers, and repeated interface text.
Static conversion with Node.js
npm install turndown cheerio
import fs from "node:fs/promises";
import TurndownService from "turndown";
import * as cheerio from "cheerio";
const html = await (await fetch("https://example.com/article")).text();
const $ = cheerio.load(html);
const region = $("article").first().html() || $("main").first().html();
if (!region) throw new Error("No article or main element in initial HTML");
const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" });
await fs.writeFile("article.md", turndown.turndown(region), "utf8");
Stage 2: Render an SPA with Playwright
Install Playwright and its browser once in the environment that will run the extractor.
npm install playwright turndown
npx playwright install chromium
The following script navigates to a URL, waits for a content-specific condition, extracts the smallest useful region, and converts it. Replace the URL, selector, and readiness test with values for the target site.
Recommended Free Tools
Rank #2
import fs from "node:fs/promises";
import { chromium } from "playwright";
import TurndownService from "turndown";
const url = "https://example.com/app/article/42";
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
userAgent: "markdown-extractor/1.0"
});
await page.goto(url, { waitUntil: "domcontentloaded", timeout: 60000 });
await page.locator("article").waitFor({ state: "visible", timeout: 30000 });
// Optional: trigger lazy content only when this page requires it.
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
await page.waitForTimeout(500);
const html = await page.locator("article").first().innerHTML();
const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" });
await fs.writeFile("article.md", turndown.turndown(html), "utf8");
} finally {
await browser.close();
}
domcontentloaded only means the document was parsed. It is not proof that an SPA has finished fetching data. Waiting for the expected article selector, a known heading, or another page-specific condition is safer. There is no universal selector or timeout that works for every application.
When the page needs interaction
- Click a “Load more” control before reading
innerHTML. - Choose a tab or expand an accordion if that state is part of the material you need.
- Supply authentication through a stored browser context, cookies, or a login flow when permitted by the site.
- Scroll only when lazy loading is known to depend on viewport exposure. Scrolling is not a guarantee that every deferred resource has loaded.
Keep the interaction steps narrowly scoped. Clicking arbitrary controls can change the content or submit forms, and automated access must comply with the site's terms and access controls.
Extract the right DOM region before converting
Prefer a semantic container such as article or a site-specific content selector. If none exists, identify a stable class or data attribute and remove obvious chrome before conversion.
const content = await page.locator("[data-testid='article-body']").first().innerHTML();
Extraction improves signal, but it can also remove context. Check that the selected region still contains the title, headings, links, lists, tables, images, and code blocks you need. Hosted products advertise automatic content cleanup, but cleanup quality varies by page and has not been established by an independent comparison here.
Preserve structure and handle difficult elements
Headings, links, lists, and code
Turndown generally maps heading levels, anchor URLs, ordered and unordered lists, block quotes, and fenced code. Inspect the result for nested lists and relative links. If links must work outside the source site, resolve relative URLs against the page URL before conversion or post-process the Markdown.
Tables
Markdown tables cannot represent every HTML table feature. Verify merged cells, nested markup, and very wide tables. For complex data, retain an HTML table alongside the Markdown or export the underlying data separately.
Images and lazy resources
Markdown usually retains an image URL, not the image bytes. Make sure the final src is the loaded URL rather than a placeholder such as data-src. If the application swaps attributes after intersection with the viewport, scroll or trigger the component before extraction.
Shadow DOM and iframes
Content inside an iframe belongs to another document; locate the frame and extract from its frame context. Shadow-root content may require querying inside the component rather than reading the light DOM. A converter cannot recover content you never extracted.
Rank #4
A production-ready static-first fallback
- Request the page with a reasonable timeout and identify a marker that proves the required content exists.
- If the marker is present, parse and convert the selected region without launching a browser.
- If it is absent, launch Playwright and navigate to the same URL.
- Wait for a page-specific selector or text condition, then perform only the interactions needed for deferred content.
- Extract the content region and convert it with the same Turndown configuration.
- Save the rendered HTML when debugging so you can distinguish a rendering failure from a conversion failure.
This design keeps ordinary pages inexpensive while still handling client-rendered routes. It also makes failures observable: you can compare the initial response, the post-render DOM, and the final Markdown.
Troubleshooting empty or incorrect output
| Symptom | Likely cause | Fix |
|---|---|---|
| Markdown contains only a root element or script tags | You converted the SPA shell | Render with Playwright, then extract after the app populates the DOM. |
| Navigation and cookie text dominate the file | You converted the whole document | Select the article or main-content container first. |
| Selector timeout | The selector is wrong, the route failed, or content is behind a login/state change | Capture a screenshot and rendered HTML, inspect the DOM, and replace the selector or authentication flow. |
| Content is visible manually but absent in automation | Different viewport, user agent, cookies, geolocation, or bot challenge | Reproduce the required context, and do not attempt to bypass access controls. |
| Some paragraphs or images are missing | Lazy loading or an interaction has not occurred | Scroll or click the specific control, wait for the resulting content, and re-extract. |
| Links or tables are malformed | Converter limitations or unusual HTML | Inspect the selected HTML, configure Turndown rules, or preserve problematic fragments as HTML. |
| Navigation succeeds but data never appears | Background API call failed or readiness was assumed too early | Wait for a content condition, inspect network/application errors, and increase timeout only after diagnosing the cause. |
Performance, reliability, and cost decisions
- Use static fetches for static pages. They avoid browser startup and are easier to scale.
- Reuse a browser process. For batches, keep Chromium running and create isolated contexts or pages per job.
- Bound every wait. Set navigation and selector timeouts, record the URL and failure stage, and retry only transient failures.
- Cache where appropriate. A rendered DOM can become stale; choose a cache policy that matches how often the source changes.
- Control concurrency. Too many pages can exhaust CPU, memory, or the target site's acceptable request rate.
- Retain diagnostics. Store response status, final URL, rendered HTML, console errors, and a screenshot for failed jobs.
Hosted services trade browser operations for a service call. Evaluate JavaScript support, readiness controls, extraction quality, interaction and authentication handling, raw-HTML access, deployment, limits, and total cost for your workload. Vendor feature pages are not substitutes for a workload-specific quality test.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is useful when you need a rendered visual capture rather than Markdown text. It runs the browser-side capture for a URL and can return PNG, JPEG, WebP, or PDF; it does not replace the extraction-and-Turndown stage for a Markdown pipeline.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for parameters. Equivalent Python and Node.js calls are:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
FAQ
Can Turndown render a React or Vue page?
No. It converts HTML that already exists. Render the route first, then pass the resulting DOM or HTML to Turndown.
Is waiting for networkidle always correct?
No. Applications with analytics, websockets, or polling may never become idle. A condition tied to the content you need is usually more meaningful.
Should I convert the entire page or only body?
Usually neither. Extract the narrowest stable content region that contains the material you intend to publish or analyze.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhy does the browser show different content from my script?
Compare viewport, cookies, user agent, authentication, locale, geolocation, and bot-check behavior. Any of these can change the rendered DOM.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

