Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To fetch a web page as Markdown or JSON, start with its URL, retrieve its HTML or render it in a browser when the page depends on JavaScript, then convert the result to the format your next step needs. Markdown is suited to readable page context; JSON is suited to named fields that downstream code can validate and consume. For a whole site, use a crawler rather than treating a single-page fetch as a crawl.

Choose the right fetch method for the page

The first decision is whether you need one known page or a collection of pages. If you already have a URL and need its content, fetch that page. If you start with a domain and need many pages, define crawl scope and use a crawler. Firecrawl documents this distinction: its Scrape endpoint is for a known URL, while Crawl starts from a domain and follows links and reads sitemaps by default. Firecrawl Scrape · Firecrawl Crawl

Direct HTTP request

Use an HTTP client followed by an HTML parser or converter when the page is accessible in its delivered HTML and you want control over parsing, cleanup, retries, and output. This approach avoids adding a browser-rendering step, but you must implement the conversion and maintain any selectors or extraction rules yourself. A basic scraper typically begins with an HTTP GET, reads the HTML, then extracts the content you need; see the O’Reilly chapter on writing a first web scraper.

Browser-rendered fetch

Use a browser engine or a service that renders the page when important content is inserted after initial HTML delivery by client-side JavaScript. You may need to wait for a selector, a page-ready condition, or a delay. Jina documents browser-engine and wait controls, while Firecrawl says its Scrape and Crawl products render pages in Chromium. Rendering can expose dynamically produced content, but it does not establish that a login-protected, region-restricted, bot-protected, or otherwise blocked page is accessible. Jina Reader documentation · Firecrawl Scrape

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose Markdown or schema-defined JSON

Markdown for readable context

Markdown is a practical target when a person or language model needs to read the page structure: headings, paragraphs, lists, and links. It is less verbose than raw HTML and easier to inspect than a large bundle of tags. Firecrawl documents Markdown as the default output for its Scrape endpoint. That is a product behavior, not a guarantee that every page will convert cleanly; check whether the main content, headings, and links survived extraction.

JSON for defined fields

Choose JSON when another program expects named values such as a page title, author, product name, or publication date. Define the fields and their types before extraction, then validate that the response parses and that required values are present. Firecrawl documents schema-based JSON extraction. If a field is absent or ambiguous on the source page, a syntactically valid JSON response can still be incomplete or wrong, so compare important values against the page.

Keep the source and the output connected

Do not treat successful conversion as proof of correctness. For a representative set of target pages, inspect the original rendered page alongside the result. Look for navigation and footer noise, missing dynamic sections, repeated text, stale metadata, and fields that do not match their schema. Record the source URL and, where relevant to your application, the retrieval time alongside stored output so later users can trace what was extracted.

DIY: fetch a page and convert its content

The simplest implementation depends on the page. For ordinary server-rendered HTML, a GET request plus an HTML-to-text or HTML-to-Markdown converter can work. The example below uses Python requests and Beautiful Soup to retrieve a page, isolate its main content when a <main> element exists, and produce a small readable Markdown-like result. It does not execute JavaScript and intentionally does not claim to implement a full Markdown converter.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Python: direct HTTP and basic Markdown conversion

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

url = "https://example.com/article"
response = requests.get(
    url,
    headers={"User-Agent": "Mozilla/5.0 (compatible; PageFetcher/1.0)"},
    timeout=(10, 30),
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
root = soup.find("main") or soup.body or soup

lines = []
for node in root.find_all(["h1", "h2", "h3", "p", "li", "a"], recursive=True):
    text = " ".join(node.get_text(" ", strip=True).split())
    if not text:
        continue
    if node.name == "h1":
        lines.append(f"# {text}")
    elif node.name == "h2":
        lines.append(f"## {text}")
    elif node.name == "h3":
        lines.append(f"### {text}")
    elif node.name == "li":
        lines.append(f"- {text}")
    elif node.name == "a":
        href = node.get("href")
        if href:
            lines.append(f"[{text}]({urljoin(url, href)})")
    else:
        lines.append(text)

markdown = "nn".join(lines)
print(markdown)

Install the dependencies with python -m pip install requests beautifulsoup4. This sample is a starting point rather than a general-purpose converter: nested links can appear both within their parent text and separately, and a site may use other containers than <main>. For production, choose a maintained converter or define and test a site-specific extraction rule. Respect the source site’s access terms, applicable law, and rate limits; those requirements vary by site and jurisdiction.

Convert extracted values to JSON

If your consumer needs defined fields, build a dictionary from explicit extraction rules and serialize it. For example, the following snippet assumes the earlier soup exists; missing values remain None rather than being silently invented.

import json

schema_output = {
    "url": url,
    "title": soup.title.get_text(" ", strip=True) if soup.title else None,
    "h1": (
        soup.find("h1").get_text(" ", strip=True)
        if soup.find("h1") else None
    ),
}
print(json.dumps(schema_output, ensure_ascii=False, indent=2))

This is ordinary rule-based extraction, not a semantic guarantee. If the page has multiple possible titles or values, define which source element wins and validate the result against known examples. A JSON Schema or equivalent validation step can catch missing keys and wrong types, but cannot prove that a value faithfully represents the page.

When direct HTTP is insufficient

If the HTML response lacks the content visible in a normal browser, a plain HTTP client cannot make client-side JavaScript run. Switch to browser automation or a rendering service, and wait for a meaningful page-ready signal when possible instead of relying on an arbitrary sleep. Jina Reader exposes browser and wait controls; Firecrawl documents Chromium rendering for its scrape and crawl flows. Neither vendor documentation establishes access to every protected page, so do not assume rendering defeats a site’s access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a hosted reader, scrape API, or crawler

Hosted services reduce the amount of browser and conversion infrastructure you need to maintain, but differ in output controls, rendering options, scope, limits, and pricing. The distinctions below are documented product descriptions, not results of an independent accuracy or performance test.

Approach Useful when Documented distinction Check before adopting
Direct HTTP plus parser A URL is reachable and you want control of extraction. GET, HTML reading, and extraction are the basic scraper workflow described in the O’Reilly chapter. JavaScript requirements, parser maintenance, retries, volume, and schema validation.
Jina Reader You want a URL-reader workflow with configurable extraction behavior. Its documentation describes the r.jina.ai interface, JSON response metadata, browser-engine choice, target selectors, wait selectors, page-ready controls, and cached-content options. Current limits, caching behavior, access behavior, and which controls your use case requires. See official documentation.
Firecrawl Scrape You have a known URL and want hosted extraction. Its product page documents Markdown as the default, schema-based JSON options, and Chromium rendering. Current credit use, concurrency, output behavior, data handling, and plan terms. See official product page.
Firecrawl Crawl You start from a domain and need pages across a site. Its product page documents sitemap reading, recursive link following by default, path and depth controls, and per-page crawl credits. Scope, exclusions, page count, concurrency, and total credit budget. See official product page.

Check volatile pricing and credits

As listed on Firecrawl’s product pages on September 29, 2026, the Free plan included 1,000 credits per month; Hobby included 5,000 credits per month and was listed at $16 per month billed yearly. The Crawl page listed one credit per page crawled, with JSON mode adding four credits per page. These are changeable plan and product-page details, not a forecast of your bill; confirm current terms and calculate against your expected pages and output mode before committing. Scrape pricing details · Crawl credit details

Firecrawl also reports a P95 latency of 3,387 ms on a 1,000-URL scrape benchmark it says was run January 13, 2026. This is a company-reported benchmark, not a comparison with Jina or a general latency guarantee. The available vendor material does not establish a universal best service or independent comparative accuracy and success rates.

For full-site collection, set crawl boundaries first

A crawl can quickly expand beyond the pages you intended. Before starting one, specify which paths are in scope, how deep link-following should go, and what should be excluded. Confirm whether sitemaps and links are followed by default, and estimate page volume and credit use. Firecrawl documents sitemap reading and recursive link following by default, with path and depth controls; its stated charging basis is per crawled page, with additional credits for JSON mode. Verify the current behavior and pricing on the Crawl product page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
  • Start with a small representative scope and inspect the discovered URLs before a large run.
  • Exclude account, search, filter, or other paths that do not belong in the dataset.
  • Decide whether you need every page, or only a known set of URLs.
  • Check rate limits and the target site’s terms and access rules before production collection.
  • Store the source URL with each result so extracted facts can be checked later.

Or skip the browser setup

If the immediate job is capturing a page as an image or PDF rather than converting it into Markdown or structured fields, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is not a substitute for Markdown or JSON extraction; it is useful when the consumer needs a visual record of the rendered page.

Its GET endpoint can return a screenshot or PDF. Example cURL request, with the target URL set to Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options and response details. Cookie banners and consent prompts, newsletter popups, and chat widgets can be removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot incomplete or unusable output

The result omits content visible in a browser

Likely cause: the content is rendered after the initial HTML response. Fix: inspect the response HTML; if the content is absent, use a browser-capable method and wait for a specific selector or readiness condition. Rendering controls help with dynamic pages but do not guarantee access to restricted pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Markdown contains menus, footers, or repeated text

Likely cause: extraction included the whole document or selected a broad container. Fix: target the main-content element or a stable selector, remove known noise deliberately, and compare the result with multiple page layouts before applying the rule site-wide.

JSON parses, but fields are missing or wrong

Likely cause: the source lacks the field, the selector is wrong, or the page has multiple competing values. Fix: validate required keys and types, retain null for absent information, define precedence rules, and check extracted values against the source page. Parsing success is not semantic validation.

A fetch times out or is blocked

Likely cause: network delay, a page that never reaches the chosen readiness condition, or a site’s access controls. Fix: set sensible connection and read timeouts, use bounded retries for transient failures, and inspect the actual response and status. Do not treat retries or browser rendering as permission to bypass site restrictions.

A crawl returns too many or too few pages

Likely cause: link-following or sitemap behavior, path scope, or depth differs from your assumptions. Fix: inspect discovered URLs on a small run, then tighten or broaden paths and depth deliberately. Recheck current crawler defaults and per-page billing before scaling.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently asked questions

Can any web page be converted to Markdown or JSON?

No. A fetch can fail or return incomplete content because of JavaScript rendering, authentication, access restrictions, or site-specific structure. Conversion formats the content you can obtain; it does not make inaccessible content available.

Is JSON always more accurate than Markdown?

No. JSON makes field names and types explicit, which can help downstream validation, but the values still need to be checked against the source. Markdown is often easier for a person to review as page context.

Is there a universally best web page extraction service?

The documented options differ in controls, rendering, scale, and billing, and the available product material does not establish a universal winner. Compare them on representative URLs and the output your application actually needs.

What resource explains the broader scraping workflow?

Ryan Mitchell’s Web Scraping with Python, 3rd Edition, published by O’Reilly in February 2024, covers HTTP GET, HTML reading, extraction, APIs, and crawling. It is a broader web-scraping resource rather than a requirement for this conversion task. Read the publisher’s chapter preview.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.