To fetch a web page as Markdown or JSON, start with its URL, retrieve its HTML or render it in a browser when the page depends on JavaScript, then convert the result to the format your next step needs. Markdown is suited to readable page context; JSON is suited to named fields that downstream code can validate and consume. For a whole site, use a crawler rather than treating a single-page fetch as a crawl.
Choose the right fetch method for the page
The first decision is whether you need one known page or a collection of pages. If you already have a URL and need its content, fetch that page. If you start with a domain and need many pages, define crawl scope and use a crawler. Firecrawl documents this distinction: its Scrape endpoint is for a known URL, while Crawl starts from a domain and follows links and reads sitemaps by default. Firecrawl Scrape · Firecrawl Crawl
Direct HTTP request
Use an HTTP client followed by an HTML parser or converter when the page is accessible in its delivered HTML and you want control over parsing, cleanup, retries, and output. This approach avoids adding a browser-rendering step, but you must implement the conversion and maintain any selectors or extraction rules yourself. A basic scraper typically begins with an HTTP GET, reads the HTML, then extracts the content you need; see the O’Reilly chapter on writing a first web scraper.
Browser-rendered fetch
Use a browser engine or a service that renders the page when important content is inserted after initial HTML delivery by client-side JavaScript. You may need to wait for a selector, a page-ready condition, or a delay. Jina documents browser-engine and wait controls, while Firecrawl says its Scrape and Crawl products render pages in Chromium. Rendering can expose dynamically produced content, but it does not establish that a login-protected, region-restricted, bot-protected, or otherwise blocked page is accessible. Jina Reader documentation · Firecrawl Scrape
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Choose Markdown or schema-defined JSON
Markdown for readable context
Markdown is a practical target when a person or language model needs to read the page structure: headings, paragraphs, lists, and links. It is less verbose than raw HTML and easier to inspect than a large bundle of tags. Firecrawl documents Markdown as the default output for its Scrape endpoint. That is a product behavior, not a guarantee that every page will convert cleanly; check whether the main content, headings, and links survived extraction.
JSON for defined fields
Choose JSON when another program expects named values such as a page title, author, product name, or publication date. Define the fields and their types before extraction, then validate that the response parses and that required values are present. Firecrawl documents schema-based JSON extraction. If a field is absent or ambiguous on the source page, a syntactically valid JSON response can still be incomplete or wrong, so compare important values against the page.
Keep the source and the output connected
Do not treat successful conversion as proof of correctness. For a representative set of target pages, inspect the original rendered page alongside the result. Look for navigation and footer noise, missing dynamic sections, repeated text, stale metadata, and fields that do not match their schema. Record the source URL and, where relevant to your application, the retrieval time alongside stored output so later users can trace what was extracted.
DIY: fetch a page and convert its content
The simplest implementation depends on the page. For ordinary server-rendered HTML, a GET request plus an HTML-to-text or HTML-to-Markdown converter can work. The example below uses Python requests and Beautiful Soup to retrieve a page, isolate its main content when a <main> element exists, and produce a small readable Markdown-like result. It does not execute JavaScript and intentionally does not claim to implement a full Markdown converter.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Python: direct HTTP and basic Markdown conversion
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
url = "https://example.com/article"
response = requests.get(
url,
headers={"User-Agent": "Mozilla/5.0 (compatible; PageFetcher/1.0)"},
timeout=(10, 30),
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
root = soup.find("main") or soup.body or soup
lines = []
for node in root.find_all(["h1", "h2", "h3", "p", "li", "a"], recursive=True):
text = " ".join(node.get_text(" ", strip=True).split())
if not text:
continue
if node.name == "h1":
lines.append(f"# {text}")
elif node.name == "h2":
lines.append(f"## {text}")
elif node.name == "h3":
lines.append(f"### {text}")
elif node.name == "li":
lines.append(f"- {text}")
elif node.name == "a":
href = node.get("href")
if href:
lines.append(f"[{text}]({urljoin(url, href)})")
else:
lines.append(text)
markdown = "nn".join(lines)
print(markdown)
Install the dependencies with python -m pip install requests beautifulsoup4. This sample is a starting point rather than a general-purpose converter: nested links can appear both within their parent text and separately, and a site may use other containers than <main>. For production, choose a maintained converter or define and test a site-specific extraction rule. Respect the source site’s access terms, applicable law, and rate limits; those requirements vary by site and jurisdiction.
Convert extracted values to JSON
If your consumer needs defined fields, build a dictionary from explicit extraction rules and serialize it. For example, the following snippet assumes the earlier soup exists; missing values remain None rather than being silently invented.
import json
schema_output = {
"url": url,
"title": soup.title.get_text(" ", strip=True) if soup.title else None,
"h1": (
soup.find("h1").get_text(" ", strip=True)
if soup.find("h1") else None
),
}
print(json.dumps(schema_output, ensure_ascii=False, indent=2))
This is ordinary rule-based extraction, not a semantic guarantee. If the page has multiple possible titles or values, define which source element wins and validate the result against known examples. A JSON Schema or equivalent validation step can catch missing keys and wrong types, but cannot prove that a value faithfully represents the page.
When direct HTTP is insufficient
If the HTML response lacks the content visible in a normal browser, a plain HTTP client cannot make client-side JavaScript run. Switch to browser automation or a rendering service, and wait for a meaningful page-ready signal when possible instead of relying on an arbitrary sleep. Jina Reader exposes browser and wait controls; Firecrawl documents Chromium rendering for its scrape and crawl flows. Neither vendor documentation establishes access to every protected page, so do not assume rendering defeats a site’s access controls.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #3
Use a hosted reader, scrape API, or crawler
Hosted services reduce the amount of browser and conversion infrastructure you need to maintain, but differ in output controls, rendering options, scope, limits, and pricing. The distinctions below are documented product descriptions, not results of an independent accuracy or performance test.
| Approach | Useful when | Documented distinction | Check before adopting |
|---|---|---|---|
| Direct HTTP plus parser | A URL is reachable and you want control of extraction. | GET, HTML reading, and extraction are the basic scraper workflow described in the O’Reilly chapter. | JavaScript requirements, parser maintenance, retries, volume, and schema validation. |
| Jina Reader | You want a URL-reader workflow with configurable extraction behavior. | Its documentation describes the r.jina.ai interface, JSON response metadata, browser-engine choice, target selectors, wait selectors, page-ready controls, and cached-content options. |
Current limits, caching behavior, access behavior, and which controls your use case requires. See official documentation. |
| Firecrawl Scrape | You have a known URL and want hosted extraction. | Its product page documents Markdown as the default, schema-based JSON options, and Chromium rendering. | Current credit use, concurrency, output behavior, data handling, and plan terms. See official product page. |
| Firecrawl Crawl | You start from a domain and need pages across a site. | Its product page documents sitemap reading, recursive link following by default, path and depth controls, and per-page crawl credits. | Scope, exclusions, page count, concurrency, and total credit budget. See official product page. |
Check volatile pricing and credits
As listed on Firecrawl’s product pages on September 29, 2026, the Free plan included 1,000 credits per month; Hobby included 5,000 credits per month and was listed at $16 per month billed yearly. The Crawl page listed one credit per page crawled, with JSON mode adding four credits per page. These are changeable plan and product-page details, not a forecast of your bill; confirm current terms and calculate against your expected pages and output mode before committing. Scrape pricing details · Crawl credit details
Firecrawl also reports a P95 latency of 3,387 ms on a 1,000-URL scrape benchmark it says was run January 13, 2026. This is a company-reported benchmark, not a comparison with Jina or a general latency guarantee. The available vendor material does not establish a universal best service or independent comparative accuracy and success rates.
For full-site collection, set crawl boundaries first
A crawl can quickly expand beyond the pages you intended. Before starting one, specify which paths are in scope, how deep link-following should go, and what should be excluded. Confirm whether sitemaps and links are followed by default, and estimate page volume and credit use. Firecrawl documents sitemap reading and recursive link following by default, with path and depth controls; its stated charging basis is per crawled page, with additional credits for JSON mode. Verify the current behavior and pricing on the Crawl product page.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
- Start with a small representative scope and inspect the discovered URLs before a large run.
- Exclude account, search, filter, or other paths that do not belong in the dataset.
- Decide whether you need every page, or only a known set of URLs.
- Check rate limits and the target site’s terms and access rules before production collection.
- Store the source URL with each result so extracted facts can be checked later.
Or skip the browser setup
If the immediate job is capturing a page as an image or PDF rather than converting it into Markdown or structured fields, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is not a substitute for Markdown or JSON extraction; it is useful when the consumer needs a visual record of the rendered page.
Its GET endpoint can return a screenshot or PDF. Example cURL request, with the target URL set to Stripe:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options and response details. Cookie banners and consent prompts, newsletter popups, and chat widgets can be removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan.
Troubleshoot incomplete or unusable output
The result omits content visible in a browser
Likely cause: the content is rendered after the initial HTML response. Fix: inspect the response HTML; if the content is absent, use a browser-capable method and wait for a specific selector or readiness condition. Rendering controls help with dynamic pages but do not guarantee access to restricted pages.
The Markdown contains menus, footers, or repeated text
Likely cause: extraction included the whole document or selected a broad container. Fix: target the main-content element or a stable selector, remove known noise deliberately, and compare the result with multiple page layouts before applying the rule site-wide.
Best Value
JSON parses, but fields are missing or wrong
Likely cause: the source lacks the field, the selector is wrong, or the page has multiple competing values. Fix: validate required keys and types, retain null for absent information, define precedence rules, and check extracted values against the source page. Parsing success is not semantic validation.
A fetch times out or is blocked
Likely cause: network delay, a page that never reaches the chosen readiness condition, or a site’s access controls. Fix: set sensible connection and read timeouts, use bounded retries for transient failures, and inspect the actual response and status. Do not treat retries or browser rendering as permission to bypass site restrictions.
A crawl returns too many or too few pages
Likely cause: link-following or sitemap behavior, path scope, or depth differs from your assumptions. Fix: inspect discovered URLs on a small run, then tighten or broaden paths and depth deliberately. Recheck current crawler defaults and per-page billing before scaling.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Frequently asked questions
Can any web page be converted to Markdown or JSON?
No. A fetch can fail or return incomplete content because of JavaScript rendering, authentication, access restrictions, or site-specific structure. Conversion formats the content you can obtain; it does not make inaccessible content available.
Is JSON always more accurate than Markdown?
No. JSON makes field names and types explicit, which can help downstream validation, but the values still need to be checked against the source. Markdown is often easier for a person to review as page context.
Is there a universally best web page extraction service?
The documented options differ in controls, rendering, scale, and billing, and the available product material does not establish a universal winner. Compare them on representative URLs and the output your application actually needs.
What resource explains the broader scraping workflow?
Ryan Mitchell’s Web Scraping with Python, 3rd Edition, published by O’Reilly in February 2024, covers HTTP GET, HTML reading, extraction, APIs, and crawling. It is a broader web-scraping resource rather than a requirement for this conversion task. Read the publisher’s chapter preview.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

