Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract a page’s HTML, make an HTTP GET request to its URL and save the response body. For a quick check, use a browser’s View Source command; for repeatable work, use curl, Wget, or Python Requests. The result is the HTML sent in the response—not necessarily the live page you see after JavaScript runs.

Choose the right kind of page source

There are two common meanings of “the HTML from a URL,” and choosing the right one prevents a lot of confusion:

  • Original response HTML: the body returned by the server for the document request. Use View Source or an HTTP client such as curl when you want to inspect what the server sent.
  • Rendered DOM: the current document structure after the browser has parsed the response, executed scripts, and possibly loaded more data. Use developer tools or a browser automation workflow when you need what the page became after loading.

A simple GET request retrieves the document identified by a URL; curl’s documentation describes GET as returning the document body. curl documentation The browser’s Elements panel shows the current DOM, which can differ from the original response when scripts alter the page or fetch content later. Scrapy: Dynamic Content

Extract HTML in a browser

View the original response

Open the page, then use the browser’s View Source command—often available by right-clicking the page or through the browser menu. This displays the document response as source text. It is the fastest option for a one-off inspection, but it is not a convenient way to process many URLs or save structured data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Inspect the live DOM

Open developer tools and select the Elements panel. This shows the browser’s current DOM, not necessarily the exact response body. If a value appears here but not in View Source, it may have been inserted by JavaScript or fetched after the initial document request.

Find data loaded separately

In developer tools, open the Network panel, reload the page, and look for XHR or Fetch requests. Inspect the request and response that contain the data you need. Scrapy’s guidance recommends locating the source of the data and reproducing the browser request, including its method, URL, headers, and body where needed. Scrapy: Dynamic Content If your browser offers an option to copy or export a request as cURL, use that as a starting point rather than guessing at the request.

Download HTML with curl or Wget

curl

Follow redirects and write the response body to a file:

curl -L "https://example.com" -o page.html

The -L option follows redirects; -o names the output file. To inspect response headers alongside the body, use -i. To request headers only, use -I; a HEAD request returns headers without the response body. See the curl manual for option details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wget

Save a page under a chosen filename with:

wget -O page.html "https://example.com"

Wget can also retrieve linked resources in recursive mode. Recursion can turn one page into a much larger crawl, so set an appropriate depth, domain boundary, and output directory before using it. Consult the GNU Wget manual for recursive retrieval options.

Check what you downloaded

Open page.html in a text editor. If the file contains an error message, login screen, JSON, or a bot-check page, the request did not produce the target HTML even if a file was saved successfully. Check the HTTP status and content type before treating the response as the page you wanted.

Fetch and save a URL with Python Requests

Install Requests if it is not already available: python -m pip install requests. Then fetch the URL, fail visibly for HTTP errors, and save the decoded response text:

import requests

url = "https://example.com"
r = requests.get(url, timeout=20)
r.raise_for_status()
html = r.text
print(html)

with open("page.html", "w", encoding=r.encoding or "utf-8") as f:
    f.write(html)

Requests exposes decoded response text through r.text, raw response bytes through r.content, and metadata such as headers through r.headers. It supports redirects, cookies, SSL verification, and timeouts; its documentation explains these behaviors. Requests Quickstart Use r.content instead of r.text when you need the original bytes or plan to handle decoding yourself. A timeout prevents a stalled request from waiting indefinitely, while raise_for_status() ensures a 4xx or 5xx response is not silently treated as a successful fetch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse the downloaded HTML with Beautiful Soup

Fetching and parsing are separate tasks: Requests obtains the response; Beautiful Soup turns the markup into a navigable tree. Install it with python -m pip install beautifulsoup4, then parse the saved response text:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
print(soup.title.get_text(strip=True) if soup.title else "No title")
for link in soup.select("a[href]"):
    print(link.get("href"))

html.parser is included with Python. Beautiful Soup can also use lxml or html5lib when those packages are installed. Malformed markup may produce different trees with different parsers, so specify the parser if you need repeatable results. Beautiful Soup documentation

When the HTML is not the content you expected

JavaScript adds the content later

A direct HTTP fetch does not run page JavaScript. If the initial HTML lacks text that appears in the browser, inspect the Network panel for the later request that supplies it. Reproducing that request is often simpler and more stable than rendering the whole page. If there is no practical request to reproduce, use a browser-based renderer that executes JavaScript before you extract the content. Scrapy discusses both identifying the source request and handling dynamic content. Scrapy: Dynamic Content

The server responds differently to your script

Some pages require cookies, authentication, a particular user agent, or other headers. Compare the browser’s document request with your script’s request and reproduce only the necessary details. Do this only for content you are authorized to access; do not treat a login or access control as a technical obstacle to bypass.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

The response encoding looks wrong

Requests decodes response text automatically, using response information to determine an encoding. Check r.encoding and r.headers if characters look corrupted. For byte-accurate handling, retain r.content and decode it using the encoding established for that response. Requests Quickstart

The parser produces an unexpected tree

HTML parsers repair malformed markup, and their recovery can differ. Try another installed Beautiful Soup parser when necessary, then keep the chosen parser fixed in repeatable extraction code. Beautiful Soup documentation

Use Scrapy to inspect its own response

If you are already working with Scrapy, its fetch command can show the response Scrapy receives:

scrapy fetch --nolog https://example.com > response.html

Compare that file with the browser’s source. If they differ, check the request headers and user agent, then reproduce the relevant browser request as needed. Scrapy’s documentation covers inspecting source and finding dynamic data sources. Scrapy: Dynamic Content

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common extraction problems

Symptom Likely cause What to do
The command fails to resolve or connect The URL is incomplete, the network is unavailable, or the host cannot be reached. Include https:// or http://, check the hostname, and retry when connectivity is available.
The saved page is a redirect destination or an error The request did not follow a redirect or the server returned an error page. Use curl -L; check status and headers with curl -i or inspect Requests’ status_code and headers.
The response is JSON or a login page You fetched an API endpoint, an authentication flow, or a page that requires authorized credentials. Check the content type and requested URL. If authorized, reproduce the needed authentication and request details.
Some visible content is missing JavaScript added it or a later XHR/Fetch request supplied it. Inspect Network requests, retrieve the underlying data request, or use a JavaScript-capable browser workflow.
Text has garbled characters The response was decoded with an unsuitable character encoding. Check the response encoding and headers; use raw bytes and decode deliberately if needed. Requests Quickstart
Links or elements are missing after parsing The document is malformed or parser recovery differs. Try lxml or html5lib and keep the parser choice consistent. Beautiful Soup documentation

Or skip the browser setup

If what you need is a visual screenshot or PDF rather than the raw HTML response, ScreenshotNeo can capture a URL with one GET request. See the ScreenshotNeo API documentation for supported parameters and response details. For example, this cURL command saves a screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. See ScreenshotNeo for the service and its documentation for setup. Sign up for 1,000 free screenshots a month—no card required.

Frequently Asked Questions

Does downloading a URL give me the full HTML?

It gives you the response body for that request. Content inserted later by JavaScript or fetched through separate requests may not be included.

Should I use View Source or the Elements panel?

Use View Source for the original document response and Elements for the browser’s current, possibly modified DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I extract HTML from a page behind a login?

Only if you are authorized. Use the legitimate credentials and request flow for that account; an HTTP fetch does not itself grant access.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.