To extract a page’s HTML, make an HTTP GET request to its URL and save the response body. For a quick check, use a browser’s View Source command; for repeatable work, use curl, Wget, or Python Requests. The result is the HTML sent in the response—not necessarily the live page you see after JavaScript runs.
Choose the right kind of page source
There are two common meanings of “the HTML from a URL,” and choosing the right one prevents a lot of confusion:
- Original response HTML: the body returned by the server for the document request. Use View Source or an HTTP client such as curl when you want to inspect what the server sent.
- Rendered DOM: the current document structure after the browser has parsed the response, executed scripts, and possibly loaded more data. Use developer tools or a browser automation workflow when you need what the page became after loading.
A simple GET request retrieves the document identified by a URL; curl’s documentation describes GET as returning the document body. curl documentation The browser’s Elements panel shows the current DOM, which can differ from the original response when scripts alter the page or fetch content later. Scrapy: Dynamic Content
Extract HTML in a browser
View the original response
Open the page, then use the browser’s View Source command—often available by right-clicking the page or through the browser menu. This displays the document response as source text. It is the fastest option for a one-off inspection, but it is not a convenient way to process many URLs or save structured data.
#1 Best Overall
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Inspect the live DOM
Open developer tools and select the Elements panel. This shows the browser’s current DOM, not necessarily the exact response body. If a value appears here but not in View Source, it may have been inserted by JavaScript or fetched after the initial document request.
Find data loaded separately
In developer tools, open the Network panel, reload the page, and look for XHR or Fetch requests. Inspect the request and response that contain the data you need. Scrapy’s guidance recommends locating the source of the data and reproducing the browser request, including its method, URL, headers, and body where needed. Scrapy: Dynamic Content If your browser offers an option to copy or export a request as cURL, use that as a starting point rather than guessing at the request.
Download HTML with curl or Wget
curl
Follow redirects and write the response body to a file:
curl -L "https://example.com" -o page.html
The -L option follows redirects; -o names the output file. To inspect response headers alongside the body, use -i. To request headers only, use -I; a HEAD request returns headers without the response body. See the curl manual for option details.
Rank #2
Wget
Save a page under a chosen filename with:
wget -O page.html "https://example.com"
Wget can also retrieve linked resources in recursive mode. Recursion can turn one page into a much larger crawl, so set an appropriate depth, domain boundary, and output directory before using it. Consult the GNU Wget manual for recursive retrieval options.
Check what you downloaded
Open page.html in a text editor. If the file contains an error message, login screen, JSON, or a bot-check page, the request did not produce the target HTML even if a file was saved successfully. Check the HTTP status and content type before treating the response as the page you wanted.
Fetch and save a URL with Python Requests
Install Requests if it is not already available: python -m pip install requests. Then fetch the URL, fail visibly for HTTP errors, and save the decoded response text:
import requests
url = "https://example.com"
r = requests.get(url, timeout=20)
r.raise_for_status()
html = r.text
print(html)
with open("page.html", "w", encoding=r.encoding or "utf-8") as f:
f.write(html)
Requests exposes decoded response text through r.text, raw response bytes through r.content, and metadata such as headers through r.headers. It supports redirects, cookies, SSL verification, and timeouts; its documentation explains these behaviors. Requests Quickstart Use r.content instead of r.text when you need the original bytes or plan to handle decoding yourself. A timeout prevents a stalled request from waiting indefinitely, while raise_for_status() ensures a 4xx or 5xx response is not silently treated as a successful fetch.
Rank #3
Parse the downloaded HTML with Beautiful Soup
Fetching and parsing are separate tasks: Requests obtains the response; Beautiful Soup turns the markup into a navigable tree. Install it with python -m pip install beautifulsoup4, then parse the saved response text:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
print(soup.title.get_text(strip=True) if soup.title else "No title")
for link in soup.select("a[href]"):
print(link.get("href"))
html.parser is included with Python. Beautiful Soup can also use lxml or html5lib when those packages are installed. Malformed markup may produce different trees with different parsers, so specify the parser if you need repeatable results. Beautiful Soup documentation
When the HTML is not the content you expected
JavaScript adds the content later
A direct HTTP fetch does not run page JavaScript. If the initial HTML lacks text that appears in the browser, inspect the Network panel for the later request that supplies it. Reproducing that request is often simpler and more stable than rendering the whole page. If there is no practical request to reproduce, use a browser-based renderer that executes JavaScript before you extract the content. Scrapy discusses both identifying the source request and handling dynamic content. Scrapy: Dynamic Content
The server responds differently to your script
Some pages require cookies, authentication, a particular user agent, or other headers. Compare the browser’s document request with your script’s request and reproduce only the necessary details. Do this only for content you are authorized to access; do not treat a login or access control as a technical obstacle to bypass.
Recommended Free Tools
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
The response encoding looks wrong
Requests decodes response text automatically, using response information to determine an encoding. Check r.encoding and r.headers if characters look corrupted. For byte-accurate handling, retain r.content and decode it using the encoding established for that response. Requests Quickstart
The parser produces an unexpected tree
HTML parsers repair malformed markup, and their recovery can differ. Try another installed Beautiful Soup parser when necessary, then keep the chosen parser fixed in repeatable extraction code. Beautiful Soup documentation
Use Scrapy to inspect its own response
If you are already working with Scrapy, its fetch command can show the response Scrapy receives:
scrapy fetch --nolog https://example.com > response.html
Compare that file with the browser’s source. If they differ, check the request headers and user agent, then reproduce the relevant browser request as needed. Scrapy’s documentation covers inspecting source and finding dynamic data sources. Scrapy: Dynamic Content
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
Troubleshoot common extraction problems
| Symptom | Likely cause | What to do |
|---|---|---|
| The command fails to resolve or connect | The URL is incomplete, the network is unavailable, or the host cannot be reached. | Include https:// or http://, check the hostname, and retry when connectivity is available. |
| The saved page is a redirect destination or an error | The request did not follow a redirect or the server returned an error page. | Use curl -L; check status and headers with curl -i or inspect Requests’ status_code and headers. |
| The response is JSON or a login page | You fetched an API endpoint, an authentication flow, or a page that requires authorized credentials. | Check the content type and requested URL. If authorized, reproduce the needed authentication and request details. |
| Some visible content is missing | JavaScript added it or a later XHR/Fetch request supplied it. | Inspect Network requests, retrieve the underlying data request, or use a JavaScript-capable browser workflow. |
| Text has garbled characters | The response was decoded with an unsuitable character encoding. | Check the response encoding and headers; use raw bytes and decode deliberately if needed. Requests Quickstart |
| Links or elements are missing after parsing | The document is malformed or parser recovery differs. | Try lxml or html5lib and keep the parser choice consistent. Beautiful Soup documentation |
Or skip the browser setup
If what you need is a visual screenshot or PDF rather than the raw HTML response, ScreenshotNeo can capture a URL with one GET request. See the ScreenshotNeo API documentation for supported parameters and response details. For example, this cURL command saves a screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. See ScreenshotNeo for the service and its documentation for setup. Sign up for 1,000 free screenshots a month—no card required.
Frequently Asked Questions
Does downloading a URL give me the full HTML?
It gives you the response body for that request. Content inserted later by JavaScript or fetched through separate requests may not be included.
Should I use View Source or the Elements panel?
Use View Source for the original document response and Elements for the browser’s current, possibly modified DOM.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteCan I extract HTML from a page behind a login?
Only if you are authorized. Use the legitimate credentials and request flow for that account; an HTTP fetch does not itself grant access.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

