Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To capture the HTML a website has rendered, open it in a real browser, wait until the content you need is present, then serialize the live DOM. In Playwright, use await page.content() for the whole document; in Selenium, use driver.page_source in Python or driver.getPageSource() in Java. These are browser-side representations of the current DOM, not guaranteed copies of the original HTTP response bytes.

Choose what you mean by “the HTML”

A browser automation capture records the document as it exists at a particular moment. That distinction matters on JavaScript-heavy pages: the initial server response may contain only a shell, while the live DOM gains content after scripts run, data arrives, or you interact with the page.

  • Rendered full document: serialize the browser’s current page after the relevant content appears.
  • One region: serialize a selected element’s outerHTML.
  • Original response: capture the network response body separately; browser DOM serialization can change formatting or escaping.
  • Portable archive: use a resource-aware format such as a DevTools Protocol MHTML snapshot when external assets and embedded content matter.

HTML alone does not include downloaded copies of referenced images, stylesheets, fonts, or scripts. Decide whether you need just markup or a reproducible archive before choosing the capture method.

Capture the rendered document with Playwright

Playwright’s page.content() returns the full HTML contents of the page, including the doctype. The following Node.js example waits for the main content before saving the page. Install Playwright in your project first with npm install playwright; install its browser binaries with npx playwright install chromium.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { chromium } from 'playwright';

const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.locator('main').waitFor();
  const html = await page.content();
  await import('node:fs/promises').then(fs => fs.writeFile('page.html', html, 'utf8'));
} finally {
  await browser.close();
}

Save this as an ES module, for example capture.mjs, then run node capture.mjs. If the site does not use a main element, replace that locator with a selector for the content you actually need. The wait verifies that an element exists; if content is inserted later into an already-present container, wait for a more specific child or another observable condition.

Capture one element instead of the whole page

Use outerHTML to serialize a selected element and its descendants:

const sectionHtml = await page.locator('main').evaluate(el => el.outerHTML);
await import('node:fs/promises').then(fs => fs.writeFile('section.html', sectionHtml, 'utf8'));

This does not include the surrounding document’s doctype, head, or sibling elements. Choose a stable, specific selector: a broad selector can match the wrong region or more than one element.

Wait for the condition that proves readiness

domcontentloaded is a navigation milestone, not proof that a client-rendered application has finished loading its data. A locator wait is usually a better signal when the desired content has a distinct element. You can also wait for a response or application-specific state. Playwright provides navigation and evaluation APIs for sequencing these actions. A fixed delay may capture too early, or waste time after the page is ready.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The capture point defines the result. If the relevant content appears only after a click, login, scroll, or API response, perform that action and wait for its observable result before calling page.content(). A full-page HTML serialization does not itself cause every lazy-loaded image or other resource to load.

Capture the current DOM with Selenium

In Python, Selenium exposes the current page source as driver.page_source. This example waits for a visible page element, then writes the result as UTF-8:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait

with webdriver.Chrome() as driver:
    driver.get('https://example.com')
    WebDriverWait(driver, 10).until(
        lambda d: d.find_element('css selector', 'main')
    )
    html = driver.page_source
    with open('page.html', 'w', encoding='utf-8') as f:
        f.write(html)

Install Selenium with pip install selenium and make a compatible Chrome browser available. Selenium’s Java API equivalent is driver.getPageSource(). The returned source is a representation of the underlying DOM; do not treat it as a byte-for-byte copy of the server’s response or expect identical formatting and escaping.

Handle iframes and shadow DOM explicitly

A top-level page serialization is not a universal dump of every browsing context and encapsulated tree. Identify which content belongs to the page, a frame, or a shadow root before assuming it will appear in the output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Iframes

Capture a relevant frame’s document separately by locating or enumerating the frame and evaluating within that frame’s context. Do not assume an iframe’s DOM is embedded in the top-level HTML string. Cross-origin boundaries and authentication can limit what the browser session can access; record those conditions if the capture is used for debugging or audit work.

Shadow roots

Ordinary serialization may omit content held in shadow roots. MDN documents Element.getHTML() as serializing an element’s DOM to an HTML string, with options for including child shadow roots where supported. Check browser support and whether the root is open; closed shadow roots may not be available to page scripts. See MDN’s Element.getHTML() documentation.

When you need a resource-aware archive

If the goal is a portable record rather than just markup, use a DevTools Protocol MHTML snapshot where supported. Its documented snapshot format can include iframes, shadow DOM, external resources, and inline styles. HTML serialization by itself does not download and package every linked dependency, so a saved .html file may render differently or incompletely when opened later.

What browser serialization preserves—and what it does not

  • It captures the current DOM state: waits and interactions determine what is present in the output.
  • It is not necessarily raw source: Selenium specifically describes page source as a representation of the underlying DOM, not the original response’s exact formatting or escaping.
  • It is not a universal completion signal: no single wait duration guarantees readiness across sites or frameworks.
  • It may omit inaccessible content: authentication requirements, permissions, anti-bot controls, cross-origin rules, or closed shadow roots can constrain what the session can inspect.
  • It is not a dependency bundle: external images, CSS, fonts, and scripts are not automatically downloaded into the HTML file.

For debugging a rendered page, record the URL, browser, capture time, and actions or waits used. For exact response analysis, save the network response instead. For repeatable offline viewing, choose an archive or capture the required network resources as well.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and practical fixes

The saved HTML has a mostly empty app shell

The page was serialized before client-side content appeared, or the wait only checked for an outer container. Wait for the actual text, row, card, or other target element; if necessary, wait for the relevant response or perform the interaction that triggers rendering.

The target selector never appears

The selector may be wrong, the page may have navigated differently, or the content may be inside a frame. Inspect the rendered page and use a selector that exists in the correct context. If it is in an iframe, switch to or evaluate within that frame rather than waiting only in the top-level document.

The result differs from “View Source”

That is expected when you compare a live DOM serialization with the original response. Use the browser’s network response data if you need the received bytes; use the DOM capture when you need what the browser has constructed after scripts and interactions.

Some embedded content or components are missing

Check whether the content is in an iframe or shadow root, and whether the session has access to it. Capture frames separately. For shadow DOM, use supported shadow-root serialization where possible; do not assume closed roots can be extracted.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The saved file opens without its images or styling

The HTML references resources but does not package them. Use an MHTML snapshot or capture dependencies separately if offline fidelity matters.

The automation is slow or flaky

A broad fixed sleep is both inefficient and unreliable. Replace it with a wait for the specific content or event needed. Keep the navigation milestone and content-ready condition distinct: reaching domcontentloaded does not mean an application’s data has rendered.

Or skip the browser setup

If you need a screenshot or PDF rather than an HTML serialization, ScreenshotNeo is a website screenshot API and MCP server. It does not return a page’s HTML, so it is not a replacement for page.content(); it can avoid running browser infrastructure when a visual capture is the goal. One GET request returns an image or PDF. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up free for 1,000 screenshots a month with no card.

Frequently Asked Questions

Does Playwright page.content() include the doctype?

Yes. Playwright documents that page.content() returns the full HTML contents of the page, including the doctype.

Will browser automation capture HTML from a login-protected page?

Only to the extent the browser session is authenticated and permitted to access the content. The captured result reflects what that session can inspect at the time of serialization.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.