Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In headless Selenium, load the page, wait for the state your application needs, then read driver.page_source. That is Selenium’s WebDriver page-source result. If you need the live DOM after JavaScript has modified it, run document.documentElement.outerHTML with driver.execute_script() instead.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

options = Options()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    WebDriverWait(driver, 10).until(
        lambda d: d.execute_script("return document.readyState") == "complete"
    )
    html = driver.page_source
    with open("page.html", "w", encoding="utf-8") as f:
        f.write(html)
finally:
    driver.quit()

The important detail is timing: a page can report that its document is complete while a client-side application is still fetching and inserting content. Replace the generic readiness check with a selector or application-specific state when necessary.

Use driver.page_source after an explicit wait

Selenium’s Python API exposes driver.page_source as the property that “Gets the source of the current page.” In a Chrome or Firefox headless session, the property is used exactly as it is in a headed session. Selenium sends the WebDriver GET_PAGE_SOURCE command and returns a Python string.

A reliable basic sequence is:

  1. Create a headless browser instance.
  2. Navigate with driver.get().
  3. Wait for the page state or a page-specific condition.
  4. Read driver.page_source.
  5. Save or process the returned string.
  6. Close the browser in a finally block.

A complete Python example

This example writes the rendered page source to a UTF-8 file and always closes the browser, including when navigation or saving raises an exception.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com"
OUTPUT = "page.html"

options = Options()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)

try:
    driver.get(URL)

    # This is a baseline check, not a universal application-ready signal.
    WebDriverWait(driver, 10).until(
        lambda d: d.execute_script("return document.readyState") == "complete"
    )

    html = driver.page_source
    with open(OUTPUT, "w", encoding="utf-8") as output_file:
        output_file.write(html)
finally:
    driver.quit()

For a single static document, the readyState check may be sufficient. For a single-page application, wait for something that proves the required content exists:

from selenium.webdriver.common.by import By

WebDriverWait(driver, 20).until(
    lambda d: d.find_element(By.CSS_SELECTOR, "main article")
)
html = driver.page_source

Choose a selector that belongs to the content you actually need. A spinner disappearing, a result count appearing, or an application-specific JavaScript flag can be a better readiness signal than a fixed delay.

Choose between page_source and live-DOM serialization

Method What it returns Use it when Important qualification
driver.page_source The WebDriver page-source result for the current browsing context. You want Selenium’s standard page-source API with minimal code. It is not documented as byte-for-byte equality with the original HTTP response.
driver.execute_script("return document.documentElement.outerHTML;") A serialization of the document element as it exists when the script runs. You specifically want the current DOM after client-side mutations. The result depends on the DOM state and the active frame at execution time.

To serialize the live document, use Selenium’s synchronous JavaScript execution method:

html = driver.execute_script(
    "return document.documentElement.outerHTML;"
)
with open("live-dom.html", "w", encoding="utf-8") as f:
    f.write(html)

Neither method should be treated as a network capture. If your requirement is the raw response body exactly as it arrived over HTTP, use a browser-appropriate network-capture technique instead of assuming that page source is the wire payload.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for asynchronous content instead of guessing with sleep

Headless mode does not make JavaScript run synchronously. Frameworks can fetch data, render components, and replace markup after the initial navigation has completed. A hard-coded time.sleep() waits the same amount every time, even when the page is ready sooner, and can still be too short on a slow run.

Use WebDriverWait with a condition that represents the state you need:

WebDriverWait(driver, 20).until(
    lambda d: d.find_element(By.CSS_SELECTOR, "[data-loaded='true']")
)
html = driver.page_source

The exact condition is site-specific. If a request can legitimately fail, wait for either a success element or an error element and handle both outcomes. Capture only after the target content is present, not merely after the browser window has opened.

Capture markup inside an iframe

Selenium commands operate in the active browsing context. If the desired HTML is inside an iframe, switch to that frame before reading the source:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait

frame = WebDriverWait(driver, 20).until(
    lambda d: d.find_element(By.CSS_SELECTOR, "iframe#content-frame")
)
driver.switch_to.frame(frame)
try:
    WebDriverWait(driver, 20).until(
        lambda d: d.execute_script("return document.readyState") == "complete"
    )
    frame_html = driver.page_source
finally:
    driver.switch_to.default_content()

After switching, page_source refers to the frame’s document, not the top-level page. Call default_content() when you need to return to the main document and continue interacting with it.

Headless Chrome and Firefox setup

Chrome

Use Chrome options and pass them to the driver:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)

Firefox

Firefox has its own options class, but the retrieval call remains driver.page_source:

from selenium import webdriver
from selenium.webdriver.firefox.options import Options

options = Options()
options.add_argument("--headless")
driver = webdriver.Firefox(options=options)

Keep the browser, Selenium package, and driver components compatible. The source-reading code does not change when you switch between Chrome and Firefox; only browser construction and browser-specific options do.

Save and process the returned HTML safely

page_source returns text, so open the destination with an explicit encoding. UTF-8 is a practical default for modern pages:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
html = driver.page_source

with open("snapshot.html", "w", encoding="utf-8", newline="") as file:
    file.write(html)

# You can also inspect it without writing a file.
print(len(html))
print(html[:500])

For repeated captures, keep one driver alive for the batch when the site and isolation requirements allow it, navigate to each URL, wait for that URL’s readiness condition, and read the source before moving on. Always close the driver at the end of the batch. Very large documents consume memory both in the browser and in Python, so process or stream your own downstream representation rather than retaining every page indefinitely.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Symptom Likely cause Fix
TimeoutException while waiting The selector is wrong, the application uses a different readiness signal, or the page never reaches the expected state. Verify the selector in the page, wait for a success-or-error condition, and use a timeout appropriate for the application.
HTML lacks content visible in the application Capture occurred before asynchronous rendering finished. Wait for the actual content or a framework-specific ready signal rather than relying only on navigation completion.
The returned markup is for the wrong document The driver is still inside an iframe or another browsing context. Switch to the intended frame before capture, or call driver.switch_to.default_content() first.
Session cannot be created The browser, Selenium package, and driver are incompatible or the browser is unavailable in the runtime. Install a supported browser, update the Selenium package and matching driver components, and verify the same setup outside headless mode.
The process exits before a file is written An exception occurred before cleanup or file output completed. Put navigation, waiting, and writing inside try and close the driver in finally; log the exception and destination path.
A challenge or CAPTCHA appears in the source The site served a bot-check page instead of the requested document. Treat the captured HTML as the challenge response; do not assume that a successful navigation call means the intended page was delivered.

Reliability and performance considerations

  • Use deterministic readiness checks. A selector tied to the required content is usually more meaningful than a universal delay.
  • Keep the browsing context explicit. Record whether each capture is from the top-level document or a selected frame.
  • Reuse sessions carefully. Reusing a driver reduces startup overhead, but clear or isolate state when cookies and local storage could change the result.
  • Record the URL and timing. Store the requested URL, capture timestamp, and any application-level status alongside the HTML so later processing can distinguish a real page from an error document.
  • Expect source size to vary. Full rendered documents can be much larger than the initial response; avoid holding an unbounded number of them in memory.

Or skip the browser setup

If your goal is a screenshot or PDF rather than HTML source, ScreenshotNeo provides a single-call capture API. It accepts the consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the result with X-Page-Verdict and X-Billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for options and response details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots, and every feature is included on every plan. Create a free ScreenshotNeo account to get started.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

How can I save only one element instead of the entire document?

Run JavaScript against the active document and return that element’s serialization: driver.execute_script("return document.querySelector('main article').outerHTML;"). Check that the selector exists after the same readiness wait you use for the full page.

How do I read source from a different tab?

Switch to the tab’s window handle first with driver.switch_to.window(handle), then read driver.page_source. Selenium always returns source for the currently selected window and browsing context.

Can I inspect the captured text without creating a file?

Yes. The property returns a Python string, so you can parse html, search it, measure its length, or pass it to another function before deciding whether to write it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.