Use await page.content() to retrieve the page’s complete current HTML in Puppeteer, including the <!DOCTYPE html> declaration. Call it after navigation and after a readiness condition that proves the content you need has rendered:
const response = await page.goto('https://example.com');
await page.waitForSelector('main');
const html = await page.content();
console.log(html);
This is the browser’s parsed, current document—not necessarily the byte-for-byte HTTP response originally sent by the server. The distinction matters on JavaScript applications, pages with iframes, and any workflow that requires original response fidelity.
What “page source” means in Puppeteer
People use “page source” for two different outputs:
- Current browser HTML: the DOM after scripts, framework rendering, user interaction and browser parsing. Use
page.content()or DOM evaluation. - Original network response: the response body received before browser-side scripts modify the document. Capture the navigation response or use an HTTP client when byte-for-byte fidelity is required.
Puppeteer’s current API documentation (displayed version 25.12.0) describes Page.content() as returning “The full HTML contents of the page, including the DOCTYPE.”
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Get the complete current HTML with page.content()
A minimal runnable script:
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({headless: true});
try {
const page = await browser.newPage();
await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});
const html = await page.content();
console.log(html);
} finally {
await browser.close();
}
content() resolves to a string containing the serialized document. It is a getter: it reads what the page currently contains and does not change the page.
Check navigation status when HTTP success matters
goto() can return a response for valid HTTP error statuses such as 404 or 500; those statuses do not necessarily make navigation throw. Inspect the response before extracting HTML:
const response = await page.goto(url, {waitUntil: 'domcontentloaded'});
if (!response) throw new Error('No navigation response');
if (!response.ok()) {
throw new Error(`Navigation failed with HTTP ${response.status()}`);
}
const html = await page.content();
Wait for dynamic content before extraction
Navigation completion and application readiness are different events. A single-page app may render its meaningful content after DOMContentLoaded. Prefer a condition tied to the content you need rather than an arbitrary delay.
Wait for a selector
await page.goto('https://example.com/dashboard', {
waitUntil: 'domcontentloaded'
});
await page.waitForSelector('main[data-ready="true"]');
const html = await page.content();
Wait for application state
await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForFunction(() => window.app?.hydrated === true);
const html = await page.content();
Wait for network idle when that is the right signal
await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForNetworkIdle({idleTime: 500, timeout: 30000});
const html = await page.content();
Network idle alone is not proof that a page’s data is complete: analytics, polling and long-lived connections can prevent idleness, while an app may finish rendering before every request stops. A selector or application-specific condition is usually more precise. Do not substitute a fixed sleep unless the site offers no better readiness signal.
Serialize the DOM with evaluate()
Use browser-context evaluation when you need an explicit DOM serialization or want to inspect values while the page is open:
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
const html = await page.evaluate(() => document.documentElement.outerHTML);
This returns the document element’s serialized HTML. For only the body:
const bodyHtml = await page.evaluate(() => document.body.innerHTML);
evaluate() is useful when extraction is part of a larger DOM operation, but page.content() is the direct whole-document method.
Extract one element with $eval()
When the full document is unnecessary, select the first matching element and return its contents:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11const mainHtml = await page.$eval('main', element => element.innerHTML);
$eval() passes the first matching element to your function and throws if no element matches. Guard optional content explicitly:
const text = await page.$eval('main', el => el.textContent?.trim() ?? '');
Use a stable selector such as a semantic element, ID or data attribute. Avoid brittle selectors tied to generated class names.
Rank #3
Whole page versus selected region
| Need | Method | Result |
|---|---|---|
| Complete current document | page.content() |
Full HTML including DOCTYPE |
| Explicit DOM serialization | page.evaluate(() => document.documentElement.outerHTML) |
Current document element |
| One region | page.$eval(selector, ...) |
Selected element’s output |
| Body only | page.evaluate(() => document.body.innerHTML) |
Children of body |
| Original response bytes | Navigation response or HTTP client | Network payload, not a post-render DOM |
Read HTML inside an iframe
An iframe has a separate document and execution context. The top-level page’s content() does not replace the need to inspect the child frame.
await page.goto('https://example.com');
const frame = page.frames().find(f => f.url().includes('/embedded/'));
if (!frame) throw new Error('Embedded frame not found');
await frame.waitForSelector('main');
const frameHtml = await frame.content();
For a frame whose URL is not known, locate the iframe element, obtain its frame with contentFrame(), then query within that frame. Cross-origin security rules still apply to what browser JavaScript can access.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
setContent() is a setter, not a source getter
page.setContent(html) assigns markup to the page and returns Promise<void>. It does not retrieve source. Read the resulting document afterward:
await page.setContent('<main>Generated</main>');
const html = await page.content();
Current DOM is not original server source
Browser parsing can normalize markup, and scripts can add, remove or replace nodes. If you need the server’s original response, capture the navigation response separately:
const response = await page.goto(url, {waitUntil: 'domcontentloaded'});
if (!response) throw new Error('No response');
const responseText = await response.text();
const renderedHtml = await page.content();
The response text and rendered HTML answer different questions. The first is the received payload; the second is the document represented by the browser at extraction time. A response may also be compressed or transformed at the transport layer, so define whether you need decoded response text or exact wire bytes before designing the capture.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Complete extraction example with file output
import puppeteer from 'puppeteer';
import {writeFile} from 'node:fs/promises';
const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch({headless: true});
try {
const page = await browser.newPage();
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30000
});
if (!response) throw new Error('Navigation returned no response');
await page.waitForSelector('body', {timeout: 10000});
const html = await page.content();
await writeFile('page-source.html', html, 'utf8');
console.log(`Saved ${html.length} characters from ${page.url()}`);
} finally {
await browser.close();
}
Troubleshooting
The HTML is missing content rendered by JavaScript
Cause: extraction ran after navigation but before the app finished rendering. Fix: wait for a meaningful selector, state predicate or network condition, and verify the selector actually represents loaded data.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →waitForSelector() times out
Cause: wrong URL, authentication requirement, a changed selector, a failed request or content inside a different frame. Log page.url(), inspect the HTTP response status, check console and request failures, and search page.frames() for the target document.
$eval() throws “failed to find element
Cause: no element matched at the instant of evaluation. Wait for the selector first, use an optional lookup with page.$(), or correct the selector.
The output contains a DOCTYPE and more markup than expected
That is normal for page.content(). It returns the full document. Use $eval() or body evaluation when you need a narrower region.
Content is in an iframe
Find the relevant Frame and call frame methods there. Do not assume the parent document contains the child document’s nodes.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesBest Value
The result differs from “view source”
“View source” commonly represents the server response, while Puppeteer extraction usually represents the live DOM after parsing and script execution. Capture response.text() for the navigation payload and page.content() for the rendered document.
Navigation appears successful but the page is an error screen
Inspect response.status() and response.ok(). Valid 404 or 500 responses may not throw from goto().
Performance and reliability considerations
- Extract only the region you need to reduce serialization and downstream processing.
- Use explicit timeouts and always close the browser in a
finallyblock. - Reuse a browser for batches, but isolate pages when cookies, storage or authentication must not leak.
- Record the final URL, status, extraction timestamp and readiness condition alongside the HTML.
- Be careful with continuously updating pages: two calls to
content()can legitimately differ. - Respect authentication, robots policies, terms and rate limits applicable to the site you access.
Or skip the browser setup
If you need a clean screenshot rather than HTML source, ScreenshotNeo provides a single-request website capture API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server works with Claude, Cursor and other MCP clients through take_screenshot, get_page_info and capture_pdf.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options including full-page lazy-image loading, CSS-selector element capture, device presets, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture and usage data.
Recommended Free Tools
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
Frequently Asked Questions
Does Puppeteer have a method literally called getPageSource?
No. The standard whole-document equivalent is page.content(); DOM-specific alternatives are evaluate() and $eval().
Can page.content() retrieve HTML generated after a click?
Yes. Perform the click, wait for the resulting selector or application state, then call page.content().
How do I get only the original HTML sent by the server?
Capture the navigation response, such as with const response = await page.goto(url) followed by response.text(), or use a direct HTTP client.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

