iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Use await page.content() after navigation and after the page reaches the state you need. Puppeteer returns a string containing the current document’s complete HTML, including the <!DOCTYPE html> declaration. For JavaScript-heavy pages, wait for a meaningful selector, response, or application-ready signal before extracting.
The canonical Puppeteer pattern
This minimal script opens a page, waits for the DOM to be parsed, waits for a page-specific element, and then reads the rendered document:
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('main'); // Choose a selector that means the page is ready.
const html = await page.content();
console.log(html);
await browser.close();
page.content() is asynchronous, so call it with await. The result is a JavaScript string; you can print it, parse it, save it to disk, or return it from another function.
Recommended Free Tools
Save the HTML to a file
import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('main');
const html = await page.content();
await writeFile('page.html', html, 'utf8');
} finally {
await browser.close();
}
The try/finally block closes Chromium even if navigation, waiting, or writing fails.
#1 Best Overall
What “full HTML” means
Puppeteer documents Page.content() as returning “the full HTML contents of the page, including the DOCTYPE.” It serializes the document as it exists in the browser at the moment the method runs. That distinction matters: it gives you the current DOM state after scripts have changed it, not a guarantee that every future asynchronous operation has completed.
If a framework has already inserted products, article text, or other nodes into the DOM, those nodes appear in the returned string. If a lazy section has not yet been triggered, or an API request is still pending, its content may not be present yet. Wait for the condition that represents completeness for your page, then call page.content().
Choose a readiness wait that matches the page
Wait for a meaningful selector
A selector is usually the most reliable option when you know which element proves that the required content is available:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-testid="results"]', { timeout: 15000 });
const html = await page.content();
Prefer a stable application selector over a generic class name that changes between builds. If the selector can exist before its text is populated, add a separate check for the text or item count.
Wait for a specific response
When the page is complete only after an API call, coordinate navigation and the response wait:
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/articles') && response.ok()
);
await page.goto('https://example.com/articles', { waitUntil: 'domcontentloaded' });
await responsePromise;
const html = await page.content();
This ties extraction to the data your page needs rather than to an arbitrary delay.
Wait for network idle
Puppeteer provides page.waitForNetworkIdle(). It resolves only after the configured idle period, and the API notes that it always waits at least that idleTime:
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForNetworkIdle({ idleTime: 500, timeout: 10000 });
const html = await page.content();
Network idle is useful when requests settle predictably. Analytics, advertisements, polling, WebSockets, streaming, or other long-lived connections can keep the network busy indefinitely. In those cases, use a selector or response that directly represents the content you need. A short fixed delay can be a fallback, but it is less deterministic because fast and slow runs receive the same wait.
Rank #3
Use an application-ready flag
If your own application exposes a readiness marker, wait for it:
await page.waitForFunction(() => window.appReady === true, {
timeout: 15000
});
const html = await page.content();
Only define such a flag when it means the specific data you intend to capture has finished rendering.
Synchronize clicks that trigger navigation
A common source of incomplete markup is extracting while a click-triggered navigation is still underway. Start the navigation wait before the click and await both operations together:
Free tools Windows power users keep installed
One-click scans. No signup required.
await Promise.all([
page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
page.click('a.next-page')
]);
await page.waitForSelector('main');
const html = await page.content();
For single-page applications, a click may change the URL or DOM without a traditional navigation. In that case, wait for the route’s selector or its API response instead of relying on waitForNavigation().
Handling lazy-loaded and asynchronously inserted content
page.content() captures only the state available when it executes. To include content loaded after scrolling, perform the interaction first:
await page.goto('https://example.com/feed', { waitUntil: 'domcontentloaded' });
await page.evaluate(async () => {
window.scrollTo(0, document.body.scrollHeight);
});
await page.waitForSelector('.feed-item:last-child');
const html = await page.content();
For multiple lazy-load cycles, scroll in a loop and stop when the page-specific completion condition is met. Do not assume that reaching the bottom once loads every item; virtualized lists may remove off-screen nodes, so the current DOM may not contain the entire logical dataset.
page.content() versus outerHTML
| Approach | Result | Best use |
|---|---|---|
page.content() |
Serialized current page HTML, including the DOCTYPE | Normal full-document extraction through Puppeteer |
document.documentElement.outerHTML |
Markup for the current document element, evaluated inside the page | Custom browser-side processing before returning markup |
The lower-level alternative is:
const html = await page.evaluate(() => document.documentElement.outerHTML);
Use it when you need to transform or inspect the markup in page context. For the ordinary “get the full HTML” task, page.content() is clearer and preserves the DOCTYPE in the documented result.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Frames and the meaning of “full page”
The main page’s page.content() result represents its document. HTML rendered inside an iframe belongs to a different document and is not merged into the parent’s serialization. Locate the frame and extract it separately:
const frame = page.frames().find(f => f.url().includes('/embedded'));
if (!frame) throw new Error('Embedded frame was not found');
await frame.waitForSelector('.content');
const frameHtml = await frame.content();
Cross-origin restrictions still apply to what browser code can access, and a frame may load later than the parent. Wait on the frame itself, not only on the top-level page.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common failures and fixes
The HTML contains only a shell
- Cause: extraction ran before the application rendered its data.
- Fix: wait for a content selector, required response, or app-ready flag. Use network idle only when the site’s connections actually settle.
waitForSelector times out
- Cause: the selector is wrong, the page navigated elsewhere, a consent gate blocks rendering, or the element is inside a frame.
- Fix: verify the final URL, inspect a screenshot or console output, increase the timeout only when the page is legitimately slow, and search
page.frames()for iframe content.
waitForNetworkIdle never completes
- Cause: polling, analytics, streaming, ads, or open connections keep the network active.
- Fix: replace it with a selector or response wait, or set a bounded timeout and handle the timeout explicitly.
Navigation races a click
- Cause: the click was awaited before the navigation wait was registered, or extraction happened before the new route rendered.
- Fix: use
Promise.allwith the navigation wait and click, then apply the destination’s readiness condition.
Content is missing from an iframe
- Cause: the parent document does not include the iframe’s separate DOM.
- Fix: select the correct frame and call
frame.content()after waiting inside it.
The browser process remains running
- Cause: an exception bypassed
browser.close(). - Fix: put the workflow in
try/finallyand close the browser in thefinallyblock.
Practical reliability and performance notes
- Reuse one browser process for a batch of URLs, but create a fresh page when isolation is required.
- Set explicit navigation and selector timeouts so a broken site cannot stall a worker indefinitely.
- Capture after the smallest reliable readiness condition; waiting for every request can add latency without improving the HTML you need.
- Keep the URL, final URL, wait condition, and extraction timestamp with saved HTML so later debugging can distinguish redirects from rendering problems.
- Large DOMs produce large strings. Stream or write results promptly rather than retaining every page in memory during a crawl.
Or skip the browser setup
If your goal is a visual screenshot rather than the DOM string, ScreenshotNeo provides a single-request capture API. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents, including Claude and Cursor.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options and response details. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free.
Frequently Asked Questions
Does page.content() return the original server response?
No. It serializes the document currently held by the browser, so client-side changes made after navigation are included.
Can I get HTML before JavaScript runs?
Use the server response or another HTTP client for source HTML. Puppeteer’s page.content() is intended for the current browser DOM after whatever rendering has occurred.
Why is my saved HTML different on two runs?
The page may depend on timing, personalization, random data, consent state, lazy loading, or changing API responses. Use deterministic waits and browser settings where your application permits.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

