Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Use await page.content() after navigation and after the page reaches the state you need. Puppeteer returns a string containing the current document’s complete HTML, including the <!DOCTYPE html> declaration. For JavaScript-heavy pages, wait for a meaningful selector, response, or application-ready signal before extracting.

The canonical Puppeteer pattern

This minimal script opens a page, waits for the DOM to be parsed, waits for a page-specific element, and then reads the rendered document:

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();

await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('main'); // Choose a selector that means the page is ready.

const html = await page.content();
console.log(html);

await browser.close();

page.content() is asynchronous, so call it with await. The result is a JavaScript string; you can print it, parse it, save it to disk, or return it from another function.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Save the HTML to a file

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.waitForSelector('main');

  const html = await page.content();
  await writeFile('page.html', html, 'utf8');
} finally {
  await browser.close();
}

The try/finally block closes Chromium even if navigation, waiting, or writing fails.

What “full HTML” means

Puppeteer documents Page.content() as returning “the full HTML contents of the page, including the DOCTYPE.” It serializes the document as it exists in the browser at the moment the method runs. That distinction matters: it gives you the current DOM state after scripts have changed it, not a guarantee that every future asynchronous operation has completed.

If a framework has already inserted products, article text, or other nodes into the DOM, those nodes appear in the returned string. If a lazy section has not yet been triggered, or an API request is still pending, its content may not be present yet. Wait for the condition that represents completeness for your page, then call page.content().

Choose a readiness wait that matches the page

Wait for a meaningful selector

A selector is usually the most reliable option when you know which element proves that the required content is available:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-testid="results"]', { timeout: 15000 });
const html = await page.content();

Prefer a stable application selector over a generic class name that changes between builds. If the selector can exist before its text is populated, add a separate check for the text or item count.

Wait for a specific response

When the page is complete only after an API call, coordinate navigation and the response wait:

const responsePromise = page.waitForResponse(response =>
  response.url().includes('/api/articles') && response.ok()
);

await page.goto('https://example.com/articles', { waitUntil: 'domcontentloaded' });
await responsePromise;
const html = await page.content();

This ties extraction to the data your page needs rather than to an arbitrary delay.

Wait for network idle

Puppeteer provides page.waitForNetworkIdle(). It resolves only after the configured idle period, and the API notes that it always waits at least that idleTime:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForNetworkIdle({ idleTime: 500, timeout: 10000 });
const html = await page.content();

Network idle is useful when requests settle predictably. Analytics, advertisements, polling, WebSockets, streaming, or other long-lived connections can keep the network busy indefinitely. In those cases, use a selector or response that directly represents the content you need. A short fixed delay can be a fallback, but it is less deterministic because fast and slow runs receive the same wait.

Use an application-ready flag

If your own application exposes a readiness marker, wait for it:

await page.waitForFunction(() => window.appReady === true, {
  timeout: 15000
});
const html = await page.content();

Only define such a flag when it means the specific data you intend to capture has finished rendering.

Synchronize clicks that trigger navigation

A common source of incomplete markup is extracting while a click-triggered navigation is still underway. Start the navigation wait before the click and await both operations together:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await Promise.all([
  page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
  page.click('a.next-page')
]);

await page.waitForSelector('main');
const html = await page.content();

For single-page applications, a click may change the URL or DOM without a traditional navigation. In that case, wait for the route’s selector or its API response instead of relying on waitForNavigation().

Handling lazy-loaded and asynchronously inserted content

page.content() captures only the state available when it executes. To include content loaded after scrolling, perform the interaction first:

await page.goto('https://example.com/feed', { waitUntil: 'domcontentloaded' });

await page.evaluate(async () => {
  window.scrollTo(0, document.body.scrollHeight);
});

await page.waitForSelector('.feed-item:last-child');
const html = await page.content();

For multiple lazy-load cycles, scroll in a loop and stop when the page-specific completion condition is met. Do not assume that reaching the bottom once loads every item; virtualized lists may remove off-screen nodes, so the current DOM may not contain the entire logical dataset.

page.content() versus outerHTML

Approach Result Best use
page.content() Serialized current page HTML, including the DOCTYPE Normal full-document extraction through Puppeteer
document.documentElement.outerHTML Markup for the current document element, evaluated inside the page Custom browser-side processing before returning markup

The lower-level alternative is:

const html = await page.evaluate(() => document.documentElement.outerHTML);

Use it when you need to transform or inspect the markup in page context. For the ordinary “get the full HTML” task, page.content() is clearer and preserves the DOCTYPE in the documented result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frames and the meaning of “full page”

The main page’s page.content() result represents its document. HTML rendered inside an iframe belongs to a different document and is not merged into the parent’s serialization. Locate the frame and extract it separately:

const frame = page.frames().find(f => f.url().includes('/embedded'));
if (!frame) throw new Error('Embedded frame was not found');

await frame.waitForSelector('.content');
const frameHtml = await frame.content();

Cross-origin restrictions still apply to what browser code can access, and a frame may load later than the parent. Wait on the frame itself, not only on the top-level page.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

The HTML contains only a shell

  • Cause: extraction ran before the application rendered its data.
  • Fix: wait for a content selector, required response, or app-ready flag. Use network idle only when the site’s connections actually settle.

waitForSelector times out

  • Cause: the selector is wrong, the page navigated elsewhere, a consent gate blocks rendering, or the element is inside a frame.
  • Fix: verify the final URL, inspect a screenshot or console output, increase the timeout only when the page is legitimately slow, and search page.frames() for iframe content.

waitForNetworkIdle never completes

  • Cause: polling, analytics, streaming, ads, or open connections keep the network active.
  • Fix: replace it with a selector or response wait, or set a bounded timeout and handle the timeout explicitly.

Navigation races a click

  • Cause: the click was awaited before the navigation wait was registered, or extraction happened before the new route rendered.
  • Fix: use Promise.all with the navigation wait and click, then apply the destination’s readiness condition.

Content is missing from an iframe

  • Cause: the parent document does not include the iframe’s separate DOM.
  • Fix: select the correct frame and call frame.content() after waiting inside it.

The browser process remains running

  • Cause: an exception bypassed browser.close().
  • Fix: put the workflow in try/finally and close the browser in the finally block.

Practical reliability and performance notes

  • Reuse one browser process for a batch of URLs, but create a fresh page when isolation is required.
  • Set explicit navigation and selector timeouts so a broken site cannot stall a worker indefinitely.
  • Capture after the smallest reliable readiness condition; waiting for every request can add latency without improving the HTML you need.
  • Keep the URL, final URL, wait condition, and extraction timestamp with saved HTML so later debugging can distinguish redirects from rendering problems.
  • Large DOMs produce large strings. Stream or write results promptly rather than retaining every page in memory during a crawl.

Or skip the browser setup

If your goal is a visual screenshot rather than the DOM string, ScreenshotNeo provides a single-request capture API. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents, including Claude and Cursor.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for options and response details. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does page.content() return the original server response?

No. It serializes the document currently held by the browser, so client-side changes made after navigation are included.

Can I get HTML before JavaScript runs?

Use the server response or another HTTP client for source HTML. Puppeteer’s page.content() is intended for the current browser DOM after whatever rendering has occurred.

Why is my saved HTML different on two runs?

The page may depend on timing, personalization, random data, consent state, lazy loading, or changing API responses. Use deterministic waits and browser settings where your application permits.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.