Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Puppeteer to let the page’s JavaScript render, wait for the specific element or value you need, then extract it in the browser context with $eval, $$eval, or evaluate. Avoid guessing with a fixed sleep: a wait tied to the target data is usually more reliable and makes timeouts easier to diagnose.

Why JavaScript-generated values are missing from the initial HTML

A website may return a basic document first and fill in prices, totals, search results, or other values later using JavaScript. Reading only the initial response can therefore miss data that appears after scripts run. Puppeteer controls a real browser context, so page scripts can execute before your extraction code reads the rendered DOM. The official project describes Puppeteer as a JavaScript library for controlling Chrome or Firefox over the DevTools Protocol or WebDriver BiDi: Puppeteer documentation.

The key is to synchronize extraction with the page’s actual state. Navigating to a URL does not by itself guarantee that a particular value has been rendered. Identify where the value appears, wait for that condition, and validate what you retrieve.

Set up Puppeteer and navigate to the page

Install Puppeteer in a Node.js project using the official installation instructions, then create a script that launches a browser, opens a page, navigates to your target, waits for the data, and closes the browser. This example uses a fictional product page with a data-price attribute; replace the URL and selector with those from the site you are authorized to access.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage();
  await page.goto('https://example.com/product', {
    waitUntil: 'domcontentloaded'
  });

  await page.waitForSelector('[data-price]', { visible: true });
  const price = await page.$eval(
    '[data-price]',
    element => element.textContent?.trim() ?? ''
  );

  if (!price) {
    throw new Error('The price element was found but its text was empty');
  }

  console.log(price);
} finally {
  await browser.close();
}

domcontentloaded is a navigation milestone, not proof that client-side rendering has finished. The following selector wait handles that separate readiness condition. Puppeteer’s getting-started guide covers launching or connecting to a browser, creating pages, and using its API: Getting started.

Choose the right wait for the data

The best wait is the one that describes what must be true before extraction. Puppeteer’s documented default timeout for waitForSelector is 30 seconds; setting timeout: 0 disables the timeout. If a selector does not appear before the timeout, the call throws rather than silently returning no match. See the Page.waitForSelector API reference.

Wait method Use it when Important limitation
waitForSelector The element containing the result is inserted into the DOM, or you need it to be visible or hidden. A matching element can exist before its text or attribute is populated. In that case, wait for the value itself.
waitForFunction The node exists early, but its text, attribute, or application state changes later. Write a predicate that checks the actual expected condition; an overly broad predicate can resolve too soon.
waitForNetworkIdle You need network quiescence as an additional synchronization signal. It measures network activity, not whether the application has rendered your target value. Polling, analytics, WebSockets, and lazy loading can also make it late or unsuitable.
Fixed delay Only as a last-resort workaround when the site provides no observable readiness signal. A guessed delay may be too short on a slow run and waste time on a fast one.

Wait for an element with waitForSelector

Use this when the target node itself appears after rendering. Pass { visible: true } when the value must be visible to a visitor. Use { hidden: true } when you need to wait for an element to disappear or become hidden. Choose a bounded timeout appropriate for the page and handle a timeout as an explicit failure rather than extracting stale or missing data.

await page.waitForSelector('[data-total]', {
  visible: true,
  timeout: 15000
});

Wait for text or an attribute with waitForFunction

If the element is present from the start but its content is populated asynchronously, wait until the content meets a useful condition. This predicate checks for non-empty text in a value-bearing element:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.waitForFunction(() => {
  const element = document.querySelector('[data-total]');
  return element?.textContent?.trim().length > 0;
}, { timeout: 15000 });

You can apply the same pattern to an attribute, such as data-state, or to a more specific page condition. The predicate runs in the page context. Keep it self-contained or pass needed values as function arguments; Node.js variables are not automatically available inside the browser page. Page.evaluate can also run a function in the page context and wait for its returned Promise to resolve. See the Page.evaluate API reference.

Use network idle as supporting evidence, not a universal finish line

waitForNetworkIdle waits for network activity to become idle and always waits at least the configured idle time. That may be useful when requests are expected to settle, but a quiet network does not establish that a particular application value is correct or even present. Conversely, ongoing polling or long-lived connections may prevent a useful idle point. Prefer a selector or predicate tied to the data; add network-idle waiting only if it helps with a known page behavior. See the Page.waitForNetworkIdle API reference.

Extract one value, a list, or an attribute

One element with $eval

Once the readiness condition succeeds, page.$eval selects a matching element and runs the supplied function against it in the browser context. Extract the property you actually need: textContent for text, getAttribute() for an attribute, or a DOM property when the page exposes the value there.

const result = await page.$eval(
  '[data-price]',
  element => ({
    text: element.textContent?.trim() ?? '',
    raw: element.getAttribute('data-price') ?? ''
  })
);

Keep the returned value serializable. A DOM node itself is not a useful value to return to Node.js; return strings, numbers, arrays, or plain objects built from the node instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Multiple matching elements with $$eval

For repeated rows, use page.$$eval to collect and transform all matching elements in one browser-context call. The function receives the matching nodes and should return ordinary data:

const rows = await page.$$eval('[data-row]', nodes =>
  nodes.map(node => ({
    name: node.querySelector('.name')?.textContent?.trim() ?? '',
    value: node.getAttribute('data-value') ?? ''
  }))
);

console.log(rows);

The Page.$$eval API reference documents the matching-elements form. For robust selectors, prefer stable semantic attributes, labels, or roles if the site provides them instead of styling classes that may change during a redesign.

Read a value using evaluate

Use page.evaluate when extraction needs several DOM operations or a page-context calculation. For example, it can read a group of related values or derive a number from rendered text. Keep the function independent of Node-only objects and pass dynamic inputs explicitly.

Find selectors that survive frontend changes

Inspect the rendered page and choose the narrowest selector that identifies the intended data, not merely the first visually similar element. Puppeteer supports CSS selectors as well as selector features for text, accessibility attributes, XPath, and shadow-root traversal; the page interactions guide explains the available strategies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Prefer an ID, data attribute, accessible label, or role that describes the content.
  • Use a CSS class only when it is stable enough for your task; presentation-only class names can change without changing the data.
  • For repeated content, scope child lookups to each row so that a name from one row is not paired with a value from another.
  • If the value is inside a shadow root, use a supported shadow-root selector strategy rather than assuming ordinary document queries will cross the boundary.
  • If the target is in an iframe, identify the relevant frame and perform the wait and extraction there.

Validate the result before saving it

A successful selector match does not guarantee a usable value. Trim whitespace, reject empty strings, and check the format your downstream code expects before writing to a database or file. A price might need to match the site’s displayed currency format; a numeric field might need parsing and validation rather than blind conversion.

const raw = await page.$eval(
  '[data-price]',
  element => element.textContent?.trim() ?? ''
);

if (!raw) {
  throw new Error('Price was empty');
}

if (!/d/.test(raw)) {
  throw new Error(`Unexpected price format: ${raw}`);
}

Preserve the original text when formatting may contain locale-specific separators or currency symbols. If you need a normalized number, define and test the conversion rules for the target site and locale rather than assuming every displayed value uses the same format.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot empty or missing values

The selector wait times out

  • Confirm the selector against the DOM after JavaScript has run; the initial response may not contain the rendered node.
  • Check whether the value is in an iframe or shadow root rather than the top-level document.
  • Determine whether a click, scroll, consent action, pagination step, or other interaction is required before the site inserts the value.
  • Verify that the page reached the intended URL and that navigation did not fail or land on a different page.

The element exists but extraction is empty

Wait for a non-empty text or attribute with waitForFunction rather than only waiting for the node. Check whether the value is stored in an attribute or property instead of visible text, and make sure the selector identifies the data-bearing element rather than a wrapper.

Network idle never arrives or arrives too late

Do not treat network quiescence as mandatory if the site keeps polling or maintains long-lived connections. Replace it with a selector or predicate that reflects the target data, and use a bounded timeout so failure is visible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
The SQL Programming Language: .
  • Used Book in Good Condition

Compare the rendered DOM with what the script expects

During debugging, capture a screenshot or inspect await page.content() after the relevant wait. Compare that rendered DOM with the selector and expected value. If the value is still absent, inspect the frame, shadow root, and interaction sequence rather than repeatedly increasing a fixed delay.

Performance, reliability, and responsible use

Each Puppeteer run controls a browser page, so use only the browser work your extraction requires. Avoid unnecessary waits after the target condition is satisfied, and avoid loading resources or pages you do not need when the site and your use case permit it. Reuse the readiness condition and extraction logic consistently, but keep timeouts bounded so stalled pages do not leave jobs waiting indefinitely.

Reliability depends on the target site as well as the script: front-end changes can invalidate selectors, network or rendering delays can trigger timeouts, and access may depend on cookies, authentication, or an interaction. Check the site’s terms and applicable rules before collecting data, and do not use automation to bypass access controls. Treat a failed wait as a signal to inspect page state, not as proof that the data does not exist.

Or skip the browser setup

If the task is to capture a page as an image or PDF rather than extract a structured value, ScreenshotNeo offers a website screenshot API and MCP server for developers. A GET request can return a PNG, JPEG, WebP, or PDF; it is not a replacement for Puppeteer code that needs to read arbitrary values from the DOM. Its clean-capture options accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and failed or blank captures are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client. See ScreenshotNeo and the API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Replace YOUR_API_KEY with your key and change the target URL. The request returns an image file; use Puppeteer when you need to extract a rendered text value or structured data. ScreenshotNeo includes 1,000 shots per month on its free plan with no card required; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo to try the free monthly allowance.

Frequently Asked Questions

Can Puppeteer scrape a value that is not present in the original page source?

Yes. Puppeteer runs the page in a browser context, allowing its JavaScript to populate the DOM before you wait for and extract the value.

Which should I use: waitForSelector or waitForFunction?

Use waitForSelector when the node’s presence or visibility is the readiness condition; use waitForFunction when the node exists but its content, attribute, or state must change first.

Does waitForNetworkIdle guarantee that a page is ready?

No. It indicates network quiescence, not application readiness. A target-specific selector or value predicate is a more direct signal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.