Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a real browser to open the page, wait for the specific data to appear, locate it with a resilient Playwright locator, and read its text or attributes. Do not treat the page’s load event as proof that the data is ready: modern sites often render or fetch content afterward.

This guide shows a complete Playwright workflow for one element and dynamic lists, explains locator choices and readiness checks, and includes recovery steps for common failures.

What browser automation does

Browser automation controls Chromium, Firefox or WebKit as a user would: it navigates to a URL, waits for the interface to reach a useful state, finds elements, and extracts values. This is different from downloading HTML with an HTTP client because JavaScript, client-side routing, lazy loading and interaction can change what is visible after the initial response.

For this task, the reliable sequence is:

  1. Navigate to the page.
  2. Wait for the target content or state, not merely page load.
  3. Choose a locator that reflects what a user sees or understands.
  4. Extract text or attributes.
  5. For changing lists, prove the list is ready before collecting it.

Install Playwright and create a small scraper

Node.js setup

mkdir website-capture
cd website-capture
npm init -y
npm install -D playwright
npx playwright install

The browser binaries installed by the last command are required on a new machine or CI runner. The following script is an executable example; replace the URL and locators with those for your page.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();

  await page.goto('https://example.com/products', {
    waitUntil: 'domcontentloaded',
    timeout: 30_000
  });

  const heading = page.getByRole('heading', { name: 'Products' });
  await heading.waitFor({ state: 'visible', timeout: 15_000 });

  const firstPrice = page.getByRole('listitem').first().getByText(/$/);
  const priceText = await firstPrice.innerText();
  console.log({ priceText });

  await browser.close();
})();

domcontentloaded only marks an early navigation milestone. The explicit wait for the heading is the readiness condition for this example.

Navigate, then wait for the data you need

Why page load is not enough

A page can finish its initial document while JavaScript is still fetching records, an intersection observer is still triggering lazy content, or a component is still rendering. Playwright’s navigation guidance describes this distinction: a load event does not guarantee that the target data has appeared.

Wait for a meaningful condition instead:

  • A heading, table, card or status message becomes visible.
  • A loading indicator disappears.
  • A known number of rows is present.
  • A specific response has completed when the page’s data request is predictable.

Waiting for a selector or state

await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
const rows = page.locator('table tbody tr');
await rows.first().waitFor({ state: 'visible' });
const count = await rows.count();
console.log(`Rows currently rendered: ${count}`);

If an empty result is valid, wait for either the result container or an explicit “no results” message, rather than waiting forever for a row.

await page.goto('https://example.com/search?q=widget');
await page.locator('[data-testid="results"], [data-testid="no-results"]').first().waitFor({ state: 'visible' });

Choose locators that survive page changes

Playwright recommends user-facing locator attributes such as roles, visible text, labels, placeholders, alternative text and titles. Its documentation calls locators “the central piece of Playwright’s auto-waiting and retry-ability.” A locator is resolved when you use it, so it can continue to work when a framework re-renders the page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preferred locator order

Locator Example Best use
Role and accessible name page.getByRole('button', { name: 'Next' }) Buttons, links, headings, rows and other semantic controls
Label page.getByLabel('Email') Form fields
Placeholder page.getByPlaceholder('Search products') Inputs whose placeholder is stable
Visible text page.getByText('In stock') Distinct content text
Test ID page.getByTestId('product-card') A contract deliberately provided by the site
CSS or XPath page.locator('article[data-id]') When semantic or contract-based locators are unavailable

Avoid long chains such as div:nth-child(2) > div > span. They describe the current DOM implementation, not the content, and are more likely to break after a redesign. CSS and XPath remain useful for unusual structures, but keep them as short and specific as possible.

Extract text from one element

Use locator helpers when you want the rendered value:

const title = await page.getByRole('heading', { level: 1 }).innerText();
const description = await page.getByRole('article').innerText();
const link = await page.getByRole('link', { name: 'Documentation' }).getAttribute('href');

innerText() reflects rendered, visible text. textContent() includes text in descendants even when it is not rendered in the same way. getAttribute() reads attributes such as href, src, data-id and aria-label.

Use evaluate for a precise value

Playwright’s Locator API provides evaluate() for reading one matched element in the page context. The callback receives the element, so you can combine several fields without creating brittle selectors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const product = await page.getByRole('article').first().evaluate((el) => ({
  name: el.querySelector('h2')?.textContent?.trim() ?? null,
  price: el.querySelector('[data-price]')?.getAttribute('data-price') ?? null,
  url: el.querySelector('a')?.href ?? null
}));
console.log(product);

Keep the callback focused on extraction. It runs inside the browser and cannot directly use Node.js modules or variables unless you pass them as arguments.

Extract a dynamic list safely

For a collection, use evaluateAll() or iterate over a locator after readiness is established.

const cards = page.getByRole('article');
await cards.first().waitFor({ state: 'visible' });
const products = await cards.evaluateAll((elements) => elements.map((el) => ({
  name: el.querySelector('h2')?.textContent?.trim() ?? '',
  href: el.querySelector('a')?.href ?? ''
})));
console.log(products);

evaluateAll() maps the elements currently matched. Playwright specifically warns that locator.all() does not wait for matches; when a list changes while it loads, calling it too early can produce incomplete or unpredictable results.

Wait for the list to settle

If the page exposes a loading indicator, wait for it to disappear and then collect the list:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const loading = page.getByText('Loading…');
await loading.waitFor({ state: 'hidden' }).catch(() => {});
const items = page.locator('[data-testid="result-item"]');
await items.first().waitFor({ state: 'visible' });
const records = await items.evaluateAll((els) => els.map((el) => el.textContent.trim()));

For pagination or “load more,” perform the action, wait for the count or a new item to appear, and only then extract:

const items = page.getByRole('listitem');
const before = await items.count();
await page.getByRole('button', { name: 'Load more' }).click();
await page.waitForFunction((oldCount) => document.querySelectorAll('[role="listitem"]').length > oldCount, before);
const allText = await items.evaluateAll((els) => els.map((el) => el.innerText));

Build a repeatable capture function

Separate navigation, readiness and extraction so each target can be tested independently.

async function captureProducts(page, url) {
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
  const result = page.locator('[data-testid="results"]');
  const empty = page.getByText('No products found');
  await Promise.race([
    result.waitFor({ state: 'visible' }),
    empty.waitFor({ state: 'visible' })
  ]);
  if (await empty.isVisible().catch(() => false)) return [];
  return result.locator('[data-testid="product"]').evaluateAll((els) =>
    els.map((el) => ({
      name: el.querySelector('h2')?.textContent?.trim() ?? '',
      price: el.querySelector('[data-price]')?.textContent?.trim() ?? ''
    }))
  );
}

In production, write the URL, timestamp, browser version and extraction errors alongside the output. That makes a changed layout distinguishable from a temporary network failure.

Troubleshooting browser extraction

“Locator timed out”

  • Cause: the locator is wrong, the content is inside a frame, consent UI blocks the page, or the application is slower than the timeout.
  • Fix: inspect the rendered page, verify the role/name or test ID, wait for the actual ready state, and handle frames with page.frameLocator(). Increase a timeout only after correcting the readiness condition.

The script returns an empty list

  • Cause: the list is lazy-loaded or locator.all() ran before matches existed.
  • Fix: wait for a visible item, a loading indicator to disappear, or a count change before using evaluateAll() or all().

Text is present but not visible

  • Cause: textContent includes hidden descendants, while the user sees a different value.
  • Fix: use innerText() for rendered text, or target the visible child directly.

The selector broke after a redesign

  • Cause: a selector depended on DOM nesting or positional indexes.
  • Fix: switch to a role, label, accessible name, stable test ID or short attribute selector.

A navigation hangs or fails

  • Cause: DNS, TLS, authentication, a redirect loop, a blocked resource or a site that never reaches the chosen lifecycle event.
  • Fix: log the final URL, catch and classify the error, set a bounded timeout, and choose the earliest navigation milestone that still precedes your explicit content wait.

Reliability, performance and responsible operation

  • Reuse one browser process and create isolated contexts for batches instead of launching a new browser for every URL.
  • Use bounded navigation and element timeouts so one broken page cannot stall a queue.
  • Capture only the fields you need; evaluateAll() avoids transferring unnecessary markup.
  • Respect authentication, robots policies, terms, rate limits and privacy obligations. Do not bypass access controls or collect personal data without a lawful basis.
  • Record failures separately from empty results. An empty result may be valid; a timeout is an operational error.
  • Expect UI and accessibility-name changes. Keep locator checks in automated tests and update them when the site’s contract changes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean screenshot rather than structured text, ScreenshotNeo returns an image or PDF from one request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. It also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the complete parameter list in the ScreenshotNeo documentation. A direct call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Features include full-page and element capture, dark mode, device presets and custom viewports, retina scale, PDF page controls, custom CSS and JavaScript, clicks, waits, request blocking, headers and cookies, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify migration.

The Free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

When to use Playwright versus a screenshot API

Need Best fit Reason
Structured fields, links or attributes Playwright You control locators and receive data objects.
JavaScript interaction before a visual capture Playwright or ScreenshotNeo Playwright gives programmatic control; ScreenshotNeo provides hosted waits, clicks and scripts.
Clean image or PDF at scale ScreenshotNeo No browser installation, cleanup of common overlays, and usage-based API responses.
AI-agent screenshot workflows ScreenshotNeo Its MCP server exposes screenshot, page-info and PDF tools.

Frequently Asked Questions

Can Playwright extract data that appears after scrolling?

Yes. Scroll the relevant container or page, wait for the newly rendered item or loading state, then extract with a locator. The readiness condition must describe the newly loaded content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use XPath for every scraper?

No. Prefer role, label, text, placeholder, title or stable test-ID locators. Use short CSS or XPath only when those user-facing or contract-based options are unavailable.

What should I save when a capture fails?

Save the URL, final URL if navigation redirected, timestamp, error type and a screenshot or trace when permitted. This separates selector changes from transient network problems.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.