Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For JavaScript-heavy websites, Playwright can run the page’s client-side code in a real browser before you collect data. The reliable approach is to wait for the specific content or response you need, prefer semantic locators for rendered content, and validate what you extracted. When the page already receives the records in a suitable API response, capturing that response is often simpler and less brittle than rebuilding the data from the DOM.

How to scrape a JavaScript-heavy website with Playwright

A browser scraper follows the site’s normal page-loading path: create a browser context, navigate to a page, wait for a meaningful readiness condition, collect data from the rendered page or a response, validate the result, and close the browser resources. A context isolates cookies and storage for a job; it is useful to start each independent job with a fresh one.

The example below uses Node.js and Playwright’s JavaScript API. It illustrates a DOM-based approach for pages whose product records appear as article elements with headings and prices. Those locators are examples, not universal selectors: inspect the target site and adapt them to its accessible roles or stable markup. Use only on pages and data you are permitted to access.

  1. Install: In a new project, run npm init -y, then npm install playwright and npx playwright install chromium.
  2. Save the script: Put the following in scrape.mjs. Replace the example URL and, if needed, adjust the locators and expected count.
  3. Run: Use node scrape.mjs. The script prints JSON records; it exits with an error if it finds none.
import { chromium } from 'playwright';

const url = process.env.TARGET_URL ?? 'https://example.com/products';
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();

try {
  const page = await context.newPage();
  page.setDefaultNavigationTimeout(30_000);
  page.setDefaultTimeout(10_000);

  await page.goto(url, { waitUntil: 'domcontentloaded' });

  // Wait for the page-specific signal, not an arbitrary sleep.
  const cards = page.getByRole('article');
  await cards.first().waitFor({ state: 'visible' });

  const records = await cards.evaluateAll((elements) =>
    elements.map((element) => ({
      title: element.querySelector('h1, h2, h3')?.textContent?.trim() ?? null,
      text: element.textContent?.trim() ?? ''
    }))
  );

  const cleaned = records
    .filter((record) => record.title)
    .map((record) => {
      const price = record.text.match(/$s?d+(?:.d{2})?/);
      return { title: record.title, price: price?.[0] ?? null };
    });

  if (cleaned.length === 0) {
    throw new Error(`No product records found at ${url}; check the readiness condition and locators.`);
  }

  console.log(JSON.stringify(cleaned, null, 2));
} finally {
  await context.close();
  await browser.close();
}

The script waits for the first article card to become visible, then reads the cards present at that point. If the site loads records in batches, that first visible card is not proof the list is complete: add a completion condition appropriate to that site before enumeration. If the page exposes a clear total, wait for that count; if it uses pagination, request and validate each page; if it uses infinite scroll, wait for a count or page response to stabilize before reading.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should you use locators or CSS selectors?

Prefer Playwright’s user-facing locators when the page exposes useful semantics: getByRole, getByText, getByLabel, getByPlaceholder, getByAltText, getByTitle, and configured test IDs. For example:

const cards = page.getByRole('article');
const firstTitle = cards.first().getByRole('heading');
const firstPrice = cards.first().getByText(/$d+/);

These locators describe what a person or a test can identify, rather than a particular chain of container elements. Playwright describes locators as the central piece of its auto-waiting and retryability. A locator is resolved when it is used, so Playwright can find the current matching element after a re-render instead of relying on a stale element reference.

CSS and XPath selectors are still useful when a page has no stable accessible name or explicit test contract, or when you need to target a specific data attribute. Prefer stable attributes over generated class names and long paths tied to the current layout. A selector such as [data-product-id] is generally easier to reason about than a chain of positional selectors, but only use it if the target page actually provides it.

How to wait for dynamic content without sleep()

Wait for the condition that means the data you need is ready. Locator actions perform checks such as visibility and enabled state; explicit locator waits, count assertions, URL conditions, and response waits can express other readiness signals. Examples:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.getByRole('heading', { name: 'Results' }).waitFor();

// In a Playwright Test test, this retries until the count matches or times out:
await expect(page.getByRole('article')).toHaveCount(20);

// Register before the action that causes the request:
const responsePromise = page.waitForResponse((response) =>
  response.url().includes('/api/products') && response.ok()
);
await page.getByRole('button', { name: 'Load products' }).click();
const response = await responsePromise;

The assertion example uses expect from Playwright Test; if you are writing a plain Node script rather than a test, use a bounded polling loop or the API-response approach instead of copying that line without adding the test package and import.

Navigation offers wait states including commit, domcontentloaded, load, and networkidle. Do not treat networkidle as a universal signal that the page is ready for extraction: analytics, polling, streaming, or other persistent requests can keep a page active after the target content is available. Conversely, an idle network does not prove that a particular record list finished rendering. Wait for the actual heading, expected result count, URL change, or matching response instead.

Use timeouts as bounds on a real wait, not as a substitute for knowing what you are waiting for. Fixed sleeps add delay when a page is fast and can still be too short when it is slow. For dynamic lists, do not call locator.all() while the list is still changing: it returns immediately and can produce an incomplete or unpredictable result. Establish a stable count, page transition, or other completion condition first.

Can you capture the API response instead of scraping HTML?

Yes, when the page gets the records from a response that is appropriate for your use. Network extraction is often more stable than reconstructing structured records from rendered text because the response may already contain fields in a machine-readable form. DOM extraction is the better fit when the final visible state matters—for example, content revealed by an interaction or assembled from several requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To capture a response, identify the relevant request in the browser’s network activity, then wait for that response before triggering the page action that causes it. Check that the response succeeded, parse the expected format, and validate the shape before using the records. This example assumes the endpoint returns a JSON array with product objects; change the URL match and schema checks to fit the actual page:

import { chromium } from 'playwright';

const url = 'https://example.com/products';
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();

try {
  const page = await context.newPage();
  page.setDefaultTimeout(10_000);
  await page.goto(url, { waitUntil: 'domcontentloaded' });

  // Replace this with the exact, authorized endpoint identified for the page.
  const responsePromise = page.waitForResponse((response) =>
    response.url().includes('/api/products') && response.request().method() === 'GET'
  );
  await page.getByRole('button', { name: 'Load products' }).click();
  const response = await responsePromise;

  if (!response.ok()) {
    throw new Error(`Products request failed: HTTP ${response.status()} ${response.url()}`);
  }
  const payload = await response.json();
  if (!Array.isArray(payload)) {
    throw new Error('Expected the products response to be a JSON array.');
  }
  console.log(JSON.stringify(payload, null, 2));
} finally {
  await context.close();
  await browser.close();
}

The listener must be set up before the click or other action that triggers the request; otherwise, a fast response can arrive before the wait begins. If the endpoint is called during initial navigation, create the response wait before navigation. Keep the endpoint, status, and relevant request details in logs so that a future schema or site change can be diagnosed. An observed endpoint is not automatically authorized for every use; apply the same access and legal review as for the page.

Playwright versus direct HTTP requests

Use a browser when the data depends on JavaScript execution, user interactions, or the final rendered state. A direct HTTP request can have lower overhead when an authorized endpoint already returns the data you need and does not require browser behavior. The choice is not a claim that one method is always faster or more reliable: it depends on where the desired data lives and what access is permitted.

When a browser is necessary, avoid adding browser work that does not improve the result. Wait for a narrow readiness signal, extract only required fields, and close the context when finished. For larger jobs, a single script may no longer be enough operational structure; isolated jobs, capped retries, and logging make failures easier to contain and investigate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a replacement for extracting structured records from a Playwright page. Use it when the deliverable is a clean image or PDF rather than scraped data. A single request can capture a site as PNG, JPEG, WebP, or PDF; its documentation lists the supported options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For screenshot tasks, ScreenshotNeo removes cookie and consent banners from more than 60 known platforms, along with newsletter popups and chat widgets, before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots. See ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Reliability, performance, and cost controls

Production scraping needs more than a selector that works once. The following practices limit common failure modes without claiming that any particular target site or script has been tested:

  • Isolate jobs: Use a fresh browser context for independent jobs so cookies and storage do not leak between them.
  • Bound waits: Set navigation and action timeouts appropriate to the operation. Record which wait timed out rather than retrying blindly.
  • Validate output: Check that the expected fields exist and the record count is plausible. Distinguish a legitimate empty result from a failed or partial page.
  • Retry narrowly: Cap retries and apply them only to steps safe to repeat, such as an idempotent navigation or extraction. Log each attempt and its reason.
  • Track context: On failure, retain the URL, response status where available, and failure reason. That makes selector changes and intermittent loading problems easier to separate.
  • Close resources: Put context and browser cleanup in a finally block so exceptions do not leave pages or browser processes running.

Browser execution has more setup and resource overhead than a direct request. Whether that extra cost is justified depends on whether the site requires JavaScript or interaction to expose the permitted data. There is no general throughput figure here: workload, page complexity, infrastructure, and site behavior vary, so measure your own bounded workload rather than relying on an unsupported universal benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Respect robots.txt and review legal constraints

RFC 9309 defines the Robots Exclusion Protocol: a site publishes crawler instructions at the top-level /robots.txt, organized into user-agent groups and allow/disallow rules matched against URI paths. Before crawling, retrieve the target domain’s file, identify the applicable user-agent group, and honor the most-specific matching rule. The RFC also makes clear that these rules are not access authorization. A robots.txt rule does not grant permission to access restricted content, and its absence does not by itself settle whether a project is allowed.

Review the target site’s terms, authentication requirements, privacy obligations, copyright restrictions, rate limits, and applicable law for the relevant jurisdictions. A technically successful Playwright run does not establish that collection or later use of the data is lawful. There is no universal legal answer for every site, dataset, use, or location; seek qualified legal advice when the stakes warrant it.

Troubleshooting common Playwright scraping failures

  • The page loads but the locator times out: Confirm the locator matches the current page and that the page has reached the relevant state. Prefer the accessible role or label where available; inspect the rendered page and use a stable CSS selector only when no semantic contract fits.
  • The script returns too few records: The list may still be loading, paginated, or expanded by scrolling. Wait for the expected count, next-page response, or another site-specific completion signal before reading the collection.
  • locator.all() returns an inconsistent list: It does not wait for a changing list to settle. Add a condition for stable completion before enumeration.
  • networkidle never arrives: Persistent analytics, polling, or streaming may keep requests open. Wait for the specific content or matching response needed for extraction instead.
  • The response wait hangs or catches the wrong request: Match a more specific URL and, where useful, the request method or response status. Register the wait before the action that triggers the request.
  • JSON parsing fails or fields are missing: Verify the status and content format before parsing, then validate the response schema. The endpoint may have changed, or the response may not be the expected JSON payload.
  • A selector breaks after a redesign: Check whether it depends on generated classes or layout structure. Move to a role, accessible name, test ID, or stable attribute when the page offers one.
  • The script exits with no records: Treat an empty result as a signal to inspect the page state, URL, locator matches, and pagination behavior—not as proof that the source contains no records.

Frequently Asked Questions

Can Playwright scrape a page that requires authentication?

Playwright can interact with browser sessions, but whether you may use a particular account or collect particular data depends on the site’s terms, authorization, privacy requirements, and applicable law. Do not treat successful login as permission to scrape.

Should I use Playwright for every scraping job?

No. If an authorized direct endpoint already returns the needed data and no browser behavior is required, a direct request avoids browser overhead. Choose Playwright when JavaScript execution, interaction, or rendered-state fidelity is needed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.