Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a headless browser when the information you need appears only after JavaScript runs, a user action occurs, or browser state changes the response. For permitted targets, Playwright is a practical way to launch Chromium, wait for the page to become usable, inspect requests, and extract the rendered result. Start with its default bundled browser, then validate a branded Chrome or Edge channel if compatibility matters. Keep browser configuration separate from permission: a proxy, a different headless mode, or a discovered API endpoint does not authorize collection.

What headless browser scraping actually does

A headless browser runs a browser engine without displaying a normal window. It still performs the work that makes many modern sites useful: parsing HTML, executing JavaScript, maintaining cookies and storage, issuing XHR and fetch requests, applying layout, and responding to clicks or form input. Your scraper can then read the DOM after that work finishes or save the resulting page.

This differs from an HTTP-only scraper that downloads a response and parses its original markup. HTTP requests are usually simpler and faster when the data is already present in the response. A browser is justified when the target depends on browser behavior, such as:

  • content loaded by client-side JavaScript;
  • pagination, tabs, filters, or “load more” controls;
  • cookies, local storage, or a login session that you are authorized to use;
  • an interaction that causes the page to request data;
  • the exact rendered text, layout, or screenshot rather than an API response.

Do not launch a browser merely because a site is popular or difficult. Browser execution adds startup time, memory use, synchronization problems, and another class of failures. If a documented, permitted endpoint supplies the same data, an ordinary HTTP client is often the more maintainable choice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a browser mode deliberately

Playwright uses open-source Chromium builds by default for Chromium-based automation and ships a separate Chromium headless shell. Its browser guide also documents an opt-in newer headless mode through the chromium channel. Playwright warns that the shell and newer mode can behave differently.

Option When to start with it Important qualification
Bundled Chromium (default) General automation and scraping where you control the runtime Use the browser version Playwright installs; validate target behavior
Chromium headless shell Lightweight headless execution supplied by Playwright It is not behaviorally identical to the newer headless implementation
New headless mode via chromium channel High-fidelity web-app or extension testing, or when the target requires it Playwright documents differences from the shell; test before switching
Installed Chrome or Edge channel Compatibility with a particular branded browser matters Playwright does not install branded browsers by default

Begin with the default bundled browser. Compare an installed Chrome or Edge channel only when you have a concrete compatibility reason, and record which mode produced each result. Playwright’s headless launch option defaults to true. Its launch API also accepts HTTP and SOCKS proxy settings; those are routing controls, not permission or a guarantee that a request will succeed.

Set up a permitted Playwright scraper

The following Node.js example is a complete starting point for a page you are allowed to access. It waits for a meaningful selector instead of assuming that a fixed delay means the page is ready.

  1. Install Playwright: npm install playwright.
  2. Install the bundled browser: npx playwright install chromium.
  3. Save the script as scrape.mjs, replace the URL and selector, and run node scrape.mjs.
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
  locale: 'en-US',
  timezoneId: 'UTC'
});
const page = await context.newPage();

try {
  await page.goto('https://example.com/catalog', {
    waitUntil: 'domcontentloaded',
    timeout: 45_000
  });
  await page.locator('[data-product]').first().waitFor({ timeout: 15_000 });

  const products = await page.locator('[data-product]').evaluateAll(nodes =>
    nodes.map(node => ({
      name: node.querySelector('.name')?.textContent?.trim() ?? null,
      price: node.querySelector('.price')?.textContent?.trim() ?? null
    }))
  );
  console.log(JSON.stringify(products, null, 2));
} finally {
  await browser.close();
}

Use stable selectors owned by the page, such as a documented test attribute, rather than long CSS paths. Prefer a selector wait, a specific response, or a page-state assertion over waitForTimeout. A fixed delay can be useful as a last resort, but it is either wasteful on fast runs or too short on slow ones.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Interactions and rendered state

Interact only as a normal authorized user would. For example:

await page.getByRole('button', { name: 'Load more' }).click();
await page.locator('[data-product]').last().waitFor();
await page.screenshot({ path: 'catalog.png', fullPage: true });

For authenticated work, create a context with credentials or a previously authorized storage state, protect that state as a secret, and close the context after the job. Do not place passwords or session tokens in source control or logs. Locale, timezone, geolocation, viewport, user agent, cookies, and extra headers can change what the site returns; set them only when they reflect a legitimate test or collection requirement.

Inspect browser network activity

Playwright can monitor and modify HTTP and HTTPS traffic generated by a page, including XHR and fetch requests. Network observation is a diagnostic: it helps you determine whether the page receives data in its initial document or through later browser requests.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();

page.on('request', request => {
  const type = request.resourceType();
  if (type === 'xhr' || type === 'fetch') {
    console.log('REQUEST', request.method(), request.url());
  }
});

page.on('response', async response => {
  const type = response.request().resourceType();
  if (type === 'xhr' || type === 'fetch') {
    console.log('RESPONSE', response.status(), response.url());
  }
});

await page.goto('https://example.com/app', { waitUntil: 'networkidle' });
await page.waitForTimeout(1_000); // replace with a page-specific readiness check
await browser.close();

networkidle means the browser observed a period with no active network connections; it is not proof that every application component is ready. Prefer a known element or response when possible. Log status and URL, but redact authorization headers, cookies, personal data, and query parameters that contain secrets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An endpoint seen in DevTools or a Playwright event is not automatically a stable, public, or authorized API. Check the site’s documentation and terms, and honor authentication, rate limits, and contractual restrictions. If you do use a permitted endpoint, an HTTP client may be more efficient than rendering every page.

Robots.txt, permission, and security are different questions

RFC 9309 describes robots.txt as requested crawler instructions and states: “These rules are not a form of access authorization.” Google likewise explains that robots.txt does not enforce crawler behavior or secure a page; a disallowed URL can still be indexed when other pages link to it. Password protection and other actual access controls are appropriate for private content. Google’s advice about indexing is specific to Google Search and is not a universal legal rule.

Before running a job, evaluate three separate layers:

  • Permission: your contract, account rights, the site’s terms, and applicable law.
  • Technical controls: authentication, authorization, rate limits, bot challenges, and network controls.
  • Crawler guidance: robots.txt and any published scraping policy.

Follow applicable instructions and stop when access is denied. A proxy option, user-agent change, or browser mode does not bypass those responsibilities or justify evading a restriction. This guide cannot determine the legal status of scraping in a particular jurisdiction or for a particular site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability and performance practices

Make waits observable

Record the URL, browser mode, start and end times, final URL, response status, and the selector or response that declared success. Capture a screenshot and HTML snapshot on failure when permitted. These artifacts distinguish a selector change from a timeout, redirect, consent wall, or empty application state.

Control concurrency

Each browser and page consumes resources. Reuse a browser process, create isolated contexts for separate sessions, and cap concurrent pages according to the machine and the site’s published limits. Add bounded retries with backoff for transient navigation failures; do not retry authorization failures or a deliberate block indefinitely.

Keep extraction deterministic

Normalize whitespace and numbers after extraction, preserve the source URL and retrieval time, and deduplicate by a stable key. Treat missing fields as missing rather than silently copying stale values. If content is personalized or region-dependent, store the context settings alongside the record.

Troubleshooting common failures

Symptom Likely cause Fix
Browser executable not found Playwright’s browser was not installed in the current environment Run npx playwright install chromium and ensure the same user or container is running the script
Timeout waiting for a selector Wrong selector, slow data, redirect, consent screen, or blocked request Inspect the final URL and screenshot, verify the selector in the rendered page, and wait for a page-specific readiness signal
HTML is present but data is empty Data arrives through XHR/fetch or requires an interaction Attach request/response listeners, click the required control, and wait for the resulting element or response
Works headed but not headless Different browser mode, viewport, timing, or target behavior Compare bundled mode with the documented newer headless mode or an installed browser channel; do not assume equivalence
Frequent 403, CAPTCHA, or bot page The site is restricting automated access Stop and verify authorization and the site’s policy; do not treat proxy rotation as permission
Network logs contain secrets Headers, cookies, or query strings were logged wholesale Redact sensitive values before storage and restrict diagnostic files
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your goal is a rendered image or PDF rather than structured records, ScreenshotNeo provides a website screenshot API and MCP server. A single request returns PNG, JPEG, WebP, or PDF. It accepts the cookie or consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for parameters. Its 63 options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work for easier migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Every feature is on every plan, and yearly billing gives two months free. Start with 1,000 free screenshots a month—no card required.

FAQ

Is headless scraping invisible?

No. Sites can observe requests and may identify automation. Headless means no visible window, not anonymity or guaranteed access.

Should I use Playwright’s Chromium or installed Chrome?

Use bundled Chromium first. Validate an installed Chrome or Edge channel when the target’s browser-specific behavior requires it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can network inspection replace an API agreement?

No. Seeing an XHR or fetch URL explains page behavior but does not establish that the endpoint is public, stable, or authorized for your use.

Does robots.txt protect private data?

No. It requests crawler behavior; it is not an access-control mechanism. Protect private content with authentication and authorization.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.