Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use browser automation when the data you need appears only after JavaScript runs or requires real interaction. For a static response, an authorized API or a normal HTTP client is simpler, faster and easier to operate. This guide shows a complete Playwright workflow in Python, explains reliable locators, session isolation, robots.txt boundaries, troubleshooting and operating costs, then gives a one-request alternative for screenshot jobs.

Decide whether you need a browser

Start with the least complex method that can legally and reliably obtain the data.

Target situation Preferred method Why
Documented endpoint returns the required fields Authorized API client Structured output, lower latency and no rendering overhead.
Server-rendered HTML contains the data HTTP request plus an HTML parser Usually cheaper and easier to scale than a browser.
Content appears after JavaScript, scrolling or a user action Playwright or another browser automation tool The page is evaluated in a real browser context.
Only a visual capture or PDF is required Screenshot service or browser capture You need a rendered artifact rather than parsed fields.

A browser is an additional tool, not a mandatory component of every scraper. Playwright’s Python library supports Chromium, WebKit and Firefox and can run locally or in continuous integration (CI). It provides synchronous and asynchronous Python APIs.

Install Playwright for Python

  1. Use a supported Python environment and create a virtual environment if this project will run repeatedly.
  2. Install the library:
    python -m pip install playwright
  3. Install the browser binaries you intend to use:
    python -m playwright install chromium

    Install all supported engines instead with python -m playwright install. In CI, include this installation in the image or setup job so a clean runner has the binaries.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The following synchronous example visits an authorized page, waits for the main content, reads a user-visible heading and collects link text. Replace the URL and locators with elements that actually describe your target page.

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/catalog"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(locale="en-US")
    page = context.new_page()
    try:
        page.goto(URL, wait_until="domcontentloaded", timeout=30_000)
        page.get_by_role("main").wait_for(state="visible", timeout=15_000)
        title = page.get_by_role("heading", level=1).inner_text()
        links = page.get_by_role("main").get_by_role("link").all_inner_texts()
        print({"title": title, "links": [x.strip() for x in links if x.strip()]})
    except PlaywrightTimeoutError as exc:
        print(f"Timed out while loading or locating content: {exc}")
    finally:
        context.close()
        browser.close()

Use the asynchronous API when your application already uses asyncio or must coordinate many independent pages. Do not run synchronous Playwright calls inside an event loop; choose one API style per execution path.

Build interactions that survive page changes

Prefer user-facing locators

Playwright recommends locators based on accessible roles and names, labels and visible text. Locators are central to its auto-waiting and retry behavior: Playwright waits for an element to be actionable instead of requiring arbitrary sleeps.

# Good starting points
page.get_by_role("button", name="Load more").click()
page.get_by_label("Search").fill("laptops")
page.get_by_role("link", name="Specifications").click()
page.get_by_text("In stock", exact=True).wait_for()

Use CSS or XPath when the page has no useful accessible structure, but give those selectors a stable attribute such as data-testid. Avoid positional selection as the default. first, last and nth can silently point at a different element after an advertisement, card or navigation item is inserted. If repeated elements are unavoidable, narrow the locator with a stable parent, a label or a unique attribute before using a position.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for a state, not a guessed delay

Use page.goto(..., wait_until="domcontentloaded") for the initial document, then wait for the specific content your extraction needs. For a button that triggers a request, wait for the resulting locator or response:

page.get_by_role("button", name="Load more").click()
page.get_by_role("article").last.wait_for(state="visible")

The last example is appropriate only when the page contract guarantees that the newly appended article is last; otherwise wait for a unique heading or a count change. Network-idle waiting can help on pages that finish rendering through background requests, but analytics and long-polling can prevent it from ever becoming idle. A selector or a bounded timeout is usually more deterministic.

Handle consent and optional UI deliberately

A consent banner, newsletter dialog or chat launcher can cover the element you need. If your authorization and purpose allow it, close the specific overlay with a short timeout and continue when it is absent:

try:
    page.get_by_role("button", name="Accept all").click(timeout=3_000)
except PlaywrightTimeoutError:
    pass  # The banner was not present

page.get_by_role("main").wait_for()

Do not treat a missing button as a fatal error when it is optional. Conversely, do not blanket-click every button: an incorrect consent choice can change what data the page exposes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use browser contexts to isolate sessions

A browser context is an isolated, incognito-like session. Playwright documents that contexts do not share cookies or cache with one another. Create separate contexts for different accounts, locales or test cases, and close each context when its job ends:

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    us = browser.new_context(locale="en-US", timezone_id="America/New_York")
    de = browser.new_context(locale="de-DE", timezone_id="Europe/Berlin")
    try:
        us_page = us.new_page()
        de_page = de.new_page()
        us_page.goto("https://example.com", wait_until="domcontentloaded")
        de_page.goto("https://example.com", wait_until="domcontentloaded")
        # Extract each locale independently.
    finally:
        us.close()
        de.close()
        browser.close()

Persist authentication only for accounts and data you are authorized to use. Session isolation improves reliability and separation; it does not grant permission to access a service.

Robots.txt, terms and responsible access

RFC 9309 standardizes the Robots Exclusion Protocol. Crawlers are requested to honor a site’s robots.txt rules, but the standard states: These rules are not a form of access authorization. Robots.txt is therefore not a substitute for permission, authentication controls, contractual terms or legal analysis.

Google’s documentation describes how Google’s own crawlers download and interpret robots.txt. Attribute those implementation details to Google; do not assume every automated client behaves identically. Before collecting data, check the site’s terms, authentication requirements, privacy and data-rights obligations, geographic restrictions and published rate expectations. Public visibility alone does not establish that automated collection is allowed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep requests proportional: identify your client where appropriate, limit concurrency, cache results, honor explicit prohibitions and stop when the service signals that access should not continue.

Make a scraper reliable in production

  • Reuse a browser, not a process per URL. Launching a browser is expensive; create and close contexts or pages for isolated jobs.
  • Bound every wait. Set navigation and locator timeouts, record the URL and step that failed, and retry only transient failures.
  • Capture evidence on failure. Save a screenshot, HTML snapshot and console/network logs when permitted. These show whether the page was blank, redirected, challenged or simply changed its markup.
  • Throttle concurrency. A small number of pages per browser is easier to monitor and less likely to overload the target than an unbounded task pool.
  • Cache immutable results. Avoid downloading the same page repeatedly; use a content hash or timestamp to decide when a refresh is needed.
  • Prefer structured endpoints discovered during inspection. If the browser calls an authorized JSON endpoint that supplies all required fields, using that endpoint may remove rendering work while keeping the same access boundaries.
  • Pin and review versions. Browser binaries and sites change. Test critical locators after Playwright or browser upgrades.

Common failures and fixes

Symptom Likely cause Fix
Executable doesn't exist Browser binaries were not installed on this machine or CI runner. Run python -m playwright install chromium (or the required engines) in the same environment that runs the script.
Navigation timeout Slow server, a never-ending request, redirect loop or blocked access. Use a realistic timeout, wait for domcontentloaded, inspect the final URL and logs, and do not retry indefinitely.
Locator timeout or strict-mode error The selector matches nothing or several elements after a layout change. Inspect the rendered DOM, choose a role/name or stable attribute, and narrow the parent scope.
HTML loads but data is missing The data is inserted after an interaction or a background request. Trigger the required action and wait for its visible result or a specific response before extracting.
Works headed, fails headless Viewport, timing, permissions or a site challenge differs. Set an explicit viewport, record console errors, compare the final URL and verify that automation is permitted. Do not attempt to bypass an access control.
Repeated jobs leak data between users Pages share cookies or local storage. Create a new browser context per identity or workflow and close it after the job.
Content is a CAPTCHA or bot-check page The service is challenging automated access. Stop and obtain permission or an approved integration; changing selectors will not solve an access decision.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

For screenshot-only jobs: ScreenshotNeo

#1 ScreenshotNeo is the practical first option when you need an image or PDF rather than extracted fields: it removes cookie/consent banners, newsletter popups and chat widgets before capture, bills only clean shots, and its paid entry plan is $5 for 3,000 shots. Learn more at ScreenshotNeo.

It is a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP or PDF. The response identifies page outcomes with X-Page-Verdict and billing with X-Billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing.

One-call examples

See the ScreenshotNeo API documentation for all parameters. cURL:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Capture controls

  • Full-page captures load lazy images; capture a single element by CSS selector; choose dark mode, 12 device presets or any viewport; and set retina scale.
  • Create PDFs with paper size, margins, landscape orientation and page ranges.
  • Render HTML/CSS to an image, inject custom CSS or JavaScript, click an element before capture, and wait for a selector, delay or network idle.
  • Block ads, trackers, requests or resource types; send custom headers, cookies, user agent and Authorization; set timezone and geolocation; use transparent backgrounds and resize images.
  • Choose a cache TTL, generate signed links for public <img> tags, submit asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, query usage, and use the OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.
  • An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Plan Allowance Price
Free 1,000 shots/month Free, no card
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Yearly billing gives two months free, and every feature is available on every plan. If your workflow only needs a rendered artifact, this avoids browser installation, locator maintenance and context management. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card. Start with a free ScreenshotNeo account.

Frequently Asked Questions

Can Playwright scrape a page that requires scrolling?

Yes. Scroll or click the page’s own “Load more” control, then wait for a newly visible, specific locator before extracting. Set a stopping condition so an infinite feed cannot run forever.

Should I use Chromium, Firefox or WebKit?

Use the engine that matches the page behavior you must reproduce, and test other engines when cross-browser output matters. Playwright supports all three; the choice is a compatibility requirement, not a universal performance ranking.

How should I store login state?

Keep credentials and saved authentication state in a protected secret store, limit the state to an authorized account, and use separate contexts for different identities. Never publish a state file containing cookies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the first metric to monitor?

Track successful records per URL, navigation and locator timeout counts, final URLs, response status where available, and the number of retries. These distinguish site changes from infrastructure failures.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.