What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a real browser only when the data you need appears after JavaScript runs. A practical Node.js crawler can launch Chromium with Playwright, wait for a page-specific readiness signal, extract selected fields, record provenance, and close resources safely. For pages whose useful content is already in the response HTML, an HTTP parser is simpler and lighter.

This guide builds a responsible rendered crawler, explains when to avoid browser automation, and covers installation, extraction, retries, browser choice, robots.txt, troubleshooting, and a browser-free API option.

Choose rendering only when the page requires it

An ordinary HTTP client downloads HTML but does not execute the JavaScript that hydrates a single-page application, loads an article from an API, or reveals content after interaction. Crawlee describes its CheerioCrawler as fast and efficient for plain HTTP/HTML work, but unable to handle JavaScript rendering. Its browser-backed choices are PlaywrightCrawler and PuppeteerCrawler. The Crawlee quick start recommends Playwright for a new headless-browser project (Crawlee quick start).

First compare the initial response with what a normal browser displays. If the title, links, and fields you need are present in the returned HTML, use an HTTP parser. If they appear only after scripts execute, use a browser. Do not render every URL by default: browser startup, compatible binaries, memory use, and page execution add operational complexity without improving an HTML-only crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick decision table

Requirement Best starting point Reason
Content is in initial HTML HTTP parser or Crawlee CheerioCrawler Less setup and no JavaScript execution
Content is inserted by client-side JavaScript PlaywrightCrawler or a direct Playwright script Runs the page in a real browser engine
Existing Puppeteer codebase PuppeteerCrawler or Puppeteer Retains the project’s familiar API
Need Chromium, Firefox, and WebKit coverage Playwright Playwright documents all three engines and branded Chrome/Edge options

A rendered result is not proof that your crawler behaves like Googlebot or that a search engine will index the same content. Google documents JavaScript rendering, robots.txt, sitemaps, canonicalization, and crawl management as separate concerns (Google Crawling and Indexing).

Install Node.js, Playwright, and a browser

Crawlee’s current quick start states Node.js 16 or later; treat that as a source-specific, changeable requirement and verify the current documentation before deployment. You can scaffold a Crawlee project with:

npx crawlee create my-crawler

For a small focused crawler, install Playwright directly:

mkdir rendered-crawler
cd rendered-crawler
npm init -y
npm install playwright

Playwright releases are coupled to specific browser binary versions. Install the supported browsers after installing the package, and repeat the command when upgrading Playwright:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
npx playwright install chromium

Playwright also supports Firefox and WebKit; install the engines your target sites require. On supported Linux environments, use Playwright’s documented dependency-install option when the operating system lacks required libraries (Playwright browsers). Branded Chrome and Edge can be used when installed locally or through the documented CLI options, but do not assume a system browser is compatible with every Playwright release.

Build a rendered crawler with Playwright

The following program accepts URLs, opens one browser, creates an isolated page for each URL, waits for a selector that represents the application’s actual content, extracts a title and headings, and writes structured JSON. It deliberately does not treat one lifecycle event as universal readiness.

import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';

const urls = process.argv.slice(2);
if (urls.length === 0) {
  console.error('Usage: node crawl.mjs https://example.com https://example.org');
  process.exit(1);
}

const browser = await chromium.launch({ headless: true });
const results = [];

try {
  for (const url of urls) {
    const startedAt = new Date().toISOString();
    const context = await browser.newContext();
    const page = await context.newPage();

    try {
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: 45_000
      });

      // Replace this with a selector that means “content is ready” on your site.
      await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 });

      const data = await page.evaluate(() => ({
        title: document.title,
        headings: [...document.querySelectorAll('h1, h2')]
          .map((node) => node.textContent?.trim())
          .filter(Boolean),
        text: document.querySelector('main')?.innerText?.trim() ?? ''
      }));

      results.push({
        url,
        fetchedAt: startedAt,
        status: response?.status() ?? null,
        ...data
      });
    } catch (error) {
      results.push({
        url,
        fetchedAt: startedAt,
        error: error instanceof Error ? error.message : String(error)
      });
    } finally {
      await context.close();
    }
  }
} finally {
  await browser.close();
}

await writeFile('results.json', JSON.stringify(results, null, 2));

Save it as crawl.mjs and run node crawl.mjs https://example.com. The output records the source URL, crawl time, HTTP status when available, extracted fields, or an error. Keeping provenance with each record makes later review and deduplication possible.

Why the readiness signal matters

domcontentloaded means the initial document has been parsed; it does not guarantee that a framework has fetched data or finished rendering. Puppeteer and Playwright both expose page lifecycle and request events, but their documentation does not claim that one event is correct for every application (Puppeteer Page API; Playwright Page API).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Wait for a stable selector such as main article when that element appears only after data loading.
  • For a known API request, wait for the response and validate its status before reading the DOM.
  • Use a short, site-specific delay only when the page has no better signal; fixed sleeps make crawls slower and still do not prove readiness.
  • Use network-idle waiting cautiously: analytics, advertising, or long-lived connections can prevent it from occurring.

Extract narrowly and safely

  • Read only fields needed for the stated purpose rather than copying an entire page.
  • Normalize whitespace and preserve the original URL.
  • Keep extraction selectors configurable because a site redesign can invalidate them.
  • Do not execute arbitrary links or submit forms unless the crawl requires it and the site permits it.
  • Close each context and the browser in finally blocks so failures do not leak processes.

Use Crawlee when you need queues and crawler controls

A direct Playwright script is easy to understand and suitable for a small URL list. Crawlee adds URL queues, request handling, retries, concurrency controls, and crawler statistics around browser automation. Its interface lets a project choose PlaywrightCrawler or PuppeteerCrawler while retaining a similar crawler structure. Install the package and the browser integration according to the current Crawlee documentation; Playwright and Puppeteer are not bundled automatically and require explicit installation.

Use a queue when discovering links, set a conservative concurrency for the target site, and cap retries. Browser concurrency is an operational setting, not a promise of speed: the available sources do not establish a universal throughput or cost comparison. Measure your own pages, network, CPU, and memory conditions.

Respect robots.txt and site boundaries

Before crawling, identify the site’s published policy, terms, authentication requirements, and applicable law. Check https://host.example/robots.txt for crawl directives, avoid excessive rates, and use a descriptive user agent when appropriate. Google explains that robots.txt controls which URLs a crawler may request, but rules cannot enforce behavior against every crawler and disallowed URLs may still appear in search results (Google’s robots.txt guide).

robots.txt is neither authentication nor a security boundary. Use password protection for private material. If your goal is search visibility, use documented controls such as noindex rather than assuming a disallow removes a URL from results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability, performance, and cost decisions

Reuse the browser, isolate pages

Launching one browser per URL is wasteful. Reuse a browser process, create a fresh context per site or tenant when isolation matters, and close contexts promptly. Keep a page limit that fits your machine; each rendered page can load scripts, images, fonts, and third-party resources.

Control navigation work

  • Set explicit navigation and selector timeouts.
  • Block unnecessary images, fonts, ads, or trackers only when doing so does not change the content you need.
  • Prefer a narrow extraction selector over serializing the entire DOM.
  • Cache results when freshness permits, and record the cache decision.
  • Retry transient network failures with a bounded count; do not retry permanent authorization or policy errors indefinitely.

Expect partial results

Pages can return a successful HTTP status while displaying an application error, consent wall, bot check, or empty shell. Validate the extracted fields and classify failures separately: timeout, navigation error, blocked access, missing selector, parse failure, and empty content. A browser does not guarantee access, permission, or successful extraction from every site.

Troubleshoot common failures

“Executable doesn’t exist” or browser launch failure

Install the browser binary with npx playwright install chromium. After upgrading Playwright, install again because browser versions are release-specific. In Linux containers, install the operating-system dependencies documented by Playwright.

The selector timeout expires

Inspect the page manually and confirm the selector is correct for the current route. The content may be behind login, a consent dialog, a bot check, or an iframe. Wait for a meaningful API response or application-specific state instead of increasing the timeout without evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The HTML contains no data

Check whether the data is in an iframe, loaded from an API, or rendered only after scrolling or clicking. Wait for the relevant request or interaction, then verify the DOM. If the server returns an error shell, record it rather than treating an empty extraction as valid.

Navigation hangs

Use an explicit timeout, avoid waiting for network idle on pages with persistent connections, and choose domcontentloaded followed by a content selector. Capture the final URL and status so redirects and error pages are visible.

Results differ from a normal browser

Compare viewport, timezone, locale, cookies, authentication, user agent, and browser engine. Some sites intentionally vary content or challenge automation. Do not attempt to bypass access controls; obtain permission or use an authorized API.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server for developers. A single GET request returns PNG, JPEG, WebP, or PDF, while its capture pipeline accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Each response identifies whether it was a clean page and whether it was billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a direct screenshot, see the ScreenshotNeo API documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It supports full-page captures, element selectors, dark mode, device presets, custom JavaScript and CSS, clicks, waits, request blocking, headers, cookies, authorization, geolocation, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and a usage API. These options are useful when you need rendered visual evidence rather than a custom DOM data model.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.

FAQ

Does headless mode hide JavaScript from a site?

No. Headless mode changes how the browser is displayed, not the site’s authorization, bot-detection, or terms. A target can still block or challenge automation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use Puppeteer or Playwright?

Choose the API your team can maintain and the browser engines your targets require. Crawlee supports both crawler types; Playwright’s documented engine coverage is broader.

Can robots.txt tell me whether I may collect personal data?

No. It is a crawl-policy signal, not a legal authorization or privacy assessment. Review the site’s terms, applicable law, and data-minimization requirements separately.

How do I keep a crawl reproducible?

Pin package versions, install the matching browser binaries, store selectors and configuration in source control, and record URL, timestamp, browser choice, status, and extraction errors with each result.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.