Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can use Puppeteer or Playwright to collect information from pages you are authorized to access, but neither tool makes scraping “stealthy” or guarantees a site will not detect automation. Before running a browser, check the site’s current terms and published access instructions, look for an API or export, keep collection limited, and stop if the site denies access or presents a challenge. The examples below show basic, transparent browser automation—not ways to conceal it or bypass access controls.

What “stealth” means—and what it does not

Browser automation can be detected through multiple signals. In a January 23, 2026 article, Browserless describes detection that can consider inconsistencies among browser fingerprints, network hints, and behavior, and cautions against assuming a plugin will defeat advanced detection. That is Browserless’s characterization, not an independent measurement or a guarantee about any particular site. There is no dependable setting that makes a browser session universally undetectable.

Use “stealth” to mean a careful, low-impact, authorized workflow—not hiding automation from a site. Puppeteer’s security policy says that the code using its capabilities is responsible for ensuring they are used safely and as intended. A technical ability to load and inspect a page is not permission to copy its contents.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check access before collecting data

  1. Prefer an approved route. Check whether the site offers a documented API, data export, or other access channel. If browser access is necessary, obtain permission when the site’s terms or circumstances require it.
  2. Read the actual terms and access instructions. Review the target site’s current terms and any instructions it publishes for automated access. RFC 9309 describes robots.txt rules that crawlers are requested to honor; robots.txt is one input, not a complete answer to legal or contractual questions.
  3. Keep the job narrow. Collect only the fields needed for the permitted purpose, avoid unnecessary pages and repeat visits, and choose conservative request rates. There is no universal rate that is safe for every site.
  4. Identify your automation honestly. Do not disguise a scraper as an ordinary visitor or attempt to work around a site’s controls.
  5. Stop on denial or challenge. If access is blocked, a CAPTCHA appears, or the site otherwise challenges the workload, stop. Seek authorization or an approved access method instead of trying to defeat the control.

Cloudflare’s sample terms, updated May 5, 2026, give site operators example language addressing automated scraping for AI development under stated conditions. Cloudflare says the sample is informational and not legal advice; it is not a universal rule for scrapers or a substitute for the target site’s terms.

Choose an API, a local browser, or hosted infrastructure

Use an API or export when it fits

A documented API or export is usually the first option to investigate: it avoids automating a rendered page when the site already provides an approved way to obtain the data. Check its permissions, scope, and usage limits before building around it.

Use Puppeteer or Playwright locally when browser rendering is needed

Local browser automation gives your code direct control over navigation and page inspection. It is appropriate for a permitted task that genuinely depends on rendered page content. You remain responsible for access compliance, request volume, handling collected data, and deciding when to stop.

Consider hosted browser infrastructure for operational needs

Browserless documents connections for Puppeteer and Playwright, browser sessions, content scraping, and crawl APIs. It may be worth evaluating when you need hosted browser infrastructure, but the cited documentation does not establish comparative performance or pricing. Compare authorization, data handling, session needs, concurrency, reliability, observability, and cost for your workload before choosing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run a basic authorized page extraction

These examples visit one page and read its title and visible text. Replace the example URL with a page you are authorized to access. They do not disguise automation, defeat a challenge, or implement a scraping rate. Install the selected library and its supported browser before running the example; follow the official installation instructions for the version in your project.

Puppeteer (JavaScript)

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });

    const result = await page.evaluate(() => ({
      title: document.title,
      text: document.body.innerText,
    }));
    console.log(result);
  } finally {
    await browser.close();
  }
})();

The finally block closes the browser even if navigation or extraction fails. domcontentloaded waits for the initial document to be parsed; it does not guarantee that every page element or later-loaded result is present. For a permitted page that renders a particular element after navigation, wait for that element explicitly rather than adding an arbitrary long delay.

Playwright (JavaScript)

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });

    const result = await page.evaluate(() => ({
      title: document.title,
      text: document.body.innerText,
    }));
    console.log(result);
  } finally {
    await browser.close();
  }
})();

As with Puppeteer, the navigation event marks a loading milestone, not proof that all dynamically rendered content is ready. For content you are allowed to access, wait for a relevant selector when the page needs more time to render.

Keep extraction focused

For a specific field, query the page for that field rather than storing the entire document. If the value appears only after client-side rendering, wait for the relevant element and handle the case where it never appears. Do not respond to missing content by circumventing a challenge or access restriction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your task is to capture a page image or PDF rather than extract structured text, ScreenshotNeo offers a one-request screenshot API. Its clean-shot options can accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Responses identify page verdict and billing status: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client. Screenshot capture is not a substitute for permission to access a site or a structured-data API.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo includes 1,000 screenshots per month on its free plan with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot without bypassing site controls

  • Navigation times out: The page may be slow, unavailable, or waiting on resources. Check that the URL is correct and the site is accessible through an approved route; use a timeout suitable for your authorized workload. Do not treat a timeout as permission to evade a restriction.
  • Expected text is missing: The page may render content after the initial document load, or the selector may have changed. Inspect the permitted page and wait for the specific element needed. If a challenge or denial is shown, stop rather than trying to get around it.
  • The script hangs or leaves processes running: Ensure browser shutdown runs in a finally block, as in the examples, and inspect errors from launch and navigation separately.
  • Access is denied or challenged: Stop the job. Contact the site owner or use an approved API, export, or other access method.
  • Repeated requests burden the site: Reduce the scope and frequency of the job, avoid revisiting unchanged pages unnecessarily, and confirm the permitted access pattern with the site.

Reliability, performance, and cost considerations

Browser automation can be slower and more resource-intensive than retrieving data through an API because it launches and operates a browser. The actual cost and reliability depend on the pages, workload, infrastructure, and service terms; the sources cited here do not establish a universal benchmark or numeric request rate. Keep jobs small enough to monitor, record failures, avoid needless repeat work, and budget for the browser runtime or any hosted service you select. For hosted options, compare the operational and data-handling requirements directly rather than assuming one provider is faster or cheaper.

Frequently Asked Questions

Does robots.txt give permission to scrape a site?

No. RFC 9309 describes crawler instructions requested to be honored, but robots.txt does not by itself settle a site’s terms, contractual restrictions, or applicable law.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can a stealth plugin guarantee Puppeteer or Playwright will not be detected?

No. Detection can involve multiple signals, and no plugin guarantees that automation will go unnoticed.

Can I use these examples to get past a CAPTCHA?

No. The examples are for permitted page inspection. Stop when a CAPTCHA or other challenge appears and seek authorization or an approved access method.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.