Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a server-rendered page, use TypeScript with a direct HTTP client and an HTML parser such as Cheerio. When the data appears only after JavaScript runs, an interaction is required, or browser state matters, use Playwright. In both cases, wait for a condition that proves the target data is ready—not merely for the browser’s load event—then validate, deduplicate, and persist the result with enough logging to diagnose failures.

Choose the smallest tool that can access the data

Start by looking at the HTML returned by an ordinary HTTP request. If the fields you need are already present, a browser adds operational cost without improving the result. If the response is only an application shell and JavaScript later fetches the records, a real browser is the appropriate escalation.

Situation Recommended approach Reason
Server-rendered HTML and a small number of URLs Built-in fetch or Axios plus Cheerio Low overhead and direct parsing of the returned document.
JavaScript-rendered content, clicks, scrolling, or browser state Playwright Runs a browser and exposes navigation, locators, and page events.
You need to understand redirects and failed resources Playwright request events Lifecycle events reveal what loaded, failed, or redirected.
Many URLs, retries, queues, or proxy controls Crawlee or an equivalent crawler framework Framework orchestration is safer to operate at crawl scale than a hand-written loop.

Do not choose a browser simply because a page looks interactive. First confirm whether the required values are in the initial response.

Plan the scraper before writing selectors

Define the record

Write the output schema first. For example, a product record might contain name, price, currency, sourceUrl, and retrievedAt. Decide which fields are required and what an empty or malformed value means.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check access conditions

Review the site’s terms, any published API, authentication boundary, and /robots.txt before sending requests. RFC 9309 specifies that the rules must be available in a file named /robots.txt at the service’s top-level path. Treat the file as an important access signal, not as a universal legal permission. Also consider privacy, copyright, contractual restrictions, and rate limits.

Use a conservative request policy

Set a clear user agent, bound concurrency, and add backoff for temporary failures. Cache responses where the site’s rules permit it. Keep discovery, extraction, validation, and persistence separate so a selector change cannot silently corrupt stored data.

Scrape server-rendered HTML with TypeScript and Cheerio

Install a TypeScript runner and the parser:

npm install cheerio
npm install -D typescript tsx @types/node

The following program fetches article cards, checks the HTTP status, validates required fields, and writes JSON. Replace the URL and selectors with those for your target.

import * as cheerio from 'cheerio';

interface Article {
  title: string;
  href: string;
  sourceUrl: string;
  retrievedAt: string;
}

const targetUrl = 'https://example.com/news';

function absoluteUrl(value: string, base: string): string {
  return new URL(value, base).href;
}

async function scrape(): Promise<Article[]> {
  const response = await fetch(targetUrl, {
    headers: { 'user-agent': 'ExampleResearchBot/1.0' },
    signal: AbortSignal.timeout(30_000)
  });

  if (!response.ok) {
    throw new Error(`HTTP ${response.status} for ${targetUrl}`);
  }

  const html = await response.text();
  const $ = cheerio.load(html);
  const retrievedAt = new Date().toISOString();
  const rows: Article[] = [];

  $('article.card').each((_, element) => {
    const title = $(element).find('h2').first().text().trim();
    const href = $(element).find('a').first().attr('href');
    if (!title || !href) return;
    rows.push({
      title,
      href: absoluteUrl(href, targetUrl),
      sourceUrl: targetUrl,
      retrievedAt
    });
  });

  if (rows.length === 0) {
    throw new Error('No article cards matched; check for a layout change or JavaScript rendering');
  }
  return rows;
}

scrape()
  .then(rows => console.log(JSON.stringify(rows, null, 2)))
  .catch(error => {
    console.error(error);
    process.exitCode = 1;
  });

Run it with npx tsx scraper.ts. A zero-row result is an explicit failure, not a successful empty crawl. In production, persist the URL, retrieval time, parser version, and selector version with each record.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrape JavaScript-rendered pages with Playwright

Install Playwright and its browser binaries according to the installation instructions for your environment. This example uses locator-based extraction, a page-specific readiness condition, typed callback parameters, and network diagnostics.

npm install playwright
npm install -D typescript tsx @types/node
import { chromium, type Page } from 'playwright';

interface Product {
  name: string;
  price: string;
  sourceUrl: string;
  retrievedAt: string;
}

const targetUrl = 'https://example.com/catalog';

async function scrape(page: Page): Promise<Product[]> {
  page.on('request', request => {
    console.log('request', request.method(), request.url());
  });
  page.on('response', response => {
    if (response.status() >= 400) {
      console.warn('HTTP error', response.status(), response.url());
    }
  });
  page.on('requestfinished', request => {
    console.log('finished', request.url());
  });
  page.on('requestfailed', request => {
    console.warn('failed', request.url(), request.failure()?.errorText);
  });

  await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
  await page.locator('[data-testid="product-card"]').first().waitFor({
    state: 'visible',
    timeout: 30_000
  });

  const retrievedAt = new Date().toISOString();
  const products = await page.locator('[data-testid="product-card"]').evaluateAll(
    (elements): Product[] => elements.map(element => {
      const name = element.querySelector('[data-testid="name"]')?.textContent?.trim() ?? '';
      const price = element.querySelector('[data-testid="price"]')?.textContent?.trim() ?? '';
      return {
        name,
        price,
        sourceUrl: location.href,
        retrievedAt: new Date().toISOString()
      };
    }).filter(product => product.name !== '' && product.price !== '')
  );

  if (products.length === 0) throw new Error('No valid products extracted');
  return products;
}

const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  const products = await scrape(page);
  console.log(JSON.stringify(products, null, 2));
} finally {
  await browser.close();
}

Use stable attributes such as data-testid when the site provides them. Avoid selectors coupled to presentation-only class names. Playwright supports locator APIs and TypeScript annotations in element callbacks; use those types to make schema changes visible during compilation.

Wait for readiness, not just page load

Navigation has several milestones. domcontentloaded means the initial document was parsed; load means the page’s load event fired. Neither guarantees that an application has finished fetching and rendering its data. Modern pages can continue network activity after both events.

Prefer a page-specific condition

  • Wait for a known result locator to become visible.
  • Wait for a response whose URL and status identify the data request.
  • Wait for a short, justified delay only when the application offers no observable condition.
  • For pages that settle predictably, use a network-idle condition cautiously; analytics or long polling can prevent it from completing.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForResponse(response =>
  response.url().includes('/api/products') && response.ok(),
  { timeout: 30_000 }
);
await page.locator('[data-testid="product-card"]').first().waitFor();

A response event alone is not proof of success: a 404 or 503 can complete at the HTTP layer. Check response.ok() or the status code, and treat redirects as data worth inspecting rather than silently following.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make selectors and extraction resilient

  • Keep selectors narrow enough to identify one field, but not tied to incidental nesting.
  • Test against representative variants: pagination, missing images, sold-out items, localization, and logged-out versus logged-in views.
  • Normalize whitespace and currency deliberately; do not parse a localized price with assumptions about decimal separators.
  • Validate required fields and reject or quarantine records that fail validation.
  • Deduplicate using a stable key such as a canonical URL or site identifier.
  • Version selectors and parsers so a later deployment can be traced to a change in output.

Advanced users can register a custom Playwright selector engine, but content-script isolation is safer when page JavaScript could interfere. A custom engine is an optimization for a known need, not a default starting point.

Observe the network while developing

Subscribe to request, response, requestfinished, and requestfailed events during development. These reveal redirect chains, blocked resources, failed API calls, and requests that never finish. Playwright exposes redirectedFrom() and redirectedTo() for tracing a chain.

Log the request URL, method, status, elapsed time, and retry count, but avoid recording unnecessary personal data or credentials. Keep verbose event logging behind a debug setting once the scraper is stable.

Production reliability for TypeScript scrapers

Bound concurrency and retries

Use a queue with a small concurrency limit. Retry transient network errors and selected 5xx responses with exponential backoff and jitter. Do not retry deterministic 4xx responses indefinitely. Give every navigation and extraction operation a timeout and a maximum retry count.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Separate stages

  1. Discover URLs and record the discovery source.
  2. Fetch or navigate with an explicit policy.
  3. Extract into an intermediate object.
  4. Validate, normalize, and deduplicate.
  5. Persist atomically and checkpoint progress.

Scale deliberately

For sustained crawls, evaluate Crawlee or an equivalent framework for queues, retries, and proxy controls. Confirm the current package behavior and any commercial terms before adopting it. A framework does not remove the need to respect site policies or control load.

Cache and resume

Cache immutable responses where permitted and store retrieval timestamps. Persist checkpoints so a process restart resumes instead of repeating the entire crawl. Keep raw responses only as long as your privacy and retention policy allows.

Troubleshooting common failures

Symptom Likely cause Fix
HTML contains no expected records The data is rendered by JavaScript or the selector changed. Inspect the raw response; switch to Playwright if the data arrives later, otherwise update and test the selector.
Playwright returns an empty list Extraction ran before the target component appeared. Wait for a specific locator or successful API response, then validate the count.
Navigation times out Slow resources, a blocked request, a redirect loop, or an unreachable host. Inspect request events, verify redirects and DNS, set a bounded timeout, and retry only transient failures.
A page looks loaded but fields are blank The application replaced placeholder nodes after the load event. Wait for visible, non-empty fields or the data response rather than for load.
HTTP 404 or 503 appears as a completed request HTTP completion is not the same as a successful status. Check the status explicitly and route the URL to retry or failure handling.
Many duplicate records Pagination, retries, or multiple discovery paths overlap. Canonicalize URLs and deduplicate before persistence.
Selectors break after a redesign Selectors depended on visual classes or deep DOM structure. Prefer stable attributes, keep selector versions, and run fixture-based tests.

Legal and responsible scraping

Public visibility is not a blanket license to collect or reuse data. Check terms of service, authentication boundaries, privacy obligations, copyright constraints, and applicable law for your jurisdiction and purpose. Read /robots.txt and follow its applicable rules as an access signal. It controls crawler access; it does not itself remove a URL from search results. Site owners seeking exclusion should use mechanisms such as authentication or noindex, not robots.txt alone.

Identify yourself accurately, send only the requests you need, honor rate limits, and stop when the site indicates that access is not permitted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than structured records, ScreenshotNeo provides a single screenshot API request. It accepts cookie or consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each cleanup step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed.

Use the API documentation at https://screenshotneo.com/docs/ for the complete parameter list. A cURL request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports PNG, JPEG, WebP, and PDF output; full-page capture with lazy images loaded; CSS-selector element capture; dark mode; 12 device presets and custom viewports; retina scale; PDF paper size, margins, landscape, and page ranges; HTML/CSS-to-image; custom CSS and JavaScript; pre-capture clicks; hidden selectors; waits for a selector, delay, or network idle; ad, tracker, request, and resource blocking; custom headers, cookies, user agents, and Authorization; timezone and geolocation; transparent backgrounds; resizing; chosen cache TTLs; signed links for public <img> tags; asynchronous jobs with signed webhooks; bulk capture of up to 100 URLs per call; a usage API; an OpenAPI specification; and compatible parameter names used by other screenshot APIs.

It also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Every feature is available on every plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Allowance Price
Free 1,000 shots/month No card required
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Yearly billing provides two months free. Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.

FAQ

Can I use the same scraper for every website?

No. Rendering model, authentication, localization, pagination, and access rules vary by site. Keep a site-specific adapter behind a shared fetch, validation, retry, and persistence pipeline.

Should I store the complete HTML for every request?

Only when it is justified by debugging, reproducibility, or your retention policy. Otherwise store the extracted record, provenance, and a minimal diagnostic sample while avoiding unnecessary personal data.

When should I move from a script to a crawler framework?

Move when queue management, resumability, retries, proxy policy, and concurrency controls become recurring engineering work across many URLs. A framework is not required for a small, bounded collection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I use the same scraper for every website?

No. Rendering model, authentication, localization, pagination, and access rules vary by site. Keep a site-specific adapter behind a shared fetch, validation, retry, and persistence pipeline.

Should I store the complete HTML for every request?

Only when it is justified by debugging, reproducibility, or your retention policy. Otherwise store the extracted record, provenance, and a minimal diagnostic sample while avoiding unnecessary personal data.

When should I move from a script to a crawler framework?

Move when queue management, resumability, retries, proxy policy, and concurrency controls become recurring engineering work across many URLs. A framework is not required for a small, bounded collection.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.