Free tools Windows power users keep installed
One-click scans. No signup required.
For a server-rendered page, use TypeScript with a direct HTTP client and an HTML parser such as Cheerio. When the data appears only after JavaScript runs, an interaction is required, or browser state matters, use Playwright. In both cases, wait for a condition that proves the target data is ready—not merely for the browser’s load event—then validate, deduplicate, and persist the result with enough logging to diagnose failures.
Choose the smallest tool that can access the data
Start by looking at the HTML returned by an ordinary HTTP request. If the fields you need are already present, a browser adds operational cost without improving the result. If the response is only an application shell and JavaScript later fetches the records, a real browser is the appropriate escalation.
| Situation | Recommended approach | Reason |
|---|---|---|
| Server-rendered HTML and a small number of URLs | Built-in fetch or Axios plus Cheerio |
Low overhead and direct parsing of the returned document. |
| JavaScript-rendered content, clicks, scrolling, or browser state | Playwright | Runs a browser and exposes navigation, locators, and page events. |
| You need to understand redirects and failed resources | Playwright request events | Lifecycle events reveal what loaded, failed, or redirected. |
| Many URLs, retries, queues, or proxy controls | Crawlee or an equivalent crawler framework | Framework orchestration is safer to operate at crawl scale than a hand-written loop. |
Do not choose a browser simply because a page looks interactive. First confirm whether the required values are in the initial response.
Plan the scraper before writing selectors
Define the record
Write the output schema first. For example, a product record might contain name, price, currency, sourceUrl, and retrievedAt. Decide which fields are required and what an empty or malformed value means.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
Check access conditions
Review the site’s terms, any published API, authentication boundary, and /robots.txt before sending requests. RFC 9309 specifies that the rules must be available in a file named /robots.txt at the service’s top-level path. Treat the file as an important access signal, not as a universal legal permission. Also consider privacy, copyright, contractual restrictions, and rate limits.
Use a conservative request policy
Set a clear user agent, bound concurrency, and add backoff for temporary failures. Cache responses where the site’s rules permit it. Keep discovery, extraction, validation, and persistence separate so a selector change cannot silently corrupt stored data.
Scrape server-rendered HTML with TypeScript and Cheerio
Install a TypeScript runner and the parser:
npm install cheerio
npm install -D typescript tsx @types/node
The following program fetches article cards, checks the HTTP status, validates required fields, and writes JSON. Replace the URL and selectors with those for your target.
import * as cheerio from 'cheerio';
interface Article {
title: string;
href: string;
sourceUrl: string;
retrievedAt: string;
}
const targetUrl = 'https://example.com/news';
function absoluteUrl(value: string, base: string): string {
return new URL(value, base).href;
}
async function scrape(): Promise<Article[]> {
const response = await fetch(targetUrl, {
headers: { 'user-agent': 'ExampleResearchBot/1.0' },
signal: AbortSignal.timeout(30_000)
});
if (!response.ok) {
throw new Error(`HTTP ${response.status} for ${targetUrl}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const retrievedAt = new Date().toISOString();
const rows: Article[] = [];
$('article.card').each((_, element) => {
const title = $(element).find('h2').first().text().trim();
const href = $(element).find('a').first().attr('href');
if (!title || !href) return;
rows.push({
title,
href: absoluteUrl(href, targetUrl),
sourceUrl: targetUrl,
retrievedAt
});
});
if (rows.length === 0) {
throw new Error('No article cards matched; check for a layout change or JavaScript rendering');
}
return rows;
}
scrape()
.then(rows => console.log(JSON.stringify(rows, null, 2)))
.catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it with npx tsx scraper.ts. A zero-row result is an explicit failure, not a successful empty crawl. In production, persist the URL, retrieval time, parser version, and selector version with each record.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Scrape JavaScript-rendered pages with Playwright
Install Playwright and its browser binaries according to the installation instructions for your environment. This example uses locator-based extraction, a page-specific readiness condition, typed callback parameters, and network diagnostics.
npm install playwright
npm install -D typescript tsx @types/node
import { chromium, type Page } from 'playwright';
interface Product {
name: string;
price: string;
sourceUrl: string;
retrievedAt: string;
}
const targetUrl = 'https://example.com/catalog';
async function scrape(page: Page): Promise<Product[]> {
page.on('request', request => {
console.log('request', request.method(), request.url());
});
page.on('response', response => {
if (response.status() >= 400) {
console.warn('HTTP error', response.status(), response.url());
}
});
page.on('requestfinished', request => {
console.log('finished', request.url());
});
page.on('requestfailed', request => {
console.warn('failed', request.url(), request.failure()?.errorText);
});
await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('[data-testid="product-card"]').first().waitFor({
state: 'visible',
timeout: 30_000
});
const retrievedAt = new Date().toISOString();
const products = await page.locator('[data-testid="product-card"]').evaluateAll(
(elements): Product[] => elements.map(element => {
const name = element.querySelector('[data-testid="name"]')?.textContent?.trim() ?? '';
const price = element.querySelector('[data-testid="price"]')?.textContent?.trim() ?? '';
return {
name,
price,
sourceUrl: location.href,
retrievedAt: new Date().toISOString()
};
}).filter(product => product.name !== '' && product.price !== '')
);
if (products.length === 0) throw new Error('No valid products extracted');
return products;
}
const browser = await chromium.launch();
try {
const page = await browser.newPage();
const products = await scrape(page);
console.log(JSON.stringify(products, null, 2));
} finally {
await browser.close();
}
Use stable attributes such as data-testid when the site provides them. Avoid selectors coupled to presentation-only class names. Playwright supports locator APIs and TypeScript annotations in element callbacks; use those types to make schema changes visible during compilation.
Wait for readiness, not just page load
Navigation has several milestones. domcontentloaded means the initial document was parsed; load means the page’s load event fired. Neither guarantees that an application has finished fetching and rendering its data. Modern pages can continue network activity after both events.
Prefer a page-specific condition
- Wait for a known result locator to become visible.
- Wait for a response whose URL and status identify the data request.
- Wait for a short, justified delay only when the application offers no observable condition.
- For pages that settle predictably, use a network-idle condition cautiously; analytics or long polling can prevent it from completing.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForResponse(response =>
response.url().includes('/api/products') && response.ok(),
{ timeout: 30_000 }
);
await page.locator('[data-testid="product-card"]').first().waitFor();
A response event alone is not proof of success: a 404 or 503 can complete at the HTTP layer. Check response.ok() or the status code, and treat redirects as data worth inspecting rather than silently following.
Recommended Free Tools
Make selectors and extraction resilient
- Keep selectors narrow enough to identify one field, but not tied to incidental nesting.
- Test against representative variants: pagination, missing images, sold-out items, localization, and logged-out versus logged-in views.
- Normalize whitespace and currency deliberately; do not parse a localized price with assumptions about decimal separators.
- Validate required fields and reject or quarantine records that fail validation.
- Deduplicate using a stable key such as a canonical URL or site identifier.
- Version selectors and parsers so a later deployment can be traced to a change in output.
Advanced users can register a custom Playwright selector engine, but content-script isolation is safer when page JavaScript could interfere. A custom engine is an optimization for a known need, not a default starting point.
Observe the network while developing
Subscribe to request, response, requestfinished, and requestfailed events during development. These reveal redirect chains, blocked resources, failed API calls, and requests that never finish. Playwright exposes redirectedFrom() and redirectedTo() for tracing a chain.
Rank #3
Log the request URL, method, status, elapsed time, and retry count, but avoid recording unnecessary personal data or credentials. Keep verbose event logging behind a debug setting once the scraper is stable.
Production reliability for TypeScript scrapers
Bound concurrency and retries
Use a queue with a small concurrency limit. Retry transient network errors and selected 5xx responses with exponential backoff and jitter. Do not retry deterministic 4xx responses indefinitely. Give every navigation and extraction operation a timeout and a maximum retry count.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsSeparate stages
- Discover URLs and record the discovery source.
- Fetch or navigate with an explicit policy.
- Extract into an intermediate object.
- Validate, normalize, and deduplicate.
- Persist atomically and checkpoint progress.
Scale deliberately
For sustained crawls, evaluate Crawlee or an equivalent framework for queues, retries, and proxy controls. Confirm the current package behavior and any commercial terms before adopting it. A framework does not remove the need to respect site policies or control load.
Cache and resume
Cache immutable responses where permitted and store retrieval timestamps. Persist checkpoints so a process restart resumes instead of repeating the entire crawl. Keep raw responses only as long as your privacy and retention policy allows.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML contains no expected records | The data is rendered by JavaScript or the selector changed. | Inspect the raw response; switch to Playwright if the data arrives later, otherwise update and test the selector. |
| Playwright returns an empty list | Extraction ran before the target component appeared. | Wait for a specific locator or successful API response, then validate the count. |
| Navigation times out | Slow resources, a blocked request, a redirect loop, or an unreachable host. | Inspect request events, verify redirects and DNS, set a bounded timeout, and retry only transient failures. |
| A page looks loaded but fields are blank | The application replaced placeholder nodes after the load event. | Wait for visible, non-empty fields or the data response rather than for load. |
| HTTP 404 or 503 appears as a completed request | HTTP completion is not the same as a successful status. | Check the status explicitly and route the URL to retry or failure handling. |
| Many duplicate records | Pagination, retries, or multiple discovery paths overlap. | Canonicalize URLs and deduplicate before persistence. |
| Selectors break after a redesign | Selectors depended on visual classes or deep DOM structure. | Prefer stable attributes, keep selector versions, and run fixture-based tests. |
Legal and responsible scraping
Public visibility is not a blanket license to collect or reuse data. Check terms of service, authentication boundaries, privacy obligations, copyright constraints, and applicable law for your jurisdiction and purpose. Read /robots.txt and follow its applicable rules as an access signal. It controls crawler access; it does not itself remove a URL from search results. Site owners seeking exclusion should use mechanisms such as authentication or noindex, not robots.txt alone.
Identify yourself accurately, send only the requests you need, honor rate limits, and stop when the site indicates that access is not permitted.
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than structured records, ScreenshotNeo provides a single screenshot API request. It accepts cookie or consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each cleanup step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed.
Use the API documentation at https://screenshotneo.com/docs/ for the complete parameter list. A cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports PNG, JPEG, WebP, and PDF output; full-page capture with lazy images loaded; CSS-selector element capture; dark mode; 12 device presets and custom viewports; retina scale; PDF paper size, margins, landscape, and page ranges; HTML/CSS-to-image; custom CSS and JavaScript; pre-capture clicks; hidden selectors; waits for a selector, delay, or network idle; ad, tracker, request, and resource blocking; custom headers, cookies, user agents, and Authorization; timezone and geolocation; transparent backgrounds; resizing; chosen cache TTLs; signed links for public <img> tags; asynchronous jobs with signed webhooks; bulk capture of up to 100 URLs per call; a usage API; an OpenAPI specification; and compatible parameter names used by other screenshot APIs.
It also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Every feature is available on every plan.
| Plan | Allowance | Price |
|---|---|---|
| Free | 1,000 shots/month | No card required |
| Starter | 3,000 shots | $5 |
| Growth | 15,000 shots | $15 |
| Pro | 60,000 shots | $39 |
| Scale | 250,000 shots | $99 |
| Business | 1,000,000 shots | $249 |
Yearly billing provides two months free. Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.
Best Value
FAQ
Can I use the same scraper for every website?
No. Rendering model, authentication, localization, pagination, and access rules vary by site. Keep a site-specific adapter behind a shared fetch, validation, retry, and persistence pipeline.
Should I store the complete HTML for every request?
Only when it is justified by debugging, reproducibility, or your retention policy. Otherwise store the extracted record, provenance, and a minimal diagnostic sample while avoiding unnecessary personal data.
When should I move from a script to a crawler framework?
Move when queue management, resumability, retries, proxy policy, and concurrency controls become recurring engineering work across many URLs. A framework is not required for a small, bounded collection.
Frequently Asked Questions
Can I use the same scraper for every website?
No. Rendering model, authentication, localization, pagination, and access rules vary by site. Keep a site-specific adapter behind a shared fetch, validation, retry, and persistence pipeline.
Should I store the complete HTML for every request?
Only when it is justified by debugging, reproducibility, or your retention policy. Otherwise store the extracted record, provenance, and a minimal diagnostic sample while avoiding unnecessary personal data.
When should I move from a script to a crawler framework?
Move when queue management, resumability, retries, proxy policy, and concurrency controls become recurring engineering work across many URLs. A framework is not required for a small, bounded collection.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →

