Free tools Windows power users keep installed
One-click scans. No signup required.
To crawl JavaScript-heavy sites reliably, use Puppeteer behind a bounded queue and a per-origin scheduler: reuse a small browser pool, create Pages for jobs, isolate state with BrowserContexts when needed, resolve every intercepted request, and persist each result before acknowledging the queue item. There is no universal safe pages-per-browser or pages-per-second number. Measure a representative mix of your target pages, then raise concurrency only while latency, memory, error rate, and host limits remain acceptable.
What Puppeteer is—and when it is the right crawler
Puppeteer is a JavaScript library that automates Chrome and Firefox through the Chrome DevTools Protocol and WebDriver BiDi. It runs headless by default, so it can execute page JavaScript, wait for client-rendered content, click controls, and extract the DOM after hydration. That makes it useful for single-page applications and sites whose meaningful HTML does not exist in the initial HTTP response.
A browser is substantially heavier than an HTTP client. If a page is static, fetch its HTML with a normal HTTP client and parse it without launching Chromium. Reserve Puppeteer for pages that need JavaScript, browser cookies, rendering, interaction, or authenticated session state. A mixed crawler that chooses the lightest method per URL usually achieves better throughput and lower memory use than sending every request through a browser.
The architecture that scales
- Normalize and validate URLs. Canonicalize scheme, host, path, and the query parameters you actually need. Reject non-HTTP schemes and traps such as unbounded calendar, search, tracking, or session URLs. Keep a depth limit and a per-origin URL budget.
- Fetch robots.txt before navigation. Retrieve the file for each origin, select the matching user-agent group, cache the result, and test each URL against its rules. If the file cannot be fetched, treat the origin as disallowed rather than assuming permission. RFC 9309 says successfully fetched, parseable rules must be followed; it also says robots.txt is not an access-authorization mechanism.
- Put work in a durable, bounded queue. Store the URL, origin, depth, attempt count, next-eligible time, and result state. Keep queue admission separate from browser execution so retries cannot create an unbounded in-memory set. A durable queue lets a process restart without losing acknowledged work.
- Schedule by origin. Use a token bucket or equivalent host limiter, honor
Retry-After, and back off on 429 and 503 responses. Do not assume thatcrawl-delayis portable; Google documents that its robots.txt parser does not support it. - Run a controlled browser pool. Reuse browser processes where practical, create Pages as units of work, and recycle workers after measured memory growth or crashes. Use BrowserContexts when cookies or local storage must be isolated between tenants or accounts.
- Resolve every request event. Request interception can save bandwidth by aborting images, fonts, media, ads, or trackers. Once interception is enabled, every request must be continued, aborted, or answered; one unresolved event can stall a page.
- Extract and checkpoint. Record the requested URL, final URL, response status, redirect chain, title, selected content, discovered links, timings, and an error class. Persist the record before marking the queue item complete.
- Close deterministically. Close each Page in a
finallyblock, close a BrowserContext when its isolation scope ends, and close or recycle browsers during controlled shutdown.
One browser, contexts, or several processes?
| Design | Isolation | Startup cost | Failure blast radius | Best use |
|---|---|---|---|---|
| One browser, several Pages | Lowest; state must be managed carefully | Low after launch | A browser crash affects every Page | Homogeneous, trusted jobs |
| One browser, multiple BrowserContexts | Cookies and local storage are isolated per context | Moderate | A browser crash still affects all contexts | Multi-tenant or stateful jobs |
| Several browser processes | Strongest process boundary | Highest | Usually limited to one worker | Untrusted pages, memory-heavy jobs, or strict fault isolation |
These are engineering trade-offs, not throughput guarantees. The Puppeteer documentation does not publish a universal pages-per-browser benchmark or a memory-per-page number. Start with the smallest pool that meets your latency target and increase it only after measurement.
#1 Best Overall
Install and pin the browser toolchain
For a package-managed browser, install Puppeteer:
npm init -y
npm i puppeteer
The puppeteer package downloads a compatible Chrome. Use puppeteer-core when your deployment manages the browser separately and supplies an executable path. If a package manager blocks install scripts, run npx puppeteer browsers install or explicitly allow the package script. Pin both Puppeteer and the browser version in deployment, record the resolved versions with each crawl, and run a smoke crawl after upgrades because browser behavior and selectors can change.
A runnable bounded crawler in Node.js
The following example demonstrates URL normalization, a cached robots policy, a per-origin limiter, a bounded worker pool, request interception, retries, link discovery, and deterministic cleanup. It is intentionally conservative. Replace the in-memory queue with a durable queue and replace the simple robots parser with a standards-compliant implementation before operating at large scale.
import puppeteer from 'puppeteer';
import { URL } from 'node:url';
const seeds = process.argv.slice(2);
const MAX_PAGES = Number(process.env.MAX_PAGES || 100);
const MAX_DEPTH = Number(process.env.MAX_DEPTH || 2);
const WORKERS = Number(process.env.WORKERS || 3);
const USER_AGENT = 'ItechGuidesBot/1.0 (+https://example.com/crawler-policy)';
const robotsCache = new Map();
const lastRequest = new Map();
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
function normalize(raw, base) {
try {
const u = new URL(raw, base);
if (!['http:', 'https:'].includes(u.protocol)) return null;
u.hash = '';
return u.href;
} catch { return null; }
}
function parseRobots(text, agent) {
const groups = [];
let current = null;
for (const raw of text.split(String.fromCharCode(10))) {
const line = raw.trim();
if (!line || line.startsWith('#')) continue;
const cut = line.indexOf(':');
if (cut < 0) continue;
const key = line.slice(0, cut).trim().toLowerCase();
const value = line.slice(cut + 1).trim();
if (key === 'user-agent') {
current = { agents: [value.toLowerCase()], rules: [] };
groups.push(current);
} else if ((key === 'allow' || key === 'disallow') && current) {
current.rules.push({ allow: key === 'allow', path: value });
}
}
const wanted = agent.toLowerCase();
const selected = groups.filter(g => g.agents.includes(wanted) || g.agents.includes('*'));
return path => {
const rules = selected.flatMap(g => g.rules).filter(r => r.path && path.startsWith(r.path));
if (!rules.length) return true;
rules.sort((a, b) => b.path.length - a.path.length);
return rules[0].allow;
};
}
async function robotsFor(origin) {
if (robotsCache.has(origin)) return robotsCache.get(origin);
const policyUrl = origin + '/robots.txt';
try {
const response = await fetch(policyUrl, { headers: { 'User-Agent': USER_AGENT } });
if (!response.ok) throw new Error('robots HTTP ' + response.status);
const policy = parseRobots(await response.text(), 'ItechGuidesBot');
robotsCache.set(origin, policy);
return policy;
} catch {
const deny = () => false;
robotsCache.set(origin, deny);
return deny;
}
}
async function pace(origin) {
const minimumGap = 500; // tune from host policy and measurements
const previous = lastRequest.get(origin) || 0;
const wait = minimumGap - (Date.now() - previous);
if (wait > 0) await sleep(wait);
lastRequest.set(origin, Date.now());
}
const queue = [];
const seen = new Set();
for (const seed of seeds) {
const url = normalize(seed);
if (url && !seen.has(url)) { seen.add(url); queue.push({ url, depth: 0, attempt: 0 }); }
}
const browser = await puppeteer.launch({ headless: true });
let completed = 0;
async function worker() {
const page = await browser.newPage();
await page.setUserAgent(USER_AGENT);
await page.setRequestInterception(true);
page.on('request', request => {
const type = request.resourceType();
const action = ['image', 'font', 'media'].includes(type) ? request.abort() : request.continue();
action.catch(() => {});
});
try {
while (completed < MAX_PAGES) {
const job = queue.shift();
if (!job) break;
const target = new URL(job.url);
const origin = target.origin;
const allowed = await robotsFor(origin);
if (!allowed(target.pathname)) { completed++; continue; }
await pace(origin);
let record;
for (let attempt = 0; attempt < 3; attempt++) {
try {
const started = Date.now();
const response = await page.goto(job.url, { waitUntil: 'domcontentloaded', timeout: 30000 });
const status = response ? response.status() : null;
if (status && status >= 500) throw new Error('HTTP ' + status);
const data = await page.evaluate(() => ({
title: document.title,
text: document.body ? document.body.innerText.slice(0, 200000) : '',
links: Array.from(document.querySelectorAll('a[href]')).map(a => a.href)
}));
record = { url: job.url, finalUrl: page.url(), status, depth: job.depth, ms: Date.now() - started, ...data };
break;
} catch (error) {
record = { url: job.url, depth: job.depth, error: String(error), attempt: attempt + 1 };
if (attempt < 2) await sleep(500 * (2 ** attempt) + Math.random() * 250);
}
}
console.log(JSON.stringify(record));
if (record.links && job.depth < MAX_DEPTH) {
for (const link of record.links) {
const next = normalize(link, record.finalUrl || job.url);
if (next && !seen.has(next) && seen.size < MAX_PAGES * 4) {
seen.add(next); queue.push({ url: next, depth: job.depth + 1, attempt: 0 });
}
}
}
completed++;
}
} finally {
await page.close();
}
}
try {
await Promise.all(Array.from({ length: WORKERS }, worker));
} finally {
await browser.close();
}
Run it with node crawler.js https://example.com. Set WORKERS, MAX_PAGES, and MAX_DEPTH through environment variables. The sample uses one browser and several Pages. For independent sessions, create a BrowserContext per job or tenant, then create the Page inside that context and close the context when finished.
How to choose concurrency without guessing
Build a benchmark harness from representative URLs: static pages, JavaScript-heavy pages, slow pages, redirects, authenticated pages if permitted, and pages with large media. For each run, vary worker count in small increments and capture median and tail navigation time, success and timeout rates, browser RSS, open Pages, CPU, network bytes, and per-origin request rate. Stop increasing concurrency when memory grows continuously, tail latency jumps, failures rise, or the target host’s rate policy is reached. Repeat after changing Chromium, Puppeteer, network location, or page mix. A number that works for one site is not a portable capacity claim.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Recycle a browser when measured memory remains above your operational ceiling, after repeated page crashes, or after a controlled number of jobs. Recycling is a recovery mechanism, not a substitute for finding leaked Pages, event listeners, or unbounded crawl state.
Robots.txt, identity, and host etiquette
RFC 9309 defines robots.txt rules at /robots.txt. A successful fetch means parseable rules must be followed, but those rules do not authorize access to private or authenticated resources. Keep a descriptive User-Agent, publish a contact or policy page where appropriate, and enforce limits in your scheduler rather than relying on a site to provide a delay value. Cache robots policies for the period your implementation supports, refresh them when they expire, and log the policy decision for every skipped URL.
Rank #3
Do not blindly retry a 4xx response, an authentication failure, or a robots denial. Retry transient DNS/TLS failures, timeouts, and 5xx responses with bounded exponential backoff and jitter. Honor a server’s Retry-After value when present.
Reliability checklist
- Record navigation URL, final URL, status, redirect chain, elapsed time, response bytes where available, and exception class.
- Separate navigation timeout, selector timeout, DNS/TLS failure, HTTP error, blocked request, and extraction-schema failure so each retry policy is targeted.
- Set ceilings for depth, URLs per origin, response size, JavaScript wait time, and total job duration.
- Persist results before acknowledging queue items; make writes idempotent by URL and crawl version.
- Monitor queue age, open Pages, browser count, per-origin rate, success rate, timeout rate, memory, and crash/recycle count.
- Close Pages and contexts in
finallyblocks and drain workers during shutdown instead of killing browsers mid-navigation.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Pages hang after interception is enabled | A request event was neither continued, aborted, nor answered | Resolve every event, including error paths; abort only resource types you intentionally exclude. |
| Memory rises until Chromium crashes | Too many concurrent Pages, unclosed Pages/contexts, or unbounded queue state | Lower concurrency, close in finally, bound admission, and recycle workers at a measured threshold. |
| Many 429 or 503 responses | Per-origin rate is too high | Reduce the host’s token rate, honor Retry-After, and use exponential backoff with jitter. |
| Empty content on an application shell | Extraction ran before hydration | Wait for a stable selector or an application-specific readiness condition instead of adding an arbitrary long delay. |
| Robots decisions change between jobs | Policy was not cached or user-agent matching was inconsistent | Cache per origin, log the selected group, and send one consistent crawler identity. |
| Installation has no browser executable | Install scripts were blocked or puppeteer-core is being used without a managed browser |
Run npx puppeteer browsers install, allow the install script, or configure the executable path explicitly. |
| Selectors fail after an upgrade | Browser, site markup, or Puppeteer behavior changed | Pin versions, keep a smoke-crawl fixture set, and update selectors with versioned extraction tests. |
Or skip the browser setup
If your job is to obtain clean page screenshots rather than crawl links and extract data, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. The API accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options. A cURL call is:
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every feature is on every plan:
| Plan | Allowance | Price |
|---|---|---|
| Free | 1,000 shots/month | Free, no card |
| Starter | 3,000 shots | $5 |
| Growth | 15,000 shots | $15 |
| Pro | 60,000 shots | $39 |
| Scale | 250,000 shots | $99 |
| Business | 1,000,000 shots | $249 |
Yearly billing gives two months free. Start with 1,000 free screenshots a month with no card; paid plans start at $5 for 3,000.
FAQ
Can I use the crawler for authenticated pages?
Only when you have permission and an approved session design. Put credentials in a dedicated BrowserContext, never in a shared Page, and exclude private URLs from link discovery unless the owner has authorized that scope.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Should I keep the browser alive for an entire crawl?
Keep a process alive while measurements show stable memory and error rates. A controlled recycle after a bounded amount of work is safer than assuming a single process can run indefinitely.
Best Value
What should I change first when a crawl is slow?
Measure per-origin wait time, navigation time, and browser resource use separately. Then reduce unnecessary resources, correct host pacing, or add workers only where the measurements show spare capacity.
Frequently Asked Questions
Can I use the crawler for authenticated pages?
Only with permission and an approved session design. Put credentials in a dedicated BrowserContext, not a shared Page, and keep private URLs out of discovery unless that scope is authorized.
Should I keep the browser alive for an entire crawl?
Keep it alive while memory and error measurements remain stable, but recycle it after a bounded amount of work rather than assuming one process can run indefinitely.
What should I change first when a crawl is slow?
Measure host wait time, navigation time, and browser resource use separately; then tune pacing, resource blocking, or worker count according to the measured bottleneck.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

