The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use a CSV parser to read records, validate each row, then use Puppeteer to navigate to the row’s URL and save a screenshot. The example below uses CSV Parse’s synchronous API for a small file; it processes rows sequentially, reports individual failures, and closes the browser even if the batch encounters an error. For large files, use a streaming parser or async iterator instead of loading every record into memory.
What the workflow does
Each CSV record supplies the inputs for one capture—usually a URL and an identifier for the output filename. The script parses the file with a CSV-aware library, checks those fields, opens the URL in a browser page, waits for an appropriate ready state, and writes an image to disk.
Do not split CSV text on commas yourself: quoted values can contain commas, quotes, or line breaks. CSV Parse supports delimiters, quotes, escape characters, and comments, and its documented APIs include synchronous, streaming, callback, and async-iterator approaches (CSV Parse usage; CSV Parse API).
Prepare a CSV file
For the example, save a file named pages.csv in the project directory:
#1 Best Overall
id,url
home,https://example.com/
pricing,https://example.com/pricing
The script expects headers named id and url. Keep identifiers unique if you want one distinct image per row; otherwise, a later row with the same identifier could overwrite an earlier image. The output name is sanitized before it becomes a path component.
Install the packages
Use a Node.js project and install Puppeteer and CSV Parse:
npm init -y
npm install puppeteer csv-parse
Package and API compatibility depends on the versions installed in your environment. Puppeteer’s browser setup and requirements can vary by platform; use the installed package’s documentation if installation reports a missing system dependency. The code below uses the current documented Puppeteer pattern of launching a browser, opening a page, navigating, taking a screenshot, and closing the browser (Puppeteer Page API).
Rank #2
Run a complete sequential capture script
Create capture-csv.js beside pages.csv. This small-file version uses synchronous parsing, validates the required columns before launching the browser, and records row-level errors so one failed destination does not prevent later rows from being attempted.
const fs = require('node:fs');
const path = require('node:path');
const { parse } = require('csv-parse/sync');
const puppeteer = require('puppeteer');
const inputPath = path.resolve('pages.csv');
const outputDir = path.resolve('screenshots');
function safeFilename(value) {
const cleaned = String(value)
.trim()
.replace(/[^a-zA-Z0-9_-]+/g, '-').replace(/^-+|-+$/g, '');
return cleaned || 'row';
}
async function main() {
const csvText = fs.readFileSync(inputPath, 'utf8');
const rows = parse(csvText, {
columns: true,
skip_empty_lines: true,
bom: true,
trim: true
});
if (rows.some(row => !Object.hasOwn(row, 'id') || !Object.hasOwn(row, 'url'))) {
throw new Error('CSV must have id and url columns in its header.');
}
fs.mkdirSync(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const failures = [];
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
for (const [index, row] of rows.entries()) {
const label = String(row.id || `row-${index + 1}`);
const url = String(row.url || '').trim();
if (!url) {
failures.push({ row: index + 1, id: label, error: 'Missing URL' });
console.error(`Row ${index + 1} (${label}): missing URL; skipped`);
continue;
}
try {
const target = new URL(url);
if (!['http:', 'https:'].includes(target.protocol)) {
throw new Error('URL must use http or https');
}
await page.goto(target.href, {
waitUntil: 'networkidle2',
timeout: 30000
});
const outputPath = path.join(
outputDir,
`${String(index + 1).padStart(4, '0')}-${safeFilename(label)}.png`
);
await page.screenshot({ path: outputPath, fullPage: true });
console.log(`Row ${index + 1} (${label}): saved ${outputPath}`);
} catch (error) {
failures.push({ row: index + 1, id: label, error: error.message });
console.error(`Row ${index + 1} (${label}): ${error.message}`);
}
}
} finally {
await browser.close();
}
if (failures.length) {
fs.writeFileSync(
path.join(outputDir, 'errors.json'),
JSON.stringify(failures, null, 2)
);
console.error(`${failures.length} row(s) failed; see screenshots/errors.json`);
process.exitCode = 1;
}
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
Start it with node capture-csv.js. On success, PNG files appear in screenshots/. If one or more rows fail, the script continues with the next row, writes the failures to screenshots/errors.json, and exits with a nonzero status after cleanup.
Choose the right parsing mode
The synchronous example reads the whole CSV text and materializes all parsed records in memory. It is a straightforward fit when the file is small enough for that approach and the script benefits from checking its headers and rows before browser work begins.
Rank #3
- Used Book in Good Condition
| Parsing approach | Good fit | Trade-off |
|---|---|---|
| Synchronous | Small files and simple, sequential scripts | Holds the input and parsed records in memory; parsing finishes before processing starts |
| Streaming or async iteration | Large files or workflows that should process records incrementally | Requires a streaming control flow and deliberate handling of parser and row errors |
| Callback | Applications structured around callback-based processing | Control flow differs from the synchronous example |
CSV Parse documents all four API styles; select one according to dataset size, memory limits, and whether all records must be available before processing (CSV Parse API). A streaming parser does not by itself make screenshot work safe to run without limits: browser pages and destination sites still consume resources.
Set screenshot readiness and capture scope
Wait for the state the image needs
networkidle2 is a useful example readiness signal for pages that settle after network activity. Puppeteer’s screenshot guide demonstrates navigation with that wait condition, but it is not a guarantee that every site’s application content or images are ready (Puppeteer Screenshots). Some pages keep connections open, defer images, or render content after network activity quiets down.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteIf a specific component matters, wait for its selector before capturing. For application-driven content, wait for the element or state that indicates the data has rendered rather than assuming navigation completion means the page is visually complete. Use a timeout so a missing selector does not stall the entire batch indefinitely.
Rank #4
Choose full-page, viewport, or element output
The example uses fullPage: true, which asks Puppeteer to capture the whole page. Remove that option for a viewport image. If only one rendered component is needed, Puppeteer supports ElementHandle.screenshot(); this avoids capturing unrelated page areas when the target is a chart, card, or other element (Puppeteer Screenshots).
Make dimensions reproducible
Set the viewport explicitly before navigation or capture so screenshots use predictable dimensions. The sample uses a 1365 × 900 CSS-pixel viewport at scale factor 1. Different viewport dimensions, device scale factors, fonts, browser versions, and page content can change the resulting image, so record or standardize those conditions if you compare captures over time.
Improve throughput without losing control
Sequential processing is easier to debug and limits simultaneous browser work; the sample reuses one page and visits one row at a time. For more throughput, use bounded concurrency only after checking the capacity of the machine and the tolerance of the target sites. Unbounded parallel pages can exhaust memory or generate an unwanted request burst.
Best Value
- Reuse a browser across rows rather than launching one for every record, and close it in a
finallyblock. - Choose a finite navigation timeout and keep row-level errors with the row number and identifier.
- For resumable work, persist completed identifiers or results so a rerun does not need to repeat successful captures.
- Use a streaming CSV API for large inputs, then feed records into a bounded worker queue rather than creating a page for every record at once.
- Be mindful of access rules and traffic expectations for the sites you capture.
Troubleshoot common failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Parser reports missing columns or produces empty values | Header names do not match the script, or the CSV has a byte-order mark or unexpected formatting | Check the header row for id and url; the sample enables BOM handling and trims values |
| A URL row is skipped | The URL field is blank | Fill in the URL; the script records a row error and moves on |
| URL validation fails | The value is malformed or uses a scheme other than HTTP or HTTPS | Use a complete https:// or http:// URL |
| Navigation times out | The destination is slow, unreachable, or never reaches the selected network-idle condition | Check the URL and connectivity; choose a readiness condition suited to that page, and adjust the finite timeout if justified |
| Screenshot is blank or missing dynamic content | The page navigated but the needed content had not rendered, or the page did not load successfully | Wait for the relevant selector or application state; inspect the destination and the recorded row error |
| Browser launch fails | Browser installation or platform dependencies are unavailable | Review the Puppeteer installation output and the requirements for the installed version and operating system |
| Output files overwrite each other | Output naming is not unique | Keep the row index in the filename and use unique identifiers |
| Batch becomes slow or memory-heavy | Large files are held in memory, full-page captures are large, or too much browser work runs simultaneously | Use incremental parsing, sequential or bounded processing, and capture only the required page area |
Or skip the browser setup
If your goal is to capture each URL rather than control a local browser, ScreenshotNeo offers a one-request screenshot API. It accepts a URL and returns an image or PDF. See the ScreenshotNeo documentation for request options. For a CSV workflow, keep the same parser and row-validation steps, then call the endpoint once for each valid URL and save the response bytes using a unique filename.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Adapt the URL to the current row and the output filename to its identifier. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides the take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. ScreenshotNeo may suit URL-by-URL capture when its managed capture behavior fits your needs, while Puppeteer remains the choice when you need direct browser control. Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Can I use a CSV column to fill a form instead of navigating to a URL?
Yes. Read the relevant values from the row, use Puppeteer to populate the page or form, and wait for the resulting content before taking the screenshot. Keep the per-row validation and error handling so a malformed record does not stop the batch.
Does a successful screenshot prove the page loaded correctly?
No. A screenshot only shows what was rendered at capture time. For quality checks, add validation for the expected page state or selector and treat missing content as a row-level failure.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

