Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use an Apify Actor with Crawlee’s PuppeteerCrawler: accept a list of URLs, open each URL in a browser-backed request handler, call page.screenshot(), and save the returned bytes in the Actor’s default key-value store. Each screenshot gets a deterministic storage key so it can be retrieved after the run.

This guide builds that crawler, explains full-page and JPEG options, handles key collisions and failures, and shows when a snapshot containing HTML is more useful than an image alone.

What you are building

The Actor below processes one or more URL objects. For every request, Crawlee provides a browser page and an Apify request. The handler captures the rendered page and writes PNG bytes to key-value storage with Actor.setValue(). The run output remains available in the Actor’s default key-value store.

  • Input: an array such as [{"url":"https://example.com"}].
  • Browser: Puppeteer through PuppeteerCrawler.
  • Output: one image record per URL, plus run logs and statistics.
  • Scope: this is a supplied-URL crawler. Link discovery, domain restrictions and recursive crawling are separate design decisions.

Apify’s documented example describes this pattern as capturing a screenshot of a web page using Puppeteer. APIs, package versions and runtime images change, so check the current Apify SDK and Crawlee documentation when creating the Actor.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prerequisites and Actor layout

  • An Apify account and an Actor project, or a local Node.js project that can run an Apify Actor.
  • Node.js compatible with the SDK version you install.
  • The Apify SDK and Crawlee packages installed according to their current documentation.
  • URLs that your crawler is permitted to request. Respect site terms, robots policies and applicable law.

A minimal project normally contains src/main.js, an Actor input schema and (when using a custom build) a Dockerfile. Apify’s standalone browser example uses an Apify Node Puppeteer Chrome image; verify the image name against the SDK/runtime version you select rather than copying an old tag unchanged.

Define input for one or many URLs

Use an array of URL objects because it maps directly to Crawlee requests and leaves room for per-URL metadata later.

{
  "startUrls": [
    { "url": "https://example.com" },
    { "url": "https://www.example.org/docs" }
  ],
  "fullPage": true,
  "imageType": "png"
}

If you expose these fields in an Actor input schema, validate that every URL is absolute and that imageType is either png or jpeg. Keep a small test list first; increase scope only after inspecting stored records and run logs.

Complete PuppeteerCrawler implementation

The following implementation reads Actor input, creates one request per URL, captures each page in the request handler and saves the image. It also records failures in the dataset so a run can be audited without searching logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { Actor } from 'apify';
import { PuppeteerCrawler } from 'crawlee';

await Actor.init();

const input = (await Actor.getInput()) ?? {};
const startUrls = Array.isArray(input.startUrls) ? input.startUrls : [];
const fullPage = input.fullPage === true;
const imageType = input.imageType === 'jpeg' ? 'jpeg' : 'png';

if (startUrls.length === 0) {
  throw new Error('Provide at least one startUrls entry.');
}

function storageKey(url, index) {
  // Keep the URL visible while adding an index to prevent normalized-key collisions.
  const normalized = url
    .replace(/^https?:\/\//i, '')
    .replace(/[^a-zA-Z0-9._-]+/g, '_')
    .replace(/^_+|_+$/g, '')
    .slice(0, 180);
  return `SCREENSHOT_${String(index).padStart(6, '0')}_${normalized || 'page'}`;
}

const crawler = new PuppeteerCrawler({
  async requestHandler({ page, request, log }) {
    const index = Number(request.userData.index ?? 0);
    const key = storageKey(request.url, index);
    const screenshot = await page.screenshot({
      fullPage,
      type: imageType,
      ...(imageType === 'jpeg' ? { quality: 85 } : {})
    });

    await Actor.setValue(key, screenshot, {
      contentType: imageType === 'jpeg' ? 'image/jpeg' : 'image/png'
    });

    await Actor.pushData({
      url: request.url,
      key,
      contentType: imageType === 'jpeg' ? 'image/jpeg' : 'image/png',
      fullPage,
      capturedAt: new Date().toISOString()
    });
    log.info(`Saved ${request.url} as ${key}`);
  },

  async failedRequestHandler({ request, log }) {
    await Actor.pushData({
      url: request.url,
      error: request.errorMessages ?? ['Request failed']
    });
    log.error(`Failed ${request.url}`);
  }
});

await crawler.run(startUrls.map((entry, index) => ({
  url: typeof entry === 'string' ? entry : entry.url,
  userData: { index }
})));

await Actor.exit();

In JavaScript source, replace the HTML-escaped arrow in the displayed example with the normal => operator when copying from a page that escapes code. The actual source must contain => as JavaScript syntax (shown above as an HTML entity for valid article markup).

Rank #2
The Standards Real Book, C Version
  • Used Book in Good Condition

Why the key includes an index

Replacing URL punctuation with underscores is convenient and follows Apify’s multi-URL example, but two different URLs can normalize to the same string. The index prefix prevents that collision within a run. For resumable or cross-run pipelines, add a stable hash of the fully qualified URL and, if relevant, a version or crawl date. Check the current key-value-store character and length rules before choosing a production scheme.

Run it and retrieve images

  1. Paste the code into the Actor’s source and configure the input JSON.
  2. Run with two or three URLs.
  3. Open the run’s default key-value store. Records named SCREENSHOT_... contain the image; dataset records describe URL, key, type and timestamp.
  4. Open the run log and confirm that each URL produced a “Saved” message. Investigate every failed-request record before scaling up.

For local development, run the same Actor with the local Apify CLI or SDK command recommended for your installed version. Keep browser and SDK versions aligned; a container that worked with an older SDK 3.6 example may require a different runtime image today.

Choose viewport, full page and image format

Viewport versus full page

page.screenshot() captures the current rendered viewport by default. Set fullPage: true to request the complete scrollable document, as in the Academy example. Full-page images can be substantially larger, especially on long pages, and therefore consume more storage and transfer time. A viewport capture is preferable for visual regression of the above-the-fold layout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PNG versus JPEG

PNG is the documented default and preserves sharp text and transparency. JPEG generally produces smaller files but is lossy and does not preserve transparency. The example sets JPEG quality to 85; tune quality for your own visual threshold rather than treating that number as a universal recommendation.

Screenshot or snapshot?

A screenshot-only workflow is simplest when the deliverable is an image. Crawlee’s snapshot utility can save a screenshot and optionally the page HTML. Use a snapshot when debugging a discrepancy, preserving the DOM that produced an image, or reviewing page state after scripts run. HTML increases storage and may contain sensitive data; apply retention and access controls appropriate to your project.

Rank #3
Car Service Record Book Auto Repair Spiral Bound - 100 Pages/Book (Book 1)
  • 🚗 AUTOMOTIVE SERVICE-FOCUSED DESIGN: Tailored for automotive services, this Daily Car Service Record Book supports technicians and service writers in auto service shops, service truck operations, and dealership departments by organizing repair appointments, job authorizations, and maintenance tracking efficiently for professional workflow.
  • 🚗 COMPREHENSIVE LOGGING SOLUTION: With 50 sheets per book structured 8.5" × 11" size, this record book provides ample space to log customer information, auto service needs, and additional repair authorizations, making it ideal for managing detailed service jobs, tracking mileage, and maintaining vehicle maintenance records across automotive services.
  • 🚗 BUILT FOR SHOP ENVIRONMENTS: Constructed from high-quality paper and spiral-bound for durability, it withstands daily use in busy auto service bays and service truck operations. Pages are easy to flip, write on, or remove without tearing, providing a reliable solution for organized record-keeping.
  • 🚗 USER-FRIENDLY RECORD KEEPING: Designed for quick and easy use, this record book includes fields for customer names, phone numbers, technician assignments, repair notes, flat-rate hours, and mileage logs, ensuring professionals can track all service details accurately without missing important information.
  • 🚗 PROFESSIONAL AND VERSATILE: Whether scheduling jobs for a service truck, documenting auto service tasks in an independent shop, or maintaining dealership records, this car service record book functions as a daily planner, mileage log, and maintenance tracker, ensuring organized and professional workflow management for all automotive services.

Scaling from a test run

Concurrency, navigation timeout, retries, resource blocking and wait conditions are project-specific. Start conservatively and measure memory, run duration and failure rate. Increase concurrency only when the browser container remains stable. If pages render content after navigation, wait for a selector, a short delay or an application-specific readiness signal before calling screenshot(). Do not assume network idle means every visual element is complete.

For a recursive crawler, add explicit domain allowlists, URL canonicalization, depth limits and duplicate-request handling. A supplied URL list is safer and more predictable than discovering every link on a site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting

The Actor exits with no screenshots

Check that input uses startUrls, that it is an array, and that each entry has a valid absolute url. The sample throws an explicit error for an empty list.

Navigation succeeds but the image is blank

The page may require a later render step, a consent interaction or a wait for a selector. Add an application-specific wait before capture and inspect a saved snapshot. Also check redirects, authentication and bot challenges.

Some records overwrite others

Your key normalization is colliding. Include a stable hash, a run-specific index, or both; never rely on punctuation replacement alone for distinct URLs.

Full-page capture is huge or times out

Try viewport capture, reduce the URL set per run, wait for the page’s real ready state, and review browser memory. Very long pages may need a project-specific maximum-height policy.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JPEG output is rejected or unreadable

Ensure the screenshot option and storage contentType agree. Use PNG while diagnosing format issues, then re-enable JPEG after confirming the consumer supports it.

The browser fails to launch in deployment

Verify that the Actor image includes a compatible Chrome/Puppeteer runtime and that package versions match. Recheck the current Apify runtime guidance instead of depending on the SDK 3.6-era image recommendation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you only need screenshots rather than a custom crawler, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or a PDF, while options cover full-page capture, lazy images, CSS selectors, dark mode, device and retina settings, waits, custom headers and cookies, geolocation, blocking, resizing, caching, signed links, asynchronous jobs and bulk capture of up to 100 URLs per call.

Its cleaning step accepts consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example API calls (see the ScreenshotNeo documentation):

Best Value
Free Fling File Transfer Software for Windows [PC Download]
  • Intuitive interface of a conventional FTP client
  • Easy and Reliable FTP Site Maintenance.
  • FTP Automation and Synchronization
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.

Optional next steps

Once the Actor is reliable, add retention rules, a manifest dataset, image hashes for change detection, and notifications for failed requests. You can package the crawler as an Apify Store Actor, but publishing and monetization terms are time-sensitive; verify the current Apify Store and creator documentation before promising pricing or revenue.

Frequently Asked Questions

Can I use Playwright instead of Puppeteer?

Yes. Apify’s materials show that the Puppeteer example is nearly the same with Playwright. Choose the library whose APIs and runtime image match your project; the available evidence does not establish a universal speed or reliability winner.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where are the screenshots stored?

The sample writes image bytes to the Actor’s default key-value store and writes metadata to the default dataset. Open the completed run to retrieve both.

Does this crawler discover links automatically?

No. It processes the supplied URL list. Recursive discovery requires explicit domain, depth, canonicalization and duplicate-request rules.

Quick Recap

Bestseller No. 2
The Standards Real Book, C Version
The Standards Real Book, C Version
Used Book in Good Condition
$47.00
Bestseller No. 4
Bestseller No. 5
Free Fling File Transfer Software for Windows [PC Download]
Free Fling File Transfer Software for Windows [PC Download]
Intuitive interface of a conventional FTP client; Easy and Reliable FTP Site Maintenance.; FTP Automation and Synchronization

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.