Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To find a page’s assets with Puppeteer, attach network listeners before navigation, then combine the captured requests with a DOM scan and any scrolling or interactions needed to trigger lazy-loaded content. The network log finds resources that actually began loading—including API calls and assets with no visible DOM element—while the DOM scan can find declared URLs that never loaded. Neither method alone guarantees a complete inventory of everything a user could reveal or trigger.

What “all page assets” means

A page asset can mean a URL declared in markup, a resource the browser requested, or the bytes returned by a request. Those are different inventories. For example, an image URL may appear in srcset but not be requested at the current viewport size; a script may make an API request that has no corresponding asset element; and a resource may be requested but fail to load.

For a useful audit, collect network activity and page declarations together. Record the URL, resource type, method, frame context, response status and headers, cache information, redirect relationship, and any failure text. Keep distinct requests distinct: a page can request the same URL more than once, and collapsing by URL alone can hide meaningful behavior.

  • Network log: requests that took place during the observation window, including non-DOM requests.
  • DOM and CSS references: URLs declared in elements and styles, whether or not they were fetched.
  • Downloaded bytes: response bodies, which require additional handling and are not guaranteed to be readable in every case.

Puppeteer’s HTTP request API describes request lifecycle events emitted when a page requests a network resource. This walkthrough uses those events for metadata collection; it does not assume that one navigation triggers every asset the site can load.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a network asset collector

Register listeners before calling page.goto(), or early requests may be missed. This runnable Node.js example launches Chromium, records request lifecycle details, visits a URL, waits for network idle, and writes a JSON report. Install Puppeteer first with npm install puppeteer, save the script as assets.js, then run node assets.js https://example.com.

const puppeteer = require('puppeteer');
const fs = require('node:fs/promises');

async function main() {
  const targetUrl = process.argv[2];
  if (!targetUrl) throw new Error('Usage: node assets.js <url>');

  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    const records = new Map();

    function getRecord(request) {
      if (!records.has(request)) {
        let frameUrl = null;
        try { frameUrl = request.frame()?.url() ?? null; } catch {}
        records.set(request, {
          url: request.url(),
          resourceType: request.resourceType(),
          method: request.method(),
          frameUrl,
          redirectChain: request.redirectChain().map(item => item.url()),
          started: true
        });
      }
      return records.get(request);
    }

    page.on('request', request => {
      getRecord(request);
    });

    page.on('response', response => {
      const request = response.request();
      Object.assign(getRecord(request), {
        status: response.status(),
        headers: response.headers(),
        fromCache: response.fromCache()
      });
    });

    page.on('requestfinished', request => {
      getRecord(request).finished = true;
    });

    page.on('requestfailed', request => {
      getRecord(request).failure = request.failure()?.errorText ?? 'Request failed; no error text reported';
    });

    const navigationResponse = await page.goto(targetUrl, {
      waitUntil: 'networkidle0',
      timeout: 60000
    });

    const networkAssets = [...records.values()];
    const domAssets = await collectDomReferences(page);
    const report = {
      requestedUrl: targetUrl,
      finalPageUrl: page.url(),
      navigationStatus: navigationResponse?.status() ?? null,
      capturedAt: new Date().toISOString(),
      networkAssets,
      domAssets
    };

    await fs.writeFile('assets.json', JSON.stringify(report, null, 2));
    console.log(`Saved ${networkAssets.length} network requests and ${domAssets.length} DOM references to assets.json`);
  } finally {
    await browser.close();
  }
}

async function collectDomReferences(page) {
  return page.evaluate(() => {
    const found = [];
    const add = (url, source, element) => {
      if (!url) return;
      try {
        found.push({
          url: new URL(url, document.baseURI).href,
          source,
          element: element?.tagName?.toLowerCase() ?? null
        });
      } catch {}
    };
    const addSrcset = (value, source, element) => {
      for (const candidate of (value || '').split(',')) {
        const url = candidate.trim().split(/\s+/)[0];
        add(url, source, element);
      }
    };
    const addCssUrls = (value, source, element) => {
      for (const match of (value || '').matchAll(/url\(\s*(['"]?)(.*?)\1\s*\)/gi)) {
        add(match[2], source, element);
      }
    };

    for (const element of document.querySelectorAll('*')) {
      for (const attr of ['src', 'href', 'poster']) {
        if (element.hasAttribute(attr)) add(element.getAttribute(attr), attr, element);
      }
      if (element.hasAttribute('srcset')) addSrcset(element.getAttribute('srcset'), 'srcset', element);
      if (element.hasAttribute('style')) addCssUrls(element.getAttribute('style'), 'inline-style', element);
    }

    for (const sheet of document.styleSheets) {
      try {
        for (const rule of sheet.cssRules) {
          addCssUrls(rule.cssText, 'stylesheet-rule', null);
        }
      } catch {
        found.push({ url: sheet.href, source: 'stylesheet-unreadable', element: 'link' });
      }
    }
    return found;
  });
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The Map uses the request object as the key, so repeated requests to one URL are not silently merged. Each redirect is represented by its own request; redirectChain preserves earlier URLs associated with that request. The JSON report therefore represents request instances, not a unique-URL list. If a downstream report needs unique URLs, derive that view separately while retaining the original records for diagnosis.

The DOM scan resolves relative references against document.baseURI and checks src, href, poster, srcset, inline styles, and readable stylesheet rules. A stylesheet’s CSS rules may be inaccessible, for example when browser security restrictions prevent reading them; the report then preserves the stylesheet URL as unreadable instead of claiming its contents were scanned. The simple srcset parser is useful for common values but is not a complete parser for every possible URL syntax.

Capture dynamically loaded assets

networkidle0 is a documented navigation wait option and is a reasonable starting point, not a universal definition of “finished.” Pages with polling, streaming connections, ads, or other recurring activity may not reach network idle. Conversely, a lazy image below the fold may not load until scrolling, and an API call may only happen after opening a menu or submitting a form.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Choose a meaningful wait condition. Use waitUntil: 'networkidle0' for a page expected to settle. If the application exposes a reliable ready element, wait for that selector instead of relying on an arbitrary delay.
  2. Trigger deferred work. Scroll through the page in increments, wait for images or content to appear, and click relevant controls. Keep the request listeners active during all of these actions.
  3. Repeat DOM collection after interaction. A scan immediately after initial navigation cannot include elements inserted later.
  4. Watch page context changes. Inspect frames as well as the main document. Where workers participate, account for worker activity separately rather than assuming a main-frame DOM scan reveals it.
  5. Set a stopping rule. For pages that keep changing, define the interaction sequence and observation window you need. “All assets” is only meaningful relative to that tested state and period.

Chrome for Developers describes Puppeteer as a JavaScript library for automating Chrome and Firefox through the Chrome DevTools Protocol and WebDriver BiDi. The event-based approach here is useful precisely because page markup alone does not describe all browser traffic.

When to use interception, and when not to

For inventory work, passive request and response listeners are usually the right tool. Enable page.setRequestInterception(true) only when you intend to alter traffic—for example, to abort, continue, or fulfill requests. Puppeteer warns that once interception is enabled, every request stalls until it is continued, responded to, aborted, or completed using the browser cache. A handler that fails to resolve one request can hang navigation.

Interception changes page behavior and can make the observed asset set differ from an ordinary load. If you enable it for a specific experiment, resolve every intercepted request deliberately and label the resulting report as an intercepted run.

URLs and metadata versus downloaded files

The example records request metadata; it does not save response bodies. If the goal is a local asset archive, read bodies while the corresponding responses remain available and write them under collision-safe names. A filename derived only from the last path segment is unsafe: distinct URLs can share that name, and query strings can distinguish different resources. Preserve the original URL and use a stable hash or equivalent collision-resistant suffix.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not assume every response body can always be read. Opaque, cross-origin, streaming, and service-worker responses need separate consideration, and the available evidence does not establish universal body access. Keep metadata collection independent from body downloading so a body-read limitation does not erase the request record.

How to interpret the report

  • Request with status 404 or 503: an HTTP response was received. These statuses do not automatically make the request a requestfailed event; inspect the recorded status.
  • Request with failure text and no status: the browser reports a failed request rather than a completed HTTP response. Use the failure text as a diagnostic clue, not as proof of one specific cause.
  • DOM reference absent from the network log: it may not have been selected, triggered, reachable, or loaded during the observed page state.
  • Network request absent from the DOM scan: it may be an API response, script-created resource, or other request with no matching element.
  • Same URL more than once: retain the separate request records. They can represent different methods, contexts, retries, or page actions.

Troubleshooting common collection problems

The first scripts or stylesheets are missing

Attach all event listeners before page.goto(). If listeners are attached after navigation, early requests have already passed. Also confirm that the script is observing the page instance on which the navigation occurs.

Navigation times out on a busy site

A page with continuous network activity may never satisfy networkidle0. Use an application-specific ready condition or a bounded wait strategy, and continue observing requests for the period your task requires. Do not treat a timeout as evidence that the site has no more assets.

Lazy images or content are missing

Scroll or interact to cause the page to request deferred resources, then scan the DOM again. A single initial page load cannot discover assets that require later user actions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A failed load looks like a missing HTTP status

Check the request’s failure field as well as response status. HTTP error responses such as 404 and 503 are still completed HTTP requests; they are not automatically reported as request failures.

Navigation hangs after enabling interception

Remove interception if you only need observation. If interception is necessary, make sure every request path—including exceptions and conditional branches—is continued, fulfilled, or aborted. An unresolved intercepted request stalls.

A stylesheet URL is found but its rules are not

The browser may deny access to the stylesheet’s rules. Preserve the stylesheet reference and mark its contents unreadable; do not interpret that as proof the stylesheet has no asset URLs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost considerations

Network logging is less disruptive than interception, but the collector still adds work: it stores metadata for each observed request, and body capture or extensive interaction adds more. For large pages, write results after collection and avoid retaining response bodies unless downloading them is required. Bound navigation and interaction waits so an indefinitely active page does not consume a run forever.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For repeatable results, record the target URL, final page URL, observation window, wait condition, interactions performed, and whether interception was enabled. The asset list depends on viewport, application state, and actions. A fresh run may also differ from a cached run; keeping cache metadata and request instances helps explain those differences. Do not treat a single run as a universal inventory of every possible page state.

Or skip the browser setup

If the task is to produce a screenshot rather than inspect every request, ScreenshotNeo provides a one-request screenshot API. Its clean-shot flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture.

For a screenshot, this cURL call saves the returned image as WebP. See the ScreenshotNeo documentation for API options. It is a screenshot service, not a replacement for Puppeteer’s request-by-request asset inventory.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month—no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.