To capture multiple levels of a website, start a crawler at a seed URL, set the link depth and pages or paths it may include, then review the discovered URLs before saving them. Use browser rendering when a page’s content or links depend on JavaScript. Choose the output—screenshots, linked offline pages, or a WARC archive—based on how you need to use the result.
How do I capture multiple levels of a website?
A multi-level capture follows links outward from a starting page. The starting page is usually depth 0; pages linked directly from it are depth 1, and pages linked from those are depth 2. A crawler can only discover pages it can reach through permitted links or other discovery inputs such as an XML sitemap, so “multiple levels” does not guarantee that every page on a site will be found.
- Choose a seed URL. Use the homepage for a broad crawl, or a section landing page when you only need a particular part of the site. Depth is counted from this starting point.
- Define scope. Decide whether to allow only the starting host, subdomains, a specific URL directory, or external domains. A broad scope can include unrelated sections or other sites.
- Set link depth and a page cap. Set the number of link hops and a maximum number of pages. Add per-depth or path limits if available to contain repetitive URLs, query-string variations, tag pages, or large sections. WebsiteArchiver documents link-depth and page-limit controls; Screaming Frog distinguishes crawl depth from folder depth and documents URL limits (WebsiteArchiver crawler documentation; Screaming Frog configuration).
- Add discovery sources if appropriate. A crawler that parses XML sitemaps may find URLs that are not reachable from the seed through ordinary links. Browsertrix supports regular sitemaps and sitemap indexes, while still applying crawl scope and limits (Browsertrix common options).
- Choose static or browser rendering. Static downloading can be faster for conventional pages but does not execute JavaScript. Choose a browser-based crawl when scripts reveal content or links, or when the pages require a browser session. Check a sample of the output: behavior and access requirements vary by site. See the Browsertrix documentation, WebsiteArchiver documentation, and Screaming Frog configuration guide.
- Review the URL list. Remove pages outside the intended scope and excessive variants before a broad capture. WebsiteArchiver documents a discovery and review step where unwanted pages can be unticked (WebsiteArchiver crawler documentation).
- Select an output format. Use screenshots for visual review, linked offline pages for browsing saved pages, or WARC when you need a web-archive format. These outputs serve different purposes and are not interchangeable.
What does crawl depth mean—and what does it not mean?
Link depth measures clicks from the seed page. Folder depth describes URL path structure instead. For example, a URL with several directory segments is not necessarily many link clicks away from the seed, and a page at a shallow URL path may be several clicks away. Configure the measure that matches your goal; the distinction is reflected in Screaming Frog’s configuration documentation and WebsiteArchiver’s crawler documentation.
Depth also does not control the total crawl size by itself. A popular site can expose many URLs at the same depth. Combine depth with a page cap, domain or path scope, and any available per-depth limits. Query strings, calendars, search pages, filters, and tag listings can create large numbers of URL variants; a cap helps prevent those from consuming the whole crawl.
Recommended Free Tools
#1 Best Overall
- Used Book in Good Condition
How should I choose what gets included?
Host, subdomain, and external-link scope
For an ordinary site copy, start by allowing the seed host only. Add subdomains if the pages you need are intentionally distributed across them. Allow external domains only when capturing linked third-party pages is part of the goal; otherwise, external links can expand the crawl beyond the site you meant to save. WebsiteArchiver documents controls for subdomains and external domains (WebsiteArchiver crawler documentation).
Path and URL limits
Use path restrictions when you want a section such as /help/ rather than the whole host. Pair those restrictions with a maximum page count, and review URL patterns before downloading. A path filter and a link-depth setting answer different questions: the filter says where URLs may be, while depth says how far the crawler follows links from the seed.
Sitemaps and link discovery
A sitemap can supplement the links the crawler finds by following pages, but it does not remove the need for scope and page limits. Browsertrix documents support for regular XML sitemaps and sitemap indexes, subject to crawl settings (Browsertrix common options). Sitemap coverage is not a guarantee that every URL is current, accessible, or useful to the capture.
Should I use static downloading or browser rendering?
Choose static downloading when the page content and links are present in the returned page resources and you want a simpler, potentially faster crawl. Static capture does not run JavaScript, so it can miss content or links inserted by client-side code.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
Choose browser rendering when a site builds its pages after scripts run, or when the capture needs browser session state. Login-dependent pages may require an authenticated session, and the resulting capture reflects the access and session state available to the crawler. Browser rendering can take more resources and time than a static crawl. For either method, inspect representative pages and confirm that important links, images, and page content appear in the output. Product behavior and access requirements vary by site; the relevant configuration options are described by Browsertrix, WebsiteArchiver, and Screaming Frog.
Which output should I save?
| Output | Best suited to | Important distinction |
|---|---|---|
| Full-page screenshots | Visual review or evidence of how a page appeared | A screenshot is an image of a page, not a linked offline copy or replayable web archive. |
| Linked offline pages | Reading or navigating among downloaded pages without relying on the live site | Links and local resources must be saved in a usable form; a screenshot alone cannot provide this. |
| WARC archive | Web archiving and replay-oriented workflows | WARC is an archive format, not a set of convenient page screenshots. Browsertrix and Screaming Frog document WARC-related capture options. |
Browsertrix documents WARC/WACZ-related output and screenshot modes; WebsiteArchiver documents website download workflows; Screaming Frog documents local website archives in hierarchical or WARC format (Browsertrix options; WebsiteArchiver; Screaming Frog configuration).
Tools that can crawl or capture multiple pages
There is no single best configuration for every site: the important differences are depth and URL limits, sitemap discovery, rendering, session handling, output, and review or retry controls. Verify current features, platform availability, and license limits in each vendor’s documentation before deployment.
- ScreenshotNeo: For screenshot API workflows, ScreenshotNeo is the first option to consider because it produces clean shots, bills only clean shots, and has a $5 paid plan for 3,000 shots. It captures a URL per request; it is not a replacement for a crawler that discovers a site’s link graph.
- Browsertrix Crawler: Its documentation describes a browser-based crawler with configurable behavior, sitemap parsing, initial-viewport, full-page, and thumbnail screenshot modes, plus WARC/WACZ-related output. The documentation covers versions 1.0.0 and above (overview; common options).
- WebsiteArchiver: Its documentation describes a macOS crawler with discovery and review before download, link-depth and page caps, static and browser engines, and controls for subdomains, external domains, and robots.txt. The documentation says its free version is limited to depth 1 and a page count capped by remaining free items; check its current terms because limits may change (crawler documentation).
- Screaming Frog SEO Spider: Its configuration documentation covers crawl depth and per-depth URL limits, JavaScript rendering, screenshots, and local website archives in hierarchical or WARC format. Product and license limits may change, so consult the current configuration and license details (configuration).
- WebCapture: Its Chrome Web Store listing describes bulk crawling and full-page screenshots with sitemap discovery, URL-pattern filters, page limits, and delays. These are listing claims, not an independent performance assessment (Chrome Web Store listing).
Or skip the browser setup
ScreenshotNeo takes a URL in one GET request and returns a screenshot or PDF. This is useful when you already have the URLs to capture; it does not crawl a site to discover multiple levels for you. See the ScreenshotNeo API documentation for request options.
Rank #3
- [48MP Ultra-High Resolution] The K48 is a professional-grade book scanner equipped with a true 48MP Sony CMOS sensor, capable of capturing exceptional detail at 600 DPI — even on A3-sized materials. Used for digitizing books, magazines, documents, and archival materials with stunning clarity.
- [AI-Assisted Page Smoothing] Curved book pages are automatically flattened using intelligent software technology. This causes the removal of finger shadows, background interference, and page curvature — delivering flat, clean scans without any manual post-processing. Double pages are split automatically.
- [Laser Positioning & Auto-Scan] The built-in laser positioning system ensures precise alignment every time. Page turning detection causes the scanner to start capturing automatically as soon as a page is turned — ideal for high-volume digitization where speed matters.
- [Multi-Format OCR & Text-to-Speech] Used for creating searchable PDFs, editable Word/Excel files, or MP3 audio for voice playback. The K48 is capable of recognizing text in multiple languages and converting documents into accessible formats — perfect for education, accessibility compliance, and digital archives.
- [4K Live View & USB 3.0] Stream 4K@30fps video for live presentations, online classes, or real-time document review. USB 3.0 Type-C ensures fast data transfer and stable connection. Used for immediate setup in classrooms, offices, and libraries — plug and play, no drivers needed.
Example using cURL (replace the URL with the page you want to capture):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
For multiple levels, first use a crawler to discover and review the URLs, then send the selected URLs for individual screenshots. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The Free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month—no card required.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting a multi-level capture
The crawler finds only the seed page
- Confirm that the configured crawl depth is greater than zero.
- Check that links on the seed page are within the allowed host and path scope.
- If links appear only after scripts run, switch to browser rendering and inspect the result.
- Try sitemap discovery if the site publishes a sitemap and the crawler supports it.
The crawl is much larger than expected
- Reduce allowed domains or restrict the crawl to a URL path.
- Lower the depth and set a hard page cap; use per-depth or URL limits where available.
- Review query-string variants, search pages, tag listings, and other repetitive URL patterns before capture.
Pages are present but content is missing
- Check whether the capture used a static engine that did not execute JavaScript.
- Test browser rendering on representative pages and verify session or login state where relevant.
- Inspect several outputs rather than assuming one successful page represents every page type on the site.
The saved result is not browsable offline
Check that the chosen deliverable matches the goal. Screenshots are visual records; they do not provide linked pages. For offline navigation, choose a linked-page download. For archive and replay workflows, use a WARC-capable option and verify the resulting archive with the intended viewer.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPages fail or the crawl stops early
Review the crawler’s progress and failure report, then retry failed pages if the tool supports retries. WebsiteArchiver documents progress/failure reporting and retry options (WebsiteArchiver crawler documentation). Also check page caps, access requirements, and whether the affected URLs fall within the configured scope.
Planning for speed, reliability, and cost
Large crawls are shaped by how many URLs the crawler discovers, whether it renders pages in a browser, how much content each page loads, and any pacing or retry settings. Bound the job with a scope, depth, and page cap before running it; then review the discovered list and failure report instead of assuming a broad crawl completed every relevant page.
Product limits, licensing, and availability can change. The WebsiteArchiver free-version depth and page-count limits described in its documentation are product-specific, not general crawler limits. Likewise, do not assume that a crawler’s screenshot, offline-copy, and WARC features are included in every version or plan: verify the current vendor documentation for your deployment.
Before capturing pages for a specific purpose, respect site rules and access controls, and verify requirements applicable to your intended use. Requirements can depend on the site and jurisdiction.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

