Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

For a single page, save the HTML together with its required assets. For a bounded site, choose HTTrack when you want a browsable offline folder, or GNU Wget when you need a repeatable, logged command-line crawl. If the result must remain replayable and auditable, produce a WARC capture or WACZ package as well as (or instead of) a simple mirror. A successful download proves that bytes were retrieved; it does not prove that JavaScript-rendered states, authenticated content, media streams or every visual element was preserved.

Choose the archive you actually need

“Download a website” can mean several different outcomes. Decide the scope and fidelity before running a crawler, because the safest command for one page is not the safest command for an entire domain.

Need Best starting point What you get Main limitation
One article or reference page Browser save or a page-level capture HTML and selected assets Dynamic states and cross-page dependencies may be absent
A bounded, browsable copy HTTrack Local folders with rewritten links, HTML, images and other files Client-rendered, protected or interactive content may not be complete
A scripted, repeatable crawl GNU Wget Logged downloads with explicit host/path, depth and rate controls It does not execute a browser’s JavaScript state
Long-term replay and audit WARC or WACZ capture Containerized request/response evidence suitable for replay tools Requires more storage and archive tooling than a folder mirror

The Library of Congress describes the basic model as starting with a seed URL, following links, and downloading the content needed to preserve the site. Keep that model bounded: define the seed, allowed hosts and paths, recursion depth, file-size ceiling and request rate before you start.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan a defensible capture

Write down the scope

  • Record the exact seed URL, including its scheme, host, path and query string.
  • List allowed hosts and path prefixes. Decide whether a CDN, image host, download host or embedded video host is in scope.
  • Set a maximum recursion depth, maximum file size, total storage budget and delay between requests.
  • Note the capture time in UTC and the software version and command used.

Check access and permission

Read the site’s terms, access controls and any applicable copyright or database-rights rules. Ask for permission before copying private, restricted, commercially sensitive or redistribution-protected material. robots.txt is a crawler instruction, not a copyright licence. Wget documents robots-aware behavior, and HTTrack exposes its own robots option; neither changes your legal obligations. Use a descriptive user agent and throttle requests so the crawl does not impair the service.

Choose what “complete” means

For a visual reference, preserve HTML, CSS, fonts, images and downloadable documents. For evidentiary preservation, retain the HTTP responses, redirects, headers, timestamps, logs and checksums in WARC/WACZ form. For an application, document which user, locale, cookie state and interaction path were captured; an anonymous homepage is not the same artifact as a signed-in dashboard.

Save one page with its assets

  1. Open the canonical page in a browser and wait until visible images, fonts and deferred sections have loaded.
  2. Use the browser’s “Save page” command and choose the option that saves the complete page, not HTML only. Keep the generated HTML file and companion asset directory together.
  3. Save linked downloads separately when they are part of the record. Note that a page’s live URL, capture time and any visible account or locale state are metadata, not optional notes.
  4. Open the saved file with networking disabled. Check that images, styles and internal links work, then compare several sections with the live page.
  5. Hash the saved files and store the hash beside the capture. A later hash mismatch tells you the bytes changed; it does not tell you which visual state was missing.

Browser “complete page” saves are convenient but are not a substitute for a crawl or web-archive container. Pages that build their content after load, require a click, or fetch data through an API can appear intact while leaving important content out of the saved files.

Mirror a bounded site with HTTrack

HTTrack’s official description says it downloads a website recursively to a local directory, including HTML, images and other files, and rewrites links so the result can be browsed offline. Its documentation also covers resuming interrupted downloads and updating an existing mirror without fetching unchanged content.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example command

httrack "https://example.com/" 
  -O "./example-mirror" 
  --depth=3 
  --stay-on-same-domain 
  --robots=1 
  --max-rate=500000

Replace the domain and output directory. The depth limit prevents an unbounded walk; same-domain restriction prevents accidental expansion to unrelated hosts; the robots setting asks HTTrack to honor crawler instructions; and the rate limit reduces load. Check the exact option spelling in the HTTrack version installed on your system before production use.

Resume and update safely

Keep the output directory and logs. Re-running the same project lets HTTrack resume interrupted work and update an existing mirror. Do not delete the project metadata between runs if you need that behavior. Treat each update as a new capture: record its time, command and changed file list.

Create preservation containers

When your HTTrack build supports the documented archive options, add a WARC filename and rotation policy so very large captures are split into manageable files. The command guide also documents CDX indexes and WACZ packaging for replay tools. Keep the original mirror, WARC/WACZ package and logs together until validation is complete; a flattened PDF or folder alone cannot reproduce every web response.

Run a repeatable crawl with GNU Wget

GNU describes Wget as a free utility for non-interactive web downloads. Its manual documents robots-aware behavior, so combine that default with explicit boundaries and conservative timing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Conservative same-site crawl

wget 
  --mirror 
  --convert-links 
  --adjust-extension 
  --page-requisites 
  --no-parent 
  --wait=1 
  --random-wait 
  --user-agent="ArchiveBot/1.0 (contact: archivist@example.org)" 
  --domains=example.com 
  --directory-prefix=./example-mirror 
  --execute robots=on 
  --output-file=./example-wget.log 
  https://example.com/docs/

--mirror enables recursive mirroring behavior; --page-requisites requests assets needed by downloaded pages; --convert-links makes links usable locally; and --no-parent keeps a crawl seeded under /docs/ from ascending to the site root. The domain restriction is a second guard. Add --exclude-directories for known areas such as search results, calendars or user profiles. Do not use --span-hosts unless you have explicitly allowlisted every additional host.

Repeatability and logs

Save the complete command, Wget version, start and finish times, exit status and log file. Run a small sample first, inspect it, then expand the scope. A repeat run should use the same boundaries and a new destination or clearly labeled update directory so you can distinguish a changed page from a missing first capture.

Preserve WARC or WACZ for replay

A simple mirror is optimized for browsing files; it can omit response headers, redirects, duplicate resources and the context needed to audit a capture. The Digital Preservation Coalition notes that crawler collections are commonly stored in WARC containers and warns that simple mirrors and PDF output can flatten web content. WACZ packages a web-archive collection for tools that support that format.

Metadata to retain

  • Seed URL and every allowlisted host/path.
  • Capture start and finish times, time zone, crawler and browser versions.
  • Command-line options, configuration files, user-agent string and any authentication procedure.
  • WARC/WACZ filenames, CDX indexes, checksums and storage locations.
  • A plain-language statement of what was excluded: robots-denied paths, paywalls, streaming media, interactive states or private areas.

Validate replay

  1. Open representative pages from the archive with network access disabled.
  2. Check text, images, CSS, fonts, internal links, downloads, forms and media separately.
  3. Compare a few pages with the source while the capture is still fresh.
  4. Re-run a small sample after the crawl. Missing dependencies often appear only when a page is revisited.

Know what basic downloaders miss

JavaScript-rendered content

A crawler can download the initial HTML while the visible page is assembled later by JavaScript. Infinite scroll, client-side routing, consent dialogs, charts and data loaded through fetch/XHR calls may therefore be absent. Use a browser-based capture or specialist web-archiving crawler that can execute the required scripts, and document the exact interaction and resulting state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Authentication and paywalls

Do not put passwords, session cookies or bearer tokens in shell history, logs or URLs. Obtain authorization, use a dedicated account where appropriate, protect exported cookies, and state whether the archive represents an anonymous or authenticated view. A redirect to a login page is not evidence that the protected page was captured.

Bot challenges and blocked requests

CAPTCHAs, bot checks, rate limits and robots exclusions can stop a crawl or replace the intended page with an interstitial. Record the response and stop escalating request volume. A successful HTTP status only proves that a response arrived, not that it contains the content a human visitor saw.

Streaming audio and video

The UK Government Web Archive advises that streaming media should be available through progressive HTTP or HTTPS download with absolute source URLs, and that audio and video should have transcripts. If the source exposes only a transient stream, preserve the permitted downloadable representation and transcript, and note what could not be collected.

Interactive state

Search results, shopping carts, maps, logged-in dashboards and forms depend on state that a link crawler cannot infer. Capture important states deliberately with a browser workflow, screenshots or recordings, and retain the steps needed to reproduce them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, storage and reliability controls

  • Start small: test one directory or a few representative URLs before committing to a full crawl.
  • Throttle: use fixed or randomized delays, a clear user agent and a rate ceiling. Faster is not safer.
  • Bound storage: set maximum file sizes and exclude generated search, calendar and session URLs that can create near-infinite combinations.
  • Use checkpoints: preserve logs and partial output so an interruption can resume instead of restarting blindly.
  • Separate captures: use dated directories or immutable object-storage keys for each run.
  • Verify integrity: generate checksums, test extraction of WARC/WACZ files and periodically restore a sample to a separate machine.
  • Plan retention: store at least two copies in separate failure domains, protect sensitive captures with access controls and encryption, and document deletion dates.

There is no universal crawl-size or success-rate figure: storage and completeness depend on the site, asset sizes, duplication, access controls and chosen fidelity.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Fix
Only the homepage appears Depth, path or parent restriction is too narrow Check the seed path, depth and allowlist; test a linked page before widening scope.
Pages open but look unstyled CSS, fonts or absolute asset hosts were excluded Include page requisites and explicitly allow required asset hosts; inspect the log for denied URLs.
Images are missing Lazy loading or JavaScript fetches Use a browser-capable capture, or trigger the lazy-load state before capture.
Every URL becomes a login page Authentication was not preserved Use an authorized browser session and secure cookie handling; label the archive’s access state.
The crawl never finishes Calendar, search or session URLs generate an unbounded link graph Exclude those paths, cap depth and file count, and review URL patterns in logs.
Wget reports robots denial The site’s robots policy disallows the path Respect the instruction, seek permission, or use an approved archival service.
Media plays online but not offline It is a stream or uses relative/transient source URLs Preserve an allowed progressive download and transcript, or document the omission.
Local links still go online URL rewriting did not cover scripts or dynamically generated links Inspect the saved HTML and use a browser-based or replay-oriented archive for dynamic navigation.

Or skip the browser setup

For a clean screenshot or PDF of a specific URL, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

Use the API documentation at https://screenshotneo.com/docs/ for all options, including full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper size and page ranges, custom CSS and JavaScript, click-before-capture, hidden selectors, selector/delay/network-idle waits, blocked ads or resource types, custom headers/cookies/user agents, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and the OpenAPI specification.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to begin.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Archive checklist

  • Scope, permission and robots policy reviewed.
  • Seed URL, allowlist, depth, exclusions and rate limits recorded.
  • Capture time, tool versions, command and user agent saved.
  • HTML, CSS, scripts, images, fonts, documents and permitted media checked.
  • Authenticated or locale state documented.
  • Logs, checksums and WARC/WACZ or mirror files stored in separate copies.
  • Representative pages replayed offline and a sample rechecked after capture.
  • Known omissions and reasons written in the archive metadata.

Frequently Asked Questions

Is a PDF enough for a web archive?

Usually not. A PDF records a rendered document but can omit links, scripts, headers, alternate states and the original response context. Keep a WARC or WACZ capture when replay or audit matters.

Can I archive a site I do not own?

Only when your use is permitted by the site’s terms, applicable copyright and database-rights rules, and any required permission. Robots instructions guide crawlers but do not grant copying rights.

How do I prove which version I captured?

Record the exact URL, UTC timestamps, tool and command, response logs and cryptographic checksums, then preserve those records with the mirror or WARC/WACZ package.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.