The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
For a single page, save the HTML together with its required assets. For a bounded site, choose HTTrack when you want a browsable offline folder, or GNU Wget when you need a repeatable, logged command-line crawl. If the result must remain replayable and auditable, produce a WARC capture or WACZ package as well as (or instead of) a simple mirror. A successful download proves that bytes were retrieved; it does not prove that JavaScript-rendered states, authenticated content, media streams or every visual element was preserved.
Choose the archive you actually need
“Download a website” can mean several different outcomes. Decide the scope and fidelity before running a crawler, because the safest command for one page is not the safest command for an entire domain.
| Need | Best starting point | What you get | Main limitation |
|---|---|---|---|
| One article or reference page | Browser save or a page-level capture | HTML and selected assets | Dynamic states and cross-page dependencies may be absent |
| A bounded, browsable copy | HTTrack | Local folders with rewritten links, HTML, images and other files | Client-rendered, protected or interactive content may not be complete |
| A scripted, repeatable crawl | GNU Wget | Logged downloads with explicit host/path, depth and rate controls | It does not execute a browser’s JavaScript state |
| Long-term replay and audit | WARC or WACZ capture | Containerized request/response evidence suitable for replay tools | Requires more storage and archive tooling than a folder mirror |
The Library of Congress describes the basic model as starting with a seed URL, following links, and downloading the content needed to preserve the site. Keep that model bounded: define the seed, allowed hosts and paths, recursion depth, file-size ceiling and request rate before you start.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesPlan a defensible capture
Write down the scope
- Record the exact seed URL, including its scheme, host, path and query string.
- List allowed hosts and path prefixes. Decide whether a CDN, image host, download host or embedded video host is in scope.
- Set a maximum recursion depth, maximum file size, total storage budget and delay between requests.
- Note the capture time in UTC and the software version and command used.
Check access and permission
Read the site’s terms, access controls and any applicable copyright or database-rights rules. Ask for permission before copying private, restricted, commercially sensitive or redistribution-protected material. robots.txt is a crawler instruction, not a copyright licence. Wget documents robots-aware behavior, and HTTrack exposes its own robots option; neither changes your legal obligations. Use a descriptive user agent and throttle requests so the crawl does not impair the service.
#1 Best Overall
Choose what “complete” means
For a visual reference, preserve HTML, CSS, fonts, images and downloadable documents. For evidentiary preservation, retain the HTTP responses, redirects, headers, timestamps, logs and checksums in WARC/WACZ form. For an application, document which user, locale, cookie state and interaction path were captured; an anonymous homepage is not the same artifact as a signed-in dashboard.
Save one page with its assets
- Open the canonical page in a browser and wait until visible images, fonts and deferred sections have loaded.
- Use the browser’s “Save page” command and choose the option that saves the complete page, not HTML only. Keep the generated HTML file and companion asset directory together.
- Save linked downloads separately when they are part of the record. Note that a page’s live URL, capture time and any visible account or locale state are metadata, not optional notes.
- Open the saved file with networking disabled. Check that images, styles and internal links work, then compare several sections with the live page.
- Hash the saved files and store the hash beside the capture. A later hash mismatch tells you the bytes changed; it does not tell you which visual state was missing.
Browser “complete page” saves are convenient but are not a substitute for a crawl or web-archive container. Pages that build their content after load, require a click, or fetch data through an API can appear intact while leaving important content out of the saved files.
Mirror a bounded site with HTTrack
HTTrack’s official description says it downloads a website recursively to a local directory, including HTML, images and other files, and rewrites links so the result can be browsed offline. Its documentation also covers resuming interrupted downloads and updating an existing mirror without fetching unchanged content.
Free tools Windows power users keep installed
One-click scans. No signup required.
Example command
httrack "https://example.com/"
-O "./example-mirror"
--depth=3
--stay-on-same-domain
--robots=1
--max-rate=500000
Replace the domain and output directory. The depth limit prevents an unbounded walk; same-domain restriction prevents accidental expansion to unrelated hosts; the robots setting asks HTTrack to honor crawler instructions; and the rate limit reduces load. Check the exact option spelling in the HTTrack version installed on your system before production use.
Resume and update safely
Keep the output directory and logs. Re-running the same project lets HTTrack resume interrupted work and update an existing mirror. Do not delete the project metadata between runs if you need that behavior. Treat each update as a new capture: record its time, command and changed file list.
Rank #2
Create preservation containers
When your HTTrack build supports the documented archive options, add a WARC filename and rotation policy so very large captures are split into manageable files. The command guide also documents CDX indexes and WACZ packaging for replay tools. Keep the original mirror, WARC/WACZ package and logs together until validation is complete; a flattened PDF or folder alone cannot reproduce every web response.
Run a repeatable crawl with GNU Wget
GNU describes Wget as a free utility for non-interactive web downloads. Its manual documents robots-aware behavior, so combine that default with explicit boundaries and conservative timing.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchConservative same-site crawl
wget
--mirror
--convert-links
--adjust-extension
--page-requisites
--no-parent
--wait=1
--random-wait
--user-agent="ArchiveBot/1.0 (contact: archivist@example.org)"
--domains=example.com
--directory-prefix=./example-mirror
--execute robots=on
--output-file=./example-wget.log
https://example.com/docs/
--mirror enables recursive mirroring behavior; --page-requisites requests assets needed by downloaded pages; --convert-links makes links usable locally; and --no-parent keeps a crawl seeded under /docs/ from ascending to the site root. The domain restriction is a second guard. Add --exclude-directories for known areas such as search results, calendars or user profiles. Do not use --span-hosts unless you have explicitly allowlisted every additional host.
Repeatability and logs
Save the complete command, Wget version, start and finish times, exit status and log file. Run a small sample first, inspect it, then expand the scope. A repeat run should use the same boundaries and a new destination or clearly labeled update directory so you can distinguish a changed page from a missing first capture.
Preserve WARC or WACZ for replay
A simple mirror is optimized for browsing files; it can omit response headers, redirects, duplicate resources and the context needed to audit a capture. The Digital Preservation Coalition notes that crawler collections are commonly stored in WARC containers and warns that simple mirrors and PDF output can flatten web content. WACZ packages a web-archive collection for tools that support that format.
Rank #3
- Used Book in Good Condition
Metadata to retain
- Seed URL and every allowlisted host/path.
- Capture start and finish times, time zone, crawler and browser versions.
- Command-line options, configuration files, user-agent string and any authentication procedure.
- WARC/WACZ filenames, CDX indexes, checksums and storage locations.
- A plain-language statement of what was excluded: robots-denied paths, paywalls, streaming media, interactive states or private areas.
Validate replay
- Open representative pages from the archive with network access disabled.
- Check text, images, CSS, fonts, internal links, downloads, forms and media separately.
- Compare a few pages with the source while the capture is still fresh.
- Re-run a small sample after the crawl. Missing dependencies often appear only when a page is revisited.
Know what basic downloaders miss
JavaScript-rendered content
A crawler can download the initial HTML while the visible page is assembled later by JavaScript. Infinite scroll, client-side routing, consent dialogs, charts and data loaded through fetch/XHR calls may therefore be absent. Use a browser-based capture or specialist web-archiving crawler that can execute the required scripts, and document the exact interaction and resulting state.
Authentication and paywalls
Do not put passwords, session cookies or bearer tokens in shell history, logs or URLs. Obtain authorization, use a dedicated account where appropriate, protect exported cookies, and state whether the archive represents an anonymous or authenticated view. A redirect to a login page is not evidence that the protected page was captured.
Bot challenges and blocked requests
CAPTCHAs, bot checks, rate limits and robots exclusions can stop a crawl or replace the intended page with an interstitial. Record the response and stop escalating request volume. A successful HTTP status only proves that a response arrived, not that it contains the content a human visitor saw.
Streaming audio and video
The UK Government Web Archive advises that streaming media should be available through progressive HTTP or HTTPS download with absolute source URLs, and that audio and video should have transcripts. If the source exposes only a transient stream, preserve the permitted downloadable representation and transcript, and note what could not be collected.
Interactive state
Search results, shopping carts, maps, logged-in dashboards and forms depend on state that a link crawler cannot infer. Capture important states deliberately with a browser workflow, screenshots or recordings, and retain the steps needed to reproduce them.
Performance, storage and reliability controls
- Start small: test one directory or a few representative URLs before committing to a full crawl.
- Throttle: use fixed or randomized delays, a clear user agent and a rate ceiling. Faster is not safer.
- Bound storage: set maximum file sizes and exclude generated search, calendar and session URLs that can create near-infinite combinations.
- Use checkpoints: preserve logs and partial output so an interruption can resume instead of restarting blindly.
- Separate captures: use dated directories or immutable object-storage keys for each run.
- Verify integrity: generate checksums, test extraction of WARC/WACZ files and periodically restore a sample to a separate machine.
- Plan retention: store at least two copies in separate failure domains, protect sensitive captures with access controls and encryption, and document deletion dates.
There is no universal crawl-size or success-rate figure: storage and completeness depend on the site, asset sizes, duplication, access controls and chosen fidelity.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the homepage appears | Depth, path or parent restriction is too narrow | Check the seed path, depth and allowlist; test a linked page before widening scope. |
| Pages open but look unstyled | CSS, fonts or absolute asset hosts were excluded | Include page requisites and explicitly allow required asset hosts; inspect the log for denied URLs. |
| Images are missing | Lazy loading or JavaScript fetches | Use a browser-capable capture, or trigger the lazy-load state before capture. |
| Every URL becomes a login page | Authentication was not preserved | Use an authorized browser session and secure cookie handling; label the archive’s access state. |
| The crawl never finishes | Calendar, search or session URLs generate an unbounded link graph | Exclude those paths, cap depth and file count, and review URL patterns in logs. |
| Wget reports robots denial | The site’s robots policy disallows the path | Respect the instruction, seek permission, or use an approved archival service. |
| Media plays online but not offline | It is a stream or uses relative/transient source URLs | Preserve an allowed progressive download and transcript, or document the omission. |
| Local links still go online | URL rewriting did not cover scripts or dynamically generated links | Inspect the saved HTML and use a browser-based or replay-oriented archive for dynamic navigation. |
Or skip the browser setup
For a clean screenshot or PDF of a specific URL, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
Use the API documentation at https://screenshotneo.com/docs/ for all options, including full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper size and page ranges, custom CSS and JavaScript, click-before-capture, hidden selectors, selector/delay/network-idle waits, blocked ads or resource types, custom headers/cookies/user agents, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and the OpenAPI specification.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to begin.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Archive checklist
- Scope, permission and robots policy reviewed.
- Seed URL, allowlist, depth, exclusions and rate limits recorded.
- Capture time, tool versions, command and user agent saved.
- HTML, CSS, scripts, images, fonts, documents and permitted media checked.
- Authenticated or locale state documented.
- Logs, checksums and WARC/WACZ or mirror files stored in separate copies.
- Representative pages replayed offline and a sample rechecked after capture.
- Known omissions and reasons written in the archive metadata.
Frequently Asked Questions
Is a PDF enough for a web archive?
Usually not. A PDF records a rendered document but can omit links, scripts, headers, alternate states and the original response context. Keep a WARC or WACZ capture when replay or audit matters.
Can I archive a site I do not own?
Only when your use is permitted by the site’s terms, applicable copyright and database-rights rules, and any required permission. Robots instructions guide crawlers but do not grant copying rights.
How do I prove which version I captured?
Record the exact URL, UTC timestamps, tool and command, response logs and cryptographic checksums, then preserve those records with the mirror or WARC/WACZ package.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

