You can make a useful offline copy of a website with a recursive crawler, but no crawler can promise a complete copy of everything a site serves. For a browsable mirror, GNU Wget is a practical command-line choice; HTTrack offers graphical and command-line workflows. If you need a preservation-oriented collection, consider retaining WARC output as well as—or instead of—a directory of downloaded pages. Start with a small, polite crawl, inspect its log, and check key pages and assets before treating the archive as complete.
What “download an entire website” means
A crawler begins at one or more URLs, follows links and retrieves resources within the scope and access rules you set. The result is a crawl snapshot: a record of what the tool could discover and retrieve at that time. It is not a copy of a site’s underlying database or a guarantee that every page and behavior has been captured.
Some material may be inaccessible to a crawler: content behind authentication, forms or interactions that must be triggered, data fetched dynamically from APIs, resources hosted outside your crawl scope, or pages blocked by access rules. A page may download while still missing an image, stylesheet, font or other linked asset. Record the crawl date, starting URLs, tool and version, scope, and known exclusions so future readers understand what the archive represents.
Choose an archiving approach
| Approach | Best for | What to expect |
|---|---|---|
| GNU Wget | A scoped crawl from the command line | Downloads files into a directory; options can retrieve page requisites and convert links for local browsing. |
| HTTrack | A browsable mirror with GUI or command-line controls | Downloads recursively, maintains a usable local link structure, and can resume or update a mirror. |
| ArchiveBox | A self-hosted collection in several capture formats | Its project describes captures including HTML, screenshots, PDF and WARC using tools such as Chrome and wget; results depend on the site and extractor. |
| WARC-oriented capture | Preservation and replay workflows | A WARC record serves a different purpose from an ordinary directory mirror. Some workflows retain both. |
For a simple offline website you can open as files, begin with Wget or HTTrack. If the archive must support preservation or replay workflows, check whether WARC—and, where relevant, CDXJ indexing or WACZ bundling—fits your requirements. HTTrack documents WARC output alongside its browsable mirror and notes that the mirror is not a substitute for the WARC record. See the HTTrack project and its command-line guide. ArchiveBox’s supported extractors and outputs should not be assumed to work identically on every site; see the ArchiveBox project.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Make a browsable mirror with GNU Wget
Install GNU Wget for your operating system, make sure the destination drive has adequate free space, and run this starting command in a shell. Replace the example URL with a site you are permitted to archive.
wget --mirror --convert-links --page-requisites --adjust-extension --wait=1 --directory-prefix=./site-archive https://example.com/
This command is a starting point, not a completeness guarantee. It asks Wget to retrieve recursively and save the result under ./site-archive.
What each option does
--mirrorenables recursive retrieval, timestamping and infinite depth. Without mirror mode, GNU Wget documents a default recursive depth of five.--page-requisitesretrieves files needed to render pages, such as images and stylesheets.--convert-linksrewrites links in downloaded documents for local viewing.--adjust-extensionhelps save HTML responses with an appropriate extension.--wait=1adds a one-second delay between retrievals. Use a slower crawl if the site is small or its operator requests it.--directory-prefix=./site-archiveplaces the downloaded files in the named local directory.
Wget follows links in HTML and CSS during recursive retrieval and honors robots rules, according to the GNU Wget manual. Set the starting URL and use appropriate directory or domain restrictions to keep the crawl in scope. Review the manual for the installed Wget version before relying on particular switches or behavior.
Run a small trial first
- Choose a representative starting page and a narrow scope, such as the relevant section of a site rather than every linked domain.
- Run the command with the destination directory on a drive with enough free space. Watch the output for failed requests, excluded paths and unexpected hosts.
- Open several downloaded pages locally. Check navigation, images, stylesheets and other important resources—not just the landing page.
- Adjust scope or settings if needed, then run the crawl at a rate that respects the site and its operator’s rules.
- Keep the log and note the date, start URLs, tool version, scope and exclusions alongside the archive.
Use HTTrack for a mirror or archival output
HTTrack downloads a site recursively into a local directory, keeps links usable for local browsing, and can resume or update an existing mirror. Its project page lists HTTrack 3.50-4 dated 2026-09-25 and describes the software as free GPL software, with interfaces for Windows, Unix-like systems and Android. Check the project page for current downloads and platform details.
Rank #2
- Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
You can use the graphical interface to enter a project name and starting URL, choose a destination, and configure crawl limits and filters. If you prefer a shell or need options such as WARC output, consult the HTTrack command-line guide for the syntax supported by your installed version. The guide documents controls for rate, connection frequency, total transfer, time and file size. It says HTTrack identifies itself as HTTrack and follows robots.txt; do not disable limits or bypass access restrictions on a site you do not control.
Choose the output for the job: the local directory is convenient for browsing, while WARC is intended for a different preservation and replay use. HTTrack also documents CDXJ indexing and WACZ bundling. Confirm your workflow’s requirements before starting so the capture format and any indexes or bundles you retain are useful later.
What no crawler can guarantee
A mirror captures retrievable material, not necessarily all the content a human visitor can reach. Client-side interactions, login-only areas, API-fed content, geo-restricted pages, rate limits and out-of-scope hosts can all leave gaps. Crawlers can also miss resources loaded only after a particular interaction or state change.
Do not assume that a page is complete because its HTML file exists. Review crawl logs, inspect representative pages and verify the assets and sections important to your purpose. If a site permits a crawl only within certain paths, hosts or rates, keep those boundaries intact. Do not attempt to evade authentication, bot checks or other access controls.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #3
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Keep the crawl polite and protect your storage
- Respect site rules. Check the site’s stated access conditions and honor crawl exclusions. Wget’s manual warns recursive retrieval can overload a server and recommends a delay between accesses.
- Limit the scope. Start at the pages you need and restrict domains or directories where appropriate; linked external sites can greatly expand a crawl.
- Throttle requests. Keep or increase the delay between requests, and use HTTrack’s rate and connection controls where relevant. Avoid aggressive crawling.
- Check disk space. Recursive downloads can consume substantial storage. The GNU Wget manual cautions: “Of course, recursive download may cause problems on your machine. If left to run unchecked, it can easily fill up the disk.”
- Preserve context. Store the crawl date, tool/version, URLs, scope, exclusions and logs with the files so the snapshot’s limits are clear.
An external hard drive can be useful for retaining a large mirror or WARC collection, but it is optional; size the storage to the archive you actually intend to keep.
After the crawl: verify and maintain the archive
- Read the log. Identify failures, redirects, excluded URLs and unexpected scope expansion. A successful process exit alone does not establish completeness.
- Spot-check different page types. Open the home page, deep pages, and pages with important images, downloads or styles. Check local navigation and assets.
- Compare against a page list. If you have an authoritative list of URLs, check that important entries were captured and record any gaps.
- Retain evidence of scope. Save the crawl log and metadata with the snapshot; note anything that required a login or interaction and was not captured.
- Update deliberately. HTTrack can resume or update mirrors. Keep a record of each update rather than silently replacing the only copy of an earlier snapshot.
Troubleshooting common problems
Local pages show broken links or missing images
Confirm you used --page-requisites and --convert-links with Wget, or the corresponding mirror options in HTTrack. Check whether the resource was hosted on another domain outside the crawl scope, excluded by rules, or loaded only after a script or interaction. A missing file cannot be repaired by link conversion alone.
The crawl stops before reaching important pages
Check the starting URL, scope restrictions, logs and the tool’s recursion behavior. Wget’s default recursive depth is five unless changed; --mirror sets infinite depth. Infinite depth can also expand the crawl substantially, so narrow the scope rather than assuming unlimited traversal is harmless.
The archive is unexpectedly large or the crawl takes too long
Look for links to external hosts, calendars, search pages or other URL patterns that generate many distinct pages. Restrict the crawl to the intended site section and set sensible rate, time, transfer or file-size limits using the options available in your chosen tool. Check available disk space before continuing.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #4
- Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
A page needs a login, form submission or interaction
A conventional recursive mirror may not capture content requiring a session or user action. Archive only content you are authorized to access, and do not bypass controls. Document excluded areas; for a site you operate, consider an export or a purpose-built capture process for the underlying data.
Wget options behave differently than expected
Check the installed Wget version and its manual. Available switches and behavior can vary by version, and a command copied from another environment may not match your installation. The GNU Wget manual is the reference for recursive retrieval and option details.
HTTrack output is not the archival record you need
A browsable mirror and a WARC capture have different purposes. Consult the HTTrack command-line guide for WARC, CDXJ and WACZ options, and verify that the resulting files fit the replay or preservation workflow you intend to use.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup:
ScreenshotNeo is a website screenshot API and MCP server, not a replacement for a recursive website mirror or WARC archive. It can preserve a page visually as an image or PDF when that is what you need. One GET request returns a screenshot; see the ScreenshotNeo API documentation for parameters and response details.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
- Plug-and-play expandability
- SuperSpeed USB 3.2 Gen 1 (5Gbps)
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o shot.webp
With ScreenshotNeo, cookie banners are accepted and removed before capture, along with known consent banners, newsletter popups and chat widgets; those cleanup steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server exposes take_screenshot, get_page_info and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. These are screenshot captures, not an entire site’s linked-page archive. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Does downloading a website save its database and original files?
No. A crawler retrieves discoverable, accessible web resources; it does not export the site’s server-side database or guarantee the original source files.
Can I use a mirror as a permanent archival record?
A browsable directory is useful for offline access, but it is not the same as a WARC record. Choose and verify an archival format that matches your preservation or replay workflow.
Can I archive a site that requires a password?
Only capture content you are authorized to access, and do not bypass access controls. A conventional crawl may not capture session-only or interaction-dependent material.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

