Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can make a useful offline copy of a website with a recursive crawler, but no crawler can promise a complete copy of everything a site serves. For a browsable mirror, GNU Wget is a practical command-line choice; HTTrack offers graphical and command-line workflows. If you need a preservation-oriented collection, consider retaining WARC output as well as—or instead of—a directory of downloaded pages. Start with a small, polite crawl, inspect its log, and check key pages and assets before treating the archive as complete.

What “download an entire website” means

A crawler begins at one or more URLs, follows links and retrieves resources within the scope and access rules you set. The result is a crawl snapshot: a record of what the tool could discover and retrieve at that time. It is not a copy of a site’s underlying database or a guarantee that every page and behavior has been captured.

Some material may be inaccessible to a crawler: content behind authentication, forms or interactions that must be triggered, data fetched dynamically from APIs, resources hosted outside your crawl scope, or pages blocked by access rules. A page may download while still missing an image, stylesheet, font or other linked asset. Record the crawl date, starting URLs, tool and version, scope, and known exclusions so future readers understand what the archive represents.

Choose an archiving approach

Approach Best for What to expect
GNU Wget A scoped crawl from the command line Downloads files into a directory; options can retrieve page requisites and convert links for local browsing.
HTTrack A browsable mirror with GUI or command-line controls Downloads recursively, maintains a usable local link structure, and can resume or update a mirror.
ArchiveBox A self-hosted collection in several capture formats Its project describes captures including HTML, screenshots, PDF and WARC using tools such as Chrome and wget; results depend on the site and extractor.
WARC-oriented capture Preservation and replay workflows A WARC record serves a different purpose from an ordinary directory mirror. Some workflows retain both.

For a simple offline website you can open as files, begin with Wget or HTTrack. If the archive must support preservation or replay workflows, check whether WARC—and, where relevant, CDXJ indexing or WACZ bundling—fits your requirements. HTTrack documents WARC output alongside its browsable mirror and notes that the mirror is not a substitute for the WARC record. See the HTTrack project and its command-line guide. ArchiveBox’s supported extractors and outputs should not be assumed to work identically on every site; see the ArchiveBox project.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Make a browsable mirror with GNU Wget

Install GNU Wget for your operating system, make sure the destination drive has adequate free space, and run this starting command in a shell. Replace the example URL with a site you are permitted to archive.

wget --mirror --convert-links --page-requisites --adjust-extension --wait=1 --directory-prefix=./site-archive https://example.com/

This command is a starting point, not a completeness guarantee. It asks Wget to retrieve recursively and save the result under ./site-archive.

What each option does

  • --mirror enables recursive retrieval, timestamping and infinite depth. Without mirror mode, GNU Wget documents a default recursive depth of five.
  • --page-requisites retrieves files needed to render pages, such as images and stylesheets.
  • --convert-links rewrites links in downloaded documents for local viewing.
  • --adjust-extension helps save HTML responses with an appropriate extension.
  • --wait=1 adds a one-second delay between retrievals. Use a slower crawl if the site is small or its operator requests it.
  • --directory-prefix=./site-archive places the downloaded files in the named local directory.

Wget follows links in HTML and CSS during recursive retrieval and honors robots rules, according to the GNU Wget manual. Set the starting URL and use appropriate directory or domain restrictions to keep the crawl in scope. Review the manual for the installed Wget version before relying on particular switches or behavior.

Run a small trial first

  1. Choose a representative starting page and a narrow scope, such as the relevant section of a site rather than every linked domain.
  2. Run the command with the destination directory on a drive with enough free space. Watch the output for failed requests, excluded paths and unexpected hosts.
  3. Open several downloaded pages locally. Check navigation, images, stylesheets and other important resources—not just the landing page.
  4. Adjust scope or settings if needed, then run the crawl at a rate that respects the site and its operator’s rules.
  5. Keep the log and note the date, start URLs, tool version, scope and exclusions alongside the archive.

Use HTTrack for a mirror or archival output

HTTrack downloads a site recursively into a local directory, keeps links usable for local browsing, and can resume or update an existing mirror. Its project page lists HTTrack 3.50-4 dated 2026-09-25 and describes the software as free GPL software, with interfaces for Windows, Unix-like systems and Android. Check the project page for current downloads and platform details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

You can use the graphical interface to enter a project name and starting URL, choose a destination, and configure crawl limits and filters. If you prefer a shell or need options such as WARC output, consult the HTTrack command-line guide for the syntax supported by your installed version. The guide documents controls for rate, connection frequency, total transfer, time and file size. It says HTTrack identifies itself as HTTrack and follows robots.txt; do not disable limits or bypass access restrictions on a site you do not control.

Choose the output for the job: the local directory is convenient for browsing, while WARC is intended for a different preservation and replay use. HTTrack also documents CDXJ indexing and WACZ bundling. Confirm your workflow’s requirements before starting so the capture format and any indexes or bundles you retain are useful later.

What no crawler can guarantee

A mirror captures retrievable material, not necessarily all the content a human visitor can reach. Client-side interactions, login-only areas, API-fed content, geo-restricted pages, rate limits and out-of-scope hosts can all leave gaps. Crawlers can also miss resources loaded only after a particular interaction or state change.

Do not assume that a page is complete because its HTML file exists. Review crawl logs, inspect representative pages and verify the assets and sections important to your purpose. If a site permits a crawl only within certain paths, hosts or rates, keep those boundaries intact. Do not attempt to evade authentication, bot checks or other access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
  • Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Keep the crawl polite and protect your storage

  • Respect site rules. Check the site’s stated access conditions and honor crawl exclusions. Wget’s manual warns recursive retrieval can overload a server and recommends a delay between accesses.
  • Limit the scope. Start at the pages you need and restrict domains or directories where appropriate; linked external sites can greatly expand a crawl.
  • Throttle requests. Keep or increase the delay between requests, and use HTTrack’s rate and connection controls where relevant. Avoid aggressive crawling.
  • Check disk space. Recursive downloads can consume substantial storage. The GNU Wget manual cautions: “Of course, recursive download may cause problems on your machine. If left to run unchecked, it can easily fill up the disk.”
  • Preserve context. Store the crawl date, tool/version, URLs, scope, exclusions and logs with the files so the snapshot’s limits are clear.

An external hard drive can be useful for retaining a large mirror or WARC collection, but it is optional; size the storage to the archive you actually intend to keep.

After the crawl: verify and maintain the archive

  1. Read the log. Identify failures, redirects, excluded URLs and unexpected scope expansion. A successful process exit alone does not establish completeness.
  2. Spot-check different page types. Open the home page, deep pages, and pages with important images, downloads or styles. Check local navigation and assets.
  3. Compare against a page list. If you have an authoritative list of URLs, check that important entries were captured and record any gaps.
  4. Retain evidence of scope. Save the crawl log and metadata with the snapshot; note anything that required a login or interaction and was not captured.
  5. Update deliberately. HTTrack can resume or update mirrors. Keep a record of each update rather than silently replacing the only copy of an earlier snapshot.

Troubleshooting common problems

Local pages show broken links or missing images

Confirm you used --page-requisites and --convert-links with Wget, or the corresponding mirror options in HTTrack. Check whether the resource was hosted on another domain outside the crawl scope, excluded by rules, or loaded only after a script or interaction. A missing file cannot be repaired by link conversion alone.

The crawl stops before reaching important pages

Check the starting URL, scope restrictions, logs and the tool’s recursion behavior. Wget’s default recursive depth is five unless changed; --mirror sets infinite depth. Infinite depth can also expand the crawl substantially, so narrow the scope rather than assuming unlimited traversal is harmless.

The archive is unexpectedly large or the crawl takes too long

Look for links to external hosts, calendars, search pages or other URL patterns that generate many distinct pages. Restrict the crawl to the intended site section and set sensible rate, time, transfer or file-size limits using the options available in your chosen tool. Check available disk space before continuing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
  • Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

A page needs a login, form submission or interaction

A conventional recursive mirror may not capture content requiring a session or user action. Archive only content you are authorized to access, and do not bypass controls. Document excluded areas; for a site you operate, consider an export or a purpose-built capture process for the underlying data.

Wget options behave differently than expected

Check the installed Wget version and its manual. Available switches and behavior can vary by version, and a command copied from another environment may not match your installation. The GNU Wget manual is the reference for recursive retrieval and option details.

HTTrack output is not the archival record you need

A browsable mirror and a WARC capture have different purposes. Consult the HTTrack command-line guide for WARC, CDXJ and WACZ options, and verify that the resulting files fit the replay or preservation workflow you intend to use.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup:

ScreenshotNeo is a website screenshot API and MCP server, not a replacement for a recursive website mirror or WARC archive. It can preserve a page visually as an image or PDF when that is what you need. One GET request returns a screenshot; see the ScreenshotNeo API documentation for parameters and response details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o shot.webp

With ScreenshotNeo, cookie banners are accepted and removed before capture, along with known consent banners, newsletter popups and chat widgets; those cleanup steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server exposes take_screenshot, get_page_info and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. These are screenshot captures, not an entire site’s linked-page archive. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.

Frequently Asked Questions

Does downloading a website save its database and original files?

No. A crawler retrieves discoverable, accessible web resources; it does not export the site’s server-side database or guarantee the original source files.

Can I use a mirror as a permanent archival record?

A browsable directory is useful for offline access, but it is not the same as a WARC record. Choose and verify an archival format that matches your preservation or replay workflow.

Can I archive a site that requires a password?

Only capture content you are authorized to access, and do not bypass access controls. A conventional crawl may not capture session-only or interaction-dependent material.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.99
Bestseller No. 2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$229.99
Bestseller No. 3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.80
Bestseller No. 4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$149.84

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.