Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use private cloud object storage as the system of record for crawled pages. Write each fetch as an immutable object containing the raw response and its metadata, give it a deterministic key, and keep a database or search index that maps URLs and crawl times to object locations. Enable versioning or deletion recovery before production recrawls. Lifecycle rules can move older crawls to cheaper storage, while short-lived signed URLs let reviewers inspect a page without making the bucket public.

Amazon S3, Google Cloud Storage, and Azure Blob Storage all fit this pattern. The best choice depends less on storing HTML than on your required recovery window, region, identity system, event integrations, and retrieval and egress costs.

Use object storage for page bodies and an index for discovery

A crawler produces different kinds of evidence: raw HTML bytes, normalized HTML, response headers, status codes, screenshots, PDFs, parser output, and decisions such as robots or consent handling. Object storage is suited to these large, immutable artifacts. A database should hold the searchable facts needed to find them.

What belongs in an object

  • Raw response bytes exactly as fetched.
  • Normalized HTML, if your pipeline creates it.
  • Response headers, HTTP status, crawl timestamp, canonical URL, and content hash.
  • Parser version and robots, consent, or other fetch decisions.
  • Optional screenshots, PDFs, extracted text, and structured results.

What belongs in the index

Store URL, host, crawl time, job or run ID, status, content hash, parser version, object key, storage provider and region, retention state, and any error code. The index lets you answer “show me every successful crawl of this URL in March” without listing or scanning a bucket.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Do not use a filename supplied by the source website as your primary key. It can collide, change, contain unsafe characters, or omit the crawl time. Treat it as metadata only.

Design deterministic keys that make retrieval repeatable

A practical key template is:

host/crawl-date/job-id/content-hash/artifact-type.ext

For example, example.com/2026-09-29/run-1842/sha256-abc123/raw.html and example.com/2026-09-29/run-1842/sha256-abc123/headers.json describe artifacts from one fetch. Keep the full URL and crawl timestamp in the index because hosts can serve different paths and query strings.

Hashing and deduplication

Hash the response bytes before writing the object. The hash provides an integrity check and identifies identical content across crawls. You may point multiple index rows at one content object when the bytes are identical, while still retaining each crawl’s URL, timestamp, status, and decision metadata.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Separate raw evidence from derived data

Never overwrite raw HTML with cleaned or parsed output. Put normalized HTML, extracted text, and parser results under separate artifact names and record the parser version. When a parser changes, you can regenerate derived objects from the original response.

A provider-neutral crawl storage workflow

  1. Capture. Fetch the page and record response bytes, status, headers, canonical URL, crawl time, content hash, parser version, and robots or consent decision.
  2. Build the key. Create a deterministic path from host, date, run ID, hash, and artifact type. Keep keys stable even if a display title changes.
  3. Write the object. Upload raw bytes and metadata to a private bucket or container. Use retries and resumable or multipart uploads when objects are large or networks are unreliable.
  4. Write the index row. Commit the object location and crawl facts only after the upload is confirmed. Include an upload checksum or provider version ID where available.
  5. Emit an event. Publish an object-created event to a queue or Pub/Sub equivalent so indexing, parsing, and alerting are decoupled from the fetch worker.
  6. Apply lifecycle policy. Keep recent data in a hot tier. Transition older objects by age or access pattern, and delete only after the retention period and legal requirements allow it.
  7. Review safely. Generate a signed URL with the shortest practical lifetime instead of granting reviewers bucket credentials.

Amazon S3, Google Cloud Storage, or Azure Blob Storage?

All three are API-accessible object stores. Compare the controls around the storage, not just the upload call.

Rank #2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
Provider Documented capabilities relevant to crawls Important qualification
Amazon S3 S3 REST API for storing and retrieving objects; versioning; Object Lock; replication; encryption; and least-privilege IAM controls. AWS states a designed durability of 99.999999999% for S3 Standard objects over a given year. That is a durability design target, not a guarantee that your application has 99.99% availability.
Google Cloud Storage Managed buckets with Standard, Nearline, Coldline, Archive, and Rapid classes; object versioning; soft delete; retention policies; lifecycle management; signed URLs; and event notifications. Google documents read-after-write and listing as strongly consistent. Its current overview states a seven-day default soft-delete retention for new buckets; verify the setting when creating a bucket because defaults can change. Google Support lists up to 5 TB per object; confirm the limit for your selected API and object type.
Azure Blob Storage Encryption by default, customer-managed keys, soft delete for blobs and containers, and resource locks to reduce accidental deletion. Verify current Azure tier names and retention defaults when you configure the account; those policies can change.

How to choose

  • Choose the provider that matches your crawler’s runtime, identity provider, and region requirements.
  • Prefer strong read-after-write behavior when a worker must index an object immediately after upload.
  • Check event delivery options if parsing should start automatically after capture.
  • Model minimum-storage charges, retrieval fees, and egress for your access pattern; exact prices vary by region and tier.
  • Confirm that versioning, retention, and delete recovery meet the period during which a crawl may need to be restored.

Protect crawl history before the first production recrawl

Versioning and retention

Enable the provider’s object versioning or equivalent before workers begin overwriting keys. S3 Versioning preserves, retrieves, and restores object versions; Object Lock can enforce retention. Google Cloud Storage provides object versioning and retention controls, while soft delete protects recently deleted data. Azure provides soft delete for blobs and containers and resource locks. Select a recovery window based on operational and legal needs rather than leaving the default unexamined.

Encryption and identity

Keep buckets and containers private. Use the narrowest identity permissions that let a fetch worker write its prefix and let readers retrieve only approved objects. Encrypt in transit and at rest; use customer-managed keys when your compliance requirements demand control over key lifecycle. Do not place long-lived provider credentials in crawler URLs or page metadata.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Signed review links

A signed URL grants temporary access to one object without making the bucket public. Set an expiry measured in minutes or hours, restrict the object and HTTP method, and avoid embedding secrets in the object itself. Revoke broad access by changing the signing policy rather than distributing permanent links.

Lifecycle and cost control

Keep active analysis data in a hot class, then transition by age or access frequency. A common policy is to retain recent crawls for fast investigation, move older immutable artifacts to a lower-cost class, and delete only after the documented retention period. Apply the policy to raw and derived artifacts separately if parsed output can be regenerated.

Storage cost is only one line item. Include request charges, minimum-storage-duration rules, retrieval charges for cold tiers, cross-region replication, and internet egress. A crawl that is cheap to store can become expensive if analysts repeatedly restore archival objects. Measure access patterns before selecting an archive tier, and obtain current regional pricing from the provider you choose.

Retrieval patterns that survive real crawler failures

Fetch by index, then object key

Resolve a URL and time range in the index, select the desired crawl row, and retrieve the recorded object key. Do not enumerate an entire bucket to find a page. If the object is missing, compare the index’s provider version ID or checksum with lifecycle and deletion logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
  • Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Handle interrupted transfers

Use retries with exponential backoff. For large HTML bundles, media, or PDFs, use resumable or multipart uploads so a transient failure does not restart the entire transfer. Record an upload as complete only after the provider confirms it; otherwise mark the crawl as incomplete and make it visible to retry workers.

Verify integrity

Recompute the content hash after download and compare it with the index. A mismatch indicates corruption, an incorrect key, or that you retrieved a different object version. Keep the original response bytes so parser bugs can be corrected without recrawling a site.

Troubleshooting common failures

The index row exists but the object is missing

Cause: the index was committed before upload completion, a lifecycle rule deleted the object, or a worker used the wrong region or prefix.

Fix: make object confirmation a prerequisite for the index commit, inspect object-version and deletion logs, and compare the recorded key byte-for-byte with the upload request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every recrawl appears to overwrite history

Cause: keys contain only the URL or filename, and versioning is disabled.

Fix: add crawl date, run ID, and content hash to the key, then enable versioning or delete recovery before the next run.

Rank #4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
  • Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Reviewers receive an access-denied error

Cause: the bucket is private, the signed URL expired, or the URL was generated for the wrong object version.

Fix: issue a new short-lived URL for the exact key and version. Do not solve the problem by making the entire bucket public.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Uploads fail during traffic bursts

Cause: workers open too many connections, retry immediately, or send large objects as one request.

Fix: cap concurrency, use exponential backoff, and switch large transfers to resumable or multipart upload. Queue excess work instead of dropping crawl records.

Cold-tier retrieval is unexpectedly slow or costly

Cause: lifecycle rules moved frequently reviewed pages into an archive class.

Fix: keep investigation windows hot, restore in batches where supported, and revise transitions using measured access frequency and current regional pricing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
UnionSine 500GB Ultra Slim Portable External Hard Drive HDD-USB 3.0
  • [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
  • 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
  • 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
  • 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
  • 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Adding screenshots without making the bucket your browser

For screenshot capture, ScreenshotNeo is the first option to try because it removes common consent banners, popups, and chat widgets before capture, bills only clean shots, and has the lowest paid plan among the stated options. Store the returned PNG, JPEG, WebP, or PDF beside the HTML object using the same deterministic key.

Or skip the browser setup

One GET request captures a URL without you managing a browser. See the ScreenshotNeo documentation for all parameters.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo can load lazy images in full-page captures, select one element by CSS selector, emulate dark mode or one of 12 device presets, use any viewport and retina scale, produce PDFs with paper size, margins, orientation, and page ranges, render HTML/CSS, run custom JavaScript, click before capture, hide selectors, wait for a selector, delay, or network idle, block ads, trackers, requests, or resource types, set headers, cookies, user agent, Authorization, timezone, and geolocation, use transparent backgrounds, resize images, cache with a chosen TTL, create signed links for public <img> tags, submit asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, and expose usage and OpenAPI APIs. Its parameter names also match those used by other screenshot APIs, which eases migration.

For crawler accounting, the response identifies the page result and billing state with X-Page-Verdict and X-Billed headers. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. An MCP server supplies take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Free plan includes 1,000 shots per month with no card. Paid plans are Starter $5 for 3,000 shots, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Sign up for the free 1,000-shot plan.

Operational checklist

  • Private bucket or container created in the required region.
  • Deterministic key convention documented and tested for URL collisions.
  • Index schema stores URL, crawl time, status, hash, key, and version information.
  • Versioning, soft delete, Object Lock, retention, or resource locks enabled as required.
  • Encryption and least-privilege writer and reader identities configured.
  • Lifecycle transitions reviewed against retrieval frequency and current pricing.
  • Retry, resumable or multipart upload, checksum verification, and incomplete-upload handling implemented.
  • Object-created events routed to parsing and indexing workers.
  • Signed review URLs expire quickly and never expose bucket credentials.
  • Restore and accidental-delete procedures tested before large recrawls.

Frequently Asked Questions

Should I store compressed HTML or the original response?

Keep the original response as the authoritative artifact. Compression can reduce transfer and storage use, but retain enough metadata to verify the uncompressed content hash and preserve the exact bytes your crawler received.

Can one bucket serve several crawler projects?

Yes, if each project has an isolated prefix, separate identities, lifecycle policy, and index namespace. Separate buckets are preferable when retention, region, or access rules differ materially.

How long should crawled pages be retained?

Set the period from your legal, editorial, and analytical requirements. Document the window, test restoration before it expires, and apply the same policy to object versions and derived artifacts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do signed URLs make a public bucket unnecessary?

Yes for normal review workflows: keep the bucket private and issue a time-limited signed URL for the specific object. A signed link is access delegation, not a reason to grant anonymous bucket listing.

Quick Recap

SaleBestseller No. 1
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$129.99
Bestseller No. 2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$219.99
Bestseller No. 3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.80
Bestseller No. 4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$189.90

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.