Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping APIs can involve two separate caches: ordinary HTTP caches between your client and the website, and an application-level cache maintained by the scraping service itself. They may operate independently, use different keys and expiration rules, and expose different ways to force a fresh fetch. A successful API call therefore does not, by itself, prove that the target page was downloaded again.

How does caching work in a web scraping API?

When your scraper requests a URL, the request can pass through several layers:

  1. Your client cache, such as a browser, SDK, reverse proxy or local HTTP cache.
  2. Intermediary caches, including CDNs and shared proxies.
  3. The scraping provider’s fetcher, which may make a request to the origin website.
  4. An application or result cache, where the provider might retain HTML, extracted fields, screenshots or other results.

Only the first two layers are governed directly by HTTP caching semantics. A provider can add application behavior above HTTP, or use no shared result cache at all. RFC 9111 cautions that when an application caches data without making that behavior apparent or controllable, it should define its operation with respect to HTTP cache directives so users are not surprised.

Think of a cache as a decision system: it identifies an equivalent request, checks whether its stored representation is still fresh, and either serves it, validates it, or fetches a new representation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTTP cache freshness: age versus lifetime

An HTTP response has a current age and a freshness lifetime. A response is normally reusable while its age is less than its lifetime. Explicit lifetime information commonly comes from:

  • Cache-Control: max-age=... for general caches.
  • Cache-Control: s-maxage=... for shared caches.
  • Expires: ..., an absolute expiration date.

If explicit expiration is absent, a cache can sometimes apply heuristic freshness based on other response metadata. Heuristic behavior is implementation-dependent, so do not infer a precise retention period from a response that lacks an explicit lifetime.

Fresh responses

A fresh response can satisfy an equivalent request without contacting the origin. This reduces network work and latency, but the returned representation may not include changes made at the website after the cached response was stored.

Stale responses

Once a response is stale, a cache generally validates it before reusing it where the protocol and configuration permit. Validation asks the origin whether the stored representation is still current rather than downloading the complete body unconditionally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What cache keys decide whether two requests are equivalent?

For a generic HTTP cache, the primary key includes the request method and target URI. The response’s Vary header can add request-header dimensions when selecting a stored response. For example, a response varying on Accept-Language cannot safely be reused for every language.

An application cache may define additional dimensions. A scraping service could distinguish requests by query parameters, JavaScript-rendering settings, proxy location, cookies, authentication, user agent, viewport, extraction schema or other options. Those are design possibilities, not universal provider behavior. Ask the service to document exactly which fields participate in equivalence.

Two calls with the same URL can therefore produce different cache outcomes if headers, cookies, credentials, rendering instructions or extraction options differ. Conversely, a service might treat calls with different incidental parameters as equivalent if those parameters are ignored by its cache key.

Cache-Control, no-cache and no-store

no-cache

no-cache does not mean “never store.” It means a stored response must be validated before reuse. A request can send Cache-Control: no-cache to ask intermediaries to revalidate, subject to the intermediary’s supported behavior and the origin’s response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

no-store

no-store tells a compliant cache not to store the request or response. It is the directive to use when storage itself is unacceptable, such as for particularly sensitive content. It is different from no-cache, which permits storage but requires validation before reuse.

Other useful directives

  • max-age=0 makes a response immediately stale from the client’s perspective; it is commonly used to request revalidation, not guaranteed deletion of every provider-side result.
  • max-stale allows a client to accept a response older than its freshness lifetime, optionally by a stated tolerance.
  • private limits a response to private caches and prevents reuse by shared caches that honor the directive.
  • Vary identifies request headers that affect representation selection.

Request and response directives are not interchangeable. A request header expresses the caller’s preference; a response header states constraints or metadata from the origin. An application-level cache may also support only a subset of these controls.

ETag and Last-Modified revalidation

An origin can provide an ETag, an opaque version identifier. During revalidation, the cache sends If-None-Match with that value. If the representation has not changed, the origin can return 304 Not Modified, allowing the cache to reuse its stored body.

Last-Modified records a modification time. A cache can send If-Modified-Since and receive 304 Not Modified when appropriate. ETags are generally better at distinguishing versions, while timestamps can have limited precision or unreliable origin clocks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Revalidation is not the same as a guaranteed fresh scrape: it still depends on the origin, intermediary and scraper implementation honoring validators.

Application-level caches in scraping services

A scraping API may cache raw pages, rendered DOM, extracted JSON, screenshots, PDFs or job results. It may also bypass shared result caching entirely and perform a new fetch for every call. HTTP rules do not establish the application’s retention period, key, invalidation process or fresh-fetch control.

Before relying on reuse, ask these specific questions:

  • Is there an application/result cache, or only ordinary HTTP caching?
  • What request attributes form the cache key?
  • What is the retention period or TTL?
  • Are origin Cache-Control, Expires, ETag and Last-Modified honored?
  • Can a caller force bypass, revalidation or invalidation?
  • Are cache hits, misses, age and revalidation visible in headers or response metadata?
  • How are personalized, authenticated or sensitive responses handled?

What implementation examples teach

Scrapy’s documented HTTP cache (2.0.1)

Scrapy’s 2.0.1 documentation describes an HTTP cache that can return a stored response for the same request without another Internet transfer. Its documented RFC2616Policy handles directives and metadata including no-store, no-cache, max-age, Expires, Last-Modified, Age, Date, ETag and Last-Modified revalidation, plus request max-stale. The same documentation lists omissions, including Vary support and invalidation after updates or deletes. This older example demonstrates that “standards-aware” does not mean complete RFC coverage, and it should not be read as a statement about current Scrapy releases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apigee response cache

Google Apigee documents a response-cache policy that supports only a subset of Cache-Control capabilities, does not support inbound client Cache-Control headers, and supports only public caches. When configured to use response cache headers, max-age can determine cache duration, subject to other policy settings. The lesson is practical: read the specific service’s supported-controls documentation instead of assuming every HTTP directive works.

Does a scraping API cache my requests?

There is no universal answer. The cited Zyte HTTP API reference documents a single-URL extraction endpoint that blocks until the result is ready, but it does not specify cache keys, lifetimes, bypass controls or reuse of identical requests. ScrapingBee’s cited scraping API and proxy documentation likewise does not establish whether repeated calls are cached, how equivalence is defined, how long data remains available or how to bypass a cache. Treat those behaviors as undocumented until the provider publishes current, explicit details.

How to request a fresh page

  1. Check the provider’s documentation for a named cache-bypass, force-refresh, revalidation or invalidation parameter.
  2. If supported, use that control rather than inventing a query-string nonce; a nonce can change billing, analytics and origin behavior without bypassing an application cache.
  3. Send appropriate request headers, such as Cache-Control: no-cache, when the service documents that it forwards or honors them.
  4. Confirm the response’s cache indicators, age and provider metadata. Do not assume a successful response proves an origin transfer.
  5. For authenticated or personalized pages, verify whether the provider stores results and how credentials are isolated.

If no documented control exists, ask the vendor for the cache key, TTL, invalidation semantics and observability before building correctness requirements around freshness.

Performance, reliability and cost trade-offs

Benefits of reuse

  • Fewer origin requests and lower bandwidth consumption.
  • Lower latency for repeated reads.
  • Reduced load on target websites and more predictable processing.

Risks of reuse

  • Stale prices, inventory, availability or article content.
  • Cross-user exposure if personalized data enters an incorrectly shared cache.
  • Unexpected results when rendering settings or headers are absent from the key.

Choose a freshness policy per data type. Frequently changing records may require revalidation or a short TTL; immutable assets can tolerate long retention. Log the URL, relevant request dimensions, cache decision and returned age where your provider makes those fields available.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your actual need is a clean website capture rather than custom cache engineering, ScreenshotNeo provides a one-call screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the API as shown in the ScreenshotNeo documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for AI agents. Every plan includes the features; the Free plan provides 1,000 screenshots per month without a card, and paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting cache surprises

The page is old even after a successful request

Check response age and provider cache metadata. A fresh intermediary response or application result may be satisfying the call. Use the documented bypass or revalidation option; do not assume max-age=0 reaches the provider’s result cache.

Every request appears to miss

Different cookies, authorization headers, query parameters, rendering options or proxy locations may be part of the cache key. Normalize only values that are safe to share, and confirm the key definition with the provider.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

no-cache has no effect

The service may not forward inbound cache directives or may implement only a subset of HTTP behavior. Apigee’s documented policy is an example of such limitations. Look for a provider-specific refresh control.

Conditional requests return full bodies

The origin may not support validators, the validator may have changed, or an intermediary may have removed conditional headers. Inspect ETag, Last-Modified, If-None-Match and If-Modified-Since where you control the HTTP client.

Private data appears reusable

Stop sharing the response, review private and no-store handling, rotate exposed credentials and obtain the provider’s data-retention and isolation policy before continuing.

Frequently Asked Questions

Is a CDN cache the same as a scraping API result cache?

No. A CDN is an HTTP intermediary, while a scraping API may separately cache fetched pages, extracted results or rendered artifacts using its own key and TTL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I guarantee a fresh scrape with a random query parameter?

No. It may alter the target request without bypassing an application cache, and it can affect analytics, routing or billing. Use a documented refresh control.

What should I record for cache debugging?

Record the request method, URL, relevant headers and options, response age, validators, provider cache indicators and the time of each attempt.

The Bottom Line

Separate HTTP caching from provider-side result caching. Judge freshness by documented keys, TTLs, validators, bypass controls and observability—not by the fact that an API returned successfully.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.