Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTTP cookies give a scraper continuity between requests. A server sends a cookie in a Set-Cookie response header; a client stores it with its scope and lifetime, then returns eligible name-value pairs in a later Cookie request header. Using a standards-aware cookie jar or HTTP session is safer and more reliable than copying one cookie string between requests.

How do cookies work in web scraping?

HTTP requests are independent by default. Without some state mechanism, a server may treat every request as coming from a new visitor. Cookies let the server and client associate successive requests with an application session, preferences, a login state or another identifier.

The exchange has two different headers:

  • Set-Cookie: sent by the server in a response. It contains a cookie name and value plus attributes such as domain, path, expiry and whether it requires HTTPS.
  • Cookie: sent by the client on a later request. It contains applicable cookie name-value pairs only; the attributes from Set-Cookie are not repeated.

A scraper should therefore model cookies as managed state, not as a permanent authorization token. The target application decides what a cookie means, and it may require additional headers, request data, JavaScript-generated state or an anti-bot challenge.

Why cookie scope matters

A cookie is returned only when its stored rules match the outgoing request. Important factors include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Host or domain: a cookie set for one host is not automatically valid for an unrelated host.
  • Path: a cookie limited to /account should not be sent to every path on the site.
  • Expiry and lifetime: an expired cookie must be discarded; session cookies normally last only for the client session.
  • Secure transport: a cookie marked Secure is intended for HTTPS requests.

If you reduce a cookie jar to a plain dictionary of names and values, you can lose this context, send a cookie to the wrong endpoint, or retain one after it has expired. Cookie selection is also library- and policy-dependent, so verify behavior against the HTTP client you use.

How do I maintain a session when scraping a website?

Use one HTTP session for the sequence of requests. In Python, a Requests Session stores response cookies and applies matching cookies to later requests. The standard library’s http.cookiejar provides the same kind of policy-aware storage for compatible clients.

Python Requests: login, then fetch a protected page

import os
import requests

login_url = "https://example.com/login"
account_url = "https://example.com/account"

with requests.Session() as session:
    session.headers.update({"User-Agent": "MyResearchBot/1.0"})

    response = session.post(
        login_url,
        data={
            "username": os.environ["SCRAPER_USER"],
            "password": os.environ["SCRAPER_PASSWORD"],
        },
        timeout=30,
    )
    response.raise_for_status()

    page = session.get(account_url, timeout=30)
    page.raise_for_status()
    print(page.text[:500])

The session receives any Set-Cookie headers from the login response and sends eligible cookies to the account URL. Keep credentials and cookies outside source control. Do not print the jar or a complete request header in production logs.

Persisting cookies between program runs

A normal in-memory session ends when the process exits. If a legitimate workflow needs persistence, serialize a cookie jar to a protected file and reload it on the next run. Restrict file permissions and delete the file when the session is no longer needed. A persistent jar containing authentication cookies should be treated like a password.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I send cookies with Python Requests?

Preferred: let a Session manage them

import requests

with requests.Session() as session:
    first = session.get("https://example.com/start", timeout=30)
    first.raise_for_status()

    second = session.get("https://example.com/next", timeout=30)
    second.raise_for_status()

Inspect cookie metadata without exposing secret values:

for cookie in session.cookies:
    print({
        "name": cookie.name,
        "domain": cookie.domain,
        "path": cookie.path,
        "expires": cookie.expires,
        "secure": cookie.secure,
    })

Narrow case: manually supply a controlled cookie

import requests

response = requests.get(
    "https://example.com/report",
    cookies={"report_view": "compact"},
    timeout=30,
)
response.raise_for_status()

This is useful for a known, non-sensitive test value or a tightly controlled request. It is easier to make stale, send to the wrong host or omit required scope information than it is with a session. Never paste a live authentication cookie into source code, tickets or examples.

Using Python’s standard-library cookie jar

import http.cookiejar
import urllib.request

jar = http.cookiejar.CookieJar()
opener = urllib.request.build_opener(
    urllib.request.HTTPCookieProcessor(jar)
)

with opener.open("https://example.com/start", timeout=30) as response:
    response.read()

with opener.open("https://example.com/next", timeout=30) as response:
    body = response.read()
    print(body[:500])

HTTPCookieProcessor extracts cookies from responses and adds applicable cookies to later requests according to the jar’s policy. This is preferable to assembling a raw Cookie header yourself.

When cookies are not enough

Cookies can preserve a session, but they do not prove that a request is authenticated or reproduce all browser behavior. A site may require a CSRF token in a form, an authorization header, a particular request sequence, JavaScript execution, a device signal or an anti-bot check. A successful login response can also set a cookie for a different path than the page you are requesting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use browser automation only when the site’s actual interaction requires browser-side behavior that ordinary HTTP requests cannot perform. Otherwise, a policy-aware HTTP session is simpler, faster and easier to inspect. Use only cookies legitimately obtained for the task, keep TLS enabled, and follow the target site’s access rules. Cookie flags have limits: HttpOnly restricts access through non-HTTP APIs, while Secure restricts transmission to secure channels; neither makes a cookie absolutely safe.

Cookie jar versus a manually supplied header

Approach Scope and expiry Persistence Best use
Session or cookie jar Retained and evaluated automatically Cookies received during the workflow are available to later requests Normal scraping sessions, redirects and login flows
Manual dictionary or Cookie header Usually discarded unless you implement it yourself You must update and replace values manually Small, controlled debugging requests

Debugging cookies step by step

  1. Inspect the response’s Set-Cookie headers, while redacting values in logs.
  2. List the jar’s cookie names, domains, paths, expiry times and secure flags.
  3. Compare the exact outgoing host and path with those attributes.
  4. Confirm the request uses HTTPS when the cookie is marked Secure.
  5. Check whether the cookie expired or was replaced by a later response.
  6. Compare the complete request flow, not just the cookie: method, redirects, CSRF fields, authorization and required headers may also matter.

Common symptoms and fixes

  • Logged out on every request: you created a new client for each request, discarded the jar, or did not complete the site’s login sequence. Reuse one session and check the login response.
  • Cookie appears stored but is not sent: host, path, expiry or HTTPS requirements do not match. Inspect those attributes rather than forcing a header.
  • 403 or bot-check response: cookies alone are not a general-purpose bypass. Stop, check permission and determine what additional application controls are required.
  • Unexpected stale behavior: an old persisted jar may contain expired or invalid authentication state. Remove it securely and establish a fresh authorized session.
  • Works in a browser but not in the script: the browser may be running JavaScript, sending other state, or following a different request sequence. Capture the minimum legitimate flow and reproduce only what the application requires.

Performance, reliability and privacy considerations

Reusing a session avoids repeatedly establishing application state and can reduce unnecessary login requests. Set explicit connect and read timeouts, handle redirects according to the site’s behavior, and add bounded retries only for transient transport failures. Do not retry a login or state-changing request blindly.

Cookies can identify a user across requests and, in third-party contexts, can support tracking. Limit collection to the cookies required for the task, protect storage and logs, and provide a deletion path. The legality and permission status of scraping depends on the target, your authorization and applicable rules; cookie mechanics alone do not answer that question.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean visual capture rather than parsing HTML, ScreenshotNeo provides a website screenshot API and MCP server. Its preprocessing accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing result.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for PNG, JPEG, WebP and PDF options, cookies and headers, waits, selectors, device settings and other controls. ScreenshotNeo also has an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Further learning

For a broader treatment of scraper login and cookie workflows, O’Reilly lists Web Scraping with Python, 3rd Edition by Ryan Mitchell, published in February 2024. It is an intermediate-to-advanced, 352-page book that includes a section on handling logins and cookies; it is not devoted exclusively to cookies.

Frequently Asked Questions

Does every website use cookies for login?

No. A site may combine cookies with authorization headers, CSRF tokens, JavaScript state or other controls, and some services use different session mechanisms.

Can I share one cookie jar between unrelated domains?

Do not assume that is safe or correct. Cookie policies use host and domain scope; keep jars separated when workflows have different trust or authorization boundaries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I log the Cookie header to diagnose a scraper?

Avoid logging live values. Inspect and log names plus non-secret metadata such as domain, path, expiry and secure status, and redact credentials.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.