Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Most failures blamed on BeautifulSoup happen before or after parsing. Requests or urllib can fail during DNS, TLS, connection, redirects, or HTTP status handling; BeautifulSoup can fail when its parser is unavailable or misconfigured; and extraction code can fail when a selector returns None. Split a scraper into request, response validation, parsing, extraction, and storage stages, then handle each stage with the narrowest useful recovery behavior.

What BeautifulSoup handles—and what it does not

BeautifulSoup accepts markup and a parser, builds a navigable tree, and provides search and extraction methods. It does not fetch URLs, retry requests, rotate proxies, enforce rate limits, or execute JavaScript. The official documentation describes its parser choices and tree-navigation behavior at BeautifulSoup documentation.

Stage Typical library Typical failures
URL construction Python and urllib.parse ValueError, malformed URLs
DNS, TCP, TLS, proxy connection Requests or urllib ConnectionError, SSLError, URLError
HTTP response Requests or urllib 401, 403, 404, 429, 500-series responses
HTML/XML parsing BeautifulSoup plus a backend parser FeatureNotFound, parser-specific errors, encoding problems
Element lookup BeautifulSoup Usually None or [], not an exception
Conversion and normalization Your Python code AttributeError, TypeError, ValueError, KeyError
Storage and export CSV, JSON, database libraries OSError, UnicodeEncodeError, database exceptions

A minimal safe request-and-parse sequence

import requests
from bs4 import BeautifulSoup

response = requests.get("https://example.com", timeout=15)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")

The timeout prevents an unbounded wait, raise_for_status() turns unsuccessful HTTP responses into HTTPError, and response.content gives BeautifulSoup the original bytes for its own encoding detection. Requests documents timeout and status behavior in its Quickstart.

Requests exceptions before BeautifulSoup runs

Requests exception classes and their conditions are listed in the Requests API reference. Catch specific, actionable classes first, then use RequestException as a Requests-specific fallback.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

try:
    response = requests.get(
        "https://example.com",
        timeout=(5, 20),  # connect timeout, read timeout
    )
    response.raise_for_status()
except requests.exceptions.ConnectTimeout as exc:
    print(f"Connection timed out: {exc}")
except requests.exceptions.ReadTimeout as exc:
    print(f"Response read timed out: {exc}")
except requests.exceptions.Timeout as exc:
    print(f"Timeout: {exc}")
except requests.exceptions.ConnectionError as exc:
    print(f"Connection failure: {exc}")
except requests.exceptions.HTTPError as exc:
    print(f"HTTP failure: {exc}")
except requests.exceptions.TooManyRedirects as exc:
    print(f"Redirect failure: {exc}")
except requests.exceptions.RequestException as exc:
    print(f"Other Requests failure: {exc}")
else:
    print(response.status_code)

Timeouts are not optional

Requests does not time out by default. A scalar timeout such as timeout=15 controls socket waiting, not necessarily the total wall-clock time for downloading a large response. Use a tuple such as timeout=(5, 30) when connection and read limits should differ. If the application needs a total operation deadline, add deadline logic outside Requests.

HTTP status is different from a transport failure

A server can return a normal response object with status 404, 403, 429, or 500. These statuses become HTTPError only after raise_for_status(). Handle statuses according to their meaning:

  • 401 or 403: authentication, permissions, application policy, or access controls; do not assume a user-agent change will solve it.
  • 404: the resource may be gone or the URL may have changed.
  • 429: rate limiting; honor Retry-After when supplied.
  • 500–599: possibly transient server failure, but not automatically safe to retry.
  • 3xx: Requests normally follows redirects, subject to its redirect limit.

A 200 response is not proof that the expected page arrived. It can contain a login form, consent page, anti-bot challenge, error template, or JavaScript shell.

Using urllib instead of Requests

The standard library raises HTTPError for HTTP-specific failures and URLError for broader network or address-resolution failures. Because HTTPError is a subclass of URLError, catch it first, as shown in the Python urllib guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen

request = Request(
    "https://example.com",
    headers={"User-Agent": "my-scraper/1.0 (+contact@example.com)"},
)

try:
    with urlopen(request, timeout=15) as response:
        markup = response.read()
except HTTPError as exc:
    print(f"HTTP error {exc.code}: {exc.reason}")
except URLError as exc:
    print(f"Network error: {exc.reason}")
else:
    soup = BeautifulSoup(markup, "html.parser")

BeautifulSoup parser exceptions and parser choice

FeatureNotFound

BeautifulSoup raises FeatureNotFound when the requested backend is unavailable. Install the dependency, or deliberately select an installed parser:

python -m pip install lxml html5lib
from bs4 import BeautifulSoup, FeatureNotFound

def parse_html(markup):
    try:
        return BeautifulSoup(markup, "lxml")
    except FeatureNotFound:
        return BeautifulSoup(markup, "html.parser")

Automatic fallback is convenient for a script but can hide an environment problem and change the resulting tree. For reproducible jobs, declare the chosen parser as a dependency and fail clearly if it is missing.

Backend differences matter

Parser Use and trade-off
html.parser Included with Python; no separate installation, but its tree construction may differ from browser-style parsing.
lxml External dependency with HTML and XML support; output can differ from other backends.
html5lib Browser-like HTML parsing, generally heavier and slower.
xml For XML documents; requires an XML-capable parser such as lxml.

Use the same backend in development and production. Do not call lxml universally “better”; compatibility, XML requirements, deployment, and resulting tree shape determine the right choice.

Malformed markup can parse successfully

Unclosed or misnested tags often do not raise an exception because the selected parser attempts recovery:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
html = "<html><body><p>Unclosed paragraph"
soup = BeautifulSoup(html, "html.parser")
print(soup.get_text())

Successful parsing does not prove that the tree is correct. Check required nodes and use diagnose() when investigating parser or markup behavior:

from bs4.diagnose import diagnose

diagnose(html)

The diagnostic utility and parser guidance are documented at BeautifulSoup documentation.

Missing tags: why no exception appears at first

Search methods return sentinel values:

title_tag = soup.find("h1")       # None when absent
links = soup.find_all("a")        # [] when no links match

# This fails later if title_tag is None:
# title = soup.find("h1").get_text(strip=True)

Use guarded extraction for optional fields:

title_tag = soup.select_one("h1")
title = title_tag.get_text(" ", strip=True) if title_tag else None

Classify absence rather than hiding it:

  • Expected absence: an optional field is not present.
  • Schema drift: a required selector no longer matches.
  • Programming error: code calls a method on None.
  • Wrong document: the response is a login, block, consent, or JavaScript page.

For a required field, fail with context:

required = soup.select_one("main article h1")
if required is None:
    raise ValueError("Required article heading was not found")

Encoding and content validation

Transport can succeed while decoding is wrong. Compare Requests’ declared and estimated encodings during diagnosis:

print(response.encoding)
print(response.apparent_encoding)
soup = BeautifulSoup(response.content, "html.parser")

apparent_encoding is a clue, not an unquestionable correction. Also check that the response is the kind of document your parser expects:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
content_type = response.headers.get("content-type", "")
if "html" not in content_type.lower():
    raise ValueError(f"Expected HTML, received {content_type}")

Keep decoding failures separate from later output failures such as UnicodeEncodeError when writing a file or database.

Retries: only for plausible transient failures

Retry connection resets, temporary proxy or DNS failures, connect timeouts, selected 5xx responses, and 429 responses when the server permits it. Do not automatically retry malformed requests, 401/403 authorization problems, genuine 404s, parser installation errors, selector failures, or conversion bugs.

import random
import time

def backoff_delay(attempt):
    return min(2 ** attempt + random.uniform(0, 0.5), 30.0)

for attempt in range(3):
    try:
        response = requests.get(url, timeout=(5, 20))
        response.raise_for_status()
        break
    except requests.exceptions.RequestException:
        if attempt == 2:
            raise
        time.sleep(backoff_delay(attempt))

Bound attempts, add jitter, respect Retry-After, and follow the target site’s terms and rate limits. Repeated requests can worsen blocking and server load.

A complete, structured scraper result

from dataclasses import dataclass
from typing import Optional
import logging
import requests
from bs4 import BeautifulSoup, FeatureNotFound

logger = logging.getLogger(__name__)

@dataclass
class ScrapeResult:
    url: str
    title: Optional[str]
    status: str
    error: Optional[str] = None

def scrape_page(url: str) -> ScrapeResult:
    try:
        response = requests.get(
            url,
            headers={"User-Agent": "example-scraper/1.0 (+contact@example.com)"},
            timeout=(5, 20),
        )
        response.raise_for_status()
    except requests.exceptions.Timeout as exc:
        logger.warning("Timeout while fetching %s: %s", url, exc)
        return ScrapeResult(url, None, "timeout", str(exc))
    except requests.exceptions.HTTPError as exc:
        status = exc.response.status_code if exc.response else None
        logger.warning("HTTP error for %s: status=%s error=%s", url, status, exc)
        return ScrapeResult(url, None, "http_error", str(exc))
    except requests.exceptions.ConnectionError as exc:
        logger.warning("Connection error for %s: %s", url, exc)
        return ScrapeResult(url, None, "connection_error", str(exc))
    except requests.exceptions.RequestException as exc:
        logger.exception("Requests failure for %s", url)
        return ScrapeResult(url, None, "request_error", str(exc))

    try:
        soup = BeautifulSoup(response.content, "lxml")
    except FeatureNotFound as exc:
        logger.error("Configured parser is unavailable: %s", exc)
        return ScrapeResult(url, None, "parser_unavailable", str(exc))

    title_tag = soup.select_one("h1")
    if title_tag is None:
        logger.info("Required title not found at %s", url)
        return ScrapeResult(url, None, "missing_title")

    return ScrapeResult(url, title_tag.get_text(" ", strip=True), "ok")

This design distinguishes network, HTTP, parser, and extraction outcomes. Those states can be counted, retried selectively, alerted on, or persisted without converting every problem into a vague “failed” result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Diagnosing common failures

AttributeError: 'NoneType' object has no attribute ...

The selector returned None. Inspect the actual response before changing the selector:

print(response.url)
print(response.status_code)
print(response.headers.get("content-type"))
print(response.text[:500])

A block page, login form, consent screen, or JavaScript shell is often the real cause.

FeatureNotFound

The requested backend is not installed. Install it with python -m pip install lxml, choose html.parser, or correct the deployment dependency. Do not silently switch parsers in a pipeline unless output differences are acceptable.

requests.exceptions.Timeout

The server, connection, or read stream exceeded the configured wait. Increase or split the timeout only after determining which phase is slow; catch Timeout before RequestException.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

requests.exceptions.ConnectionError

DNS, proxy, refusal, reset, or another network interruption prevented a usable connection. It does not prove that the target page is nonexistent.

HTTPError: 403 Client Error

Check the URL, authorization, access policy, and whether an official API or export exists. Slow requests and respect usage rules; do not present header changes or proxy rotation as guaranteed bypasses.

Empty results without an exception

Check selectors, page-structure changes, JavaScript-rendered content, iframes, and whether the body is a challenge or consent page. BeautifulSoup parses only the markup supplied to it; it does not execute JavaScript.

Conversion errors

Extraction can succeed while normalization fails:

def parse_price(text):
    if not text:
        return None
    cleaned = text.replace("$", "").replace(",", "").strip()
    try:
        return float(cleaned)
    except ValueError:
        return None

Keep ValueError, TypeError, AttributeError, and KeyError visible enough to diagnose. Never turn all of them into an unexplained missing record.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript, block pages, and choosing another tool

If the initial HTML lacks the data because JavaScript renders it later, find a legitimately available JSON endpoint or official API, or use Playwright/Selenium when browser execution is required. Browser automation adds CPU, memory, timing, and operational complexity.

  • Official API: prefer it for a stable schema, documented authentication, and defined quotas when available.
  • urllib.request: useful when standard-library dependencies matter.
  • lxml: useful for XML, XPath, or workloads where its parser behavior fits.
  • Scrapy: suited to queues, concurrency, middleware, pipelines, and crawl-wide retry policy.
  • Managed scraping or browser APIs: appropriate when proxy, geographic, rendering, or anti-bot infrastructure—not selectors—is the primary problem.

Paid services do not fix an incorrect selector, missing parser dependency, or invalid data validation. For static, low-volume HTML, Requests plus BeautifulSoup is usually the simplest option.

Observability checklist

Record enough context to classify failures without logging credentials, cookies, authorization headers, or sensitive body content:

  • Requested and final URL, timestamp, and attempt number.
  • HTTP status, content type, and response size.
  • Selected parser and exception class/message.
  • Selector or field that failed.
  • Whether the result was missing, blocked, malformed, retryable, or successful.

Useful result categories include ok, invalid_url, timeout, connection_error, http_401, http_403, http_404, http_429, http_5xx, parser_unavailable, parse_error, missing_required_field, unexpected_content_type, blocked_or_challenge, and storage_error.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Final debugging sequence

  1. Validate the URL and request parameters.
  2. Confirm that the request connected and note timeout type.
  3. Record the final URL after redirects and the HTTP status.
  4. Check content type, encoding, size, and a safe body preview.
  5. Confirm the parser is installed and identical across environments.
  6. Parse the bytes and inspect the tree if markup is malformed.
  7. Test selectors, distinguishing optional fields from required schema elements.
  8. Handle conversion and storage exceptions separately.
  9. Retry only classified transient failures with bounded, polite backoff.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.