Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can fetch a public page, check what kind of response it is, and pull specific fields out of it using only modules that ship with Python: urllib.request to retrieve the bytes, urllib.robotparser to read the site’s robots.txt rules, and html.parser, json or csv to parse the content. No pip install is needed for any of it. The method works when the data is present in the server’s raw response. It does not render pages, run JavaScript, or make a site’s terms of use irrelevant.

What “zero dependencies” covers and what it does not

The standard library gives you the tools for the whole workflow. It does not guarantee that every site will cooperate, that every response will be HTML, or that you are allowed to collect what you are looking at. Keep those three boundaries in mind as you read the steps below.

  • Covered: HTTP retrieval, reading response headers, decoding bytes, parsing HTML, JSON and CSV, and checking robots.txt rules, all with Python’s built-in modules.
  • Not covered: pages whose data appears only after client-side JavaScript runs, sites that require login or other access controls, and any question about a site’s terms, privacy obligations or applicable law.
  • Not implied: that Beautiful Soup, pandas, requests or another third-party package is needed for the core example. None is.

Step 1: Identify the response format before you write any parser

A URL does not promise HTML. A product listing might return an HTML page, a catalogue endpoint might return JSON, and a public dataset might return CSV or a downloadable binary file. Find out which one you are getting, because the parser follows the format, not the address bar.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Response type Typical Content-Type Standard-library tool What to watch for
Static HTML text/html html.parser.HTMLParser It is an event parser, not a DOM. You track state yourself.
JSON application/json json Keys and nesting can change without notice. Check for missing keys.
CSV text/csv csv Quoted fields, delimiters, and whether the first row is a header.
Plain text text/plain String methods Line endings and any leading or trailing whitespace.
Binary file Varies (for example application/octet-stream or an image type) Write the bytes to a file Do not decode it as text.

The Content-Type header is the server’s declaration of what it sent. Read it rather than guessing from the URL, and treat an unexpected type as a reason to stop.

Step 2: Check robots.txt before fetching

urllib.robotparser.RobotFileParser reads a site’s robots.txt file and answers one question: whether a given user agent may fetch a given URL under those rules. The function below builds the robots.txt address from the page’s scheme and host, reads it, and returns the answer.

import urllib.parse
import urllib.robotparser

USER_AGENT = "static-data-reader/1.0"

def robots_allows(url, user_agent=USER_AGENT):
    parts = urllib.parse.urlsplit(url)
    robots_url = urllib.parse.urlunsplit(
        (parts.scheme, parts.netloc, "/robots.txt", "", "")
    )
    parser = urllib.robotparser.RobotFileParser()
    parser.set_url(robots_url)
    parser.read()
    return parser.can_fetch(user_agent, url)

Call read() inside your own error handling, because it makes a network request. The answer reflects only what the robots.txt file says. It is one input to a decision about whether to collect data, not the decision itself.

Step 3: Fetch the response and handle failures

Use urllib.request.Request so you can set a User-Agent header, and use urlopen as a context manager so the connection closes when you finish reading. A GET request is the default when no data argument is supplied, which is what you want for reading public pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Build a Request with a descriptive User-Agent string.
  2. Call urlopen(request, timeout=10) inside a with block. The timeout is in seconds.
  3. Inside the same block, call response.read() to get the body as bytes. Read the status, the Content-Type and the charset from the headers while the response is open.
  4. Catch HTTPError before URLError, because HTTPError is a subclass of URLError. Catch TimeoutError for reads that time out.
import urllib.error
import urllib.request

USER_AGENT = "static-data-reader/1.0"
TIMEOUT_SECONDS = 10

def fetch(url):
    request = urllib.request.Request(url, headers={"User-Agent": USER_AGENT})
    try:
        with urllib.request.urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            body = response.read()
            status = response.status
            content_type = response.headers.get_content_type()
            charset = response.headers.get_content_charset()
    except urllib.error.HTTPError as err:
        raise RuntimeError(f"Server returned HTTP {err.code} for {url}") from err
    except urllib.error.URLError as err:
        raise RuntimeError(f"Could not reach {url}: {err.reason}") from err
    except TimeoutError as err:
        raise RuntimeError(f"Timed out reading {url}") from err
    return status, content_type, charset, body

The urllib reference notes that network operations can take arbitrarily long while a connection is being established. Setting a timeout and handling its failure is therefore part of a correct script, not an optional extra. Also note that urlopen raises an error for most 4xx and 5xx responses rather than handing you a body, so a failed page will not silently parse as empty content.

Step 4: Decode bytes on purpose

urlopen returns bytes because it cannot know how those bytes are encoded. The server may declare a charset in the Content-Type header, such as text/html; charset=ISO-8859-1, and response.headers.get_content_charset() returns that value, or None when the header does not include one.

  • If a charset is declared, decode with it: body.decode(charset).
  • If no charset is declared and the content is HTML, look for a <meta charset> declaration in the document before decoding. Stop and make an explicit decision if neither is present.
  • Do not write body.decode("utf-8") as a default for every site. It works for many pages and fails or garbles text on others, and the failure can be quiet.

Decode errors raise UnicodeDecodeError. Catch them and report the URL and declared charset, so a bad page is visible instead of producing wrong text.

Step 5: Parse according to the format

Static HTML with html.parser

HTMLParser is a callback-based parser. You subclass it and override handlers such as handle_starttag, handle_endtag and handle_data, and the parser calls them as it reads the markup. The example below collects the text of every <h2> element.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from html.parser import HTMLParser

class HeadingParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.headings = []
        self._in_h2 = False
        self._buffer = []

    def handle_starttag(self, tag, attrs):
        if tag == "h2":
            self._in_h2 = True
            self._buffer = []

    def handle_endtag(self, tag):
        if tag == "h2" and self._in_h2:
            self.headings.append("".join(self._buffer).strip())
            self._in_h2 = False

    def handle_data(self, data):
        if self._in_h2:
            self._buffer.append(data)

def extract_headings(html_text):
    parser = HeadingParser()
    parser.feed(html_text)
    parser.close()
    return parser.headings

This style gives you control and very little overhead, but it asks more of you. The parser can process invalid markup, yet it does not check that end tags match start tags, and it does not invoke every callback for elements that are implicitly closed. Track your own state as the example does, and make your extraction tolerant of a missing closing event. The parser is not a browser DOM and does not execute JavaScript. The more deeply nested and irregular the page, the more state logic you will write by hand, and that is the trade-off to weigh before choosing this approach over a single structured source.

JSON with the json module

When an endpoint returns JSON, parse it directly. json.loads accepts bytes and detects the UTF-8, UTF-16 or UTF-32 encoding from the first bytes, so you do not need to decode first. Then check the structure before you use it.

import json

def extract_titles(body):
    data = json.loads(body)
    items = data.get("items", [])
    return [item["title"] for item in items if "title" in item]

Use .get() and membership checks rather than direct indexing for fields that might be absent. A missing field should produce a clear warning, not a KeyError halfway through a batch.

CSV with the csv module

For delimited tabular data, decode the body and read it through csv.DictReader. Because the response is in memory, wrap the text in io.StringIO.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import csv
import io

def extract_rows(body, charset):
    text = body.decode(charset)
    return list(csv.DictReader(io.StringIO(text)))

Confirm that the first row is a header and that the column names match what you expect. If the file uses a delimiter other than a comma, pass delimiter= to the reader.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Step 6: Validate, then save or transform

After extraction, check that the result contains what you need. An empty list can mean the page changed, the content is generated by JavaScript, or your selector is wrong. Treat each of those as a different failure and log which one you suspect. Then write the output with the standard library, for example with json.dump for structured results or csv.writer for tables.

Troubleshooting: when the data is missing

  • The page loads but your fields are empty. Save the raw body to a file and search it for a value you can see in the browser. If the value is absent, the data is probably added by client-side JavaScript after the page loads. This method reads the static response, so it cannot see that data.
  • The text contains strange characters. The charset is wrong or was guessed. Compare the declared charset with any <meta> declaration and decode again.
  • You receive a 403 or 429 status. The server refused the request or limited its rate. Check the robots.txt rules and the site’s own access policy before retrying, and do not work around the refusal.
  • The script hangs. The timeout is missing or too long. Set one and handle the exception.

Limits to plan around

  • robots.txt is not a permission system. It tells you what the site’s rules say about a user agent. It does not decide whether the site’s terms, access controls, privacy expectations or applicable law allow the collection you have in mind.
  • Static responses change. The page structure, field names and CSV columns can change at any time. Keep your selectors narrow, validate every result, and expect to update the code.
  • Standard-library behaviour depends on the Python release. Check signatures and behaviour against the documentation for the version you run. The standard library index confirms which modules exist, but the details of each function are documented per release.

Within these limits, the workflow is complete. It needs no installation step beyond a Python interpreter, and every module it uses is part of the standard library.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.