Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can extract email candidates from a permitted, mostly static page in Python by fetching its HTTP response, parsing the returned HTML, collecting mailto: links and visible text, then deduplicating and validating the results. The standard library is enough for a conservative one-page workflow: urllib.request retrieves the response, html.parser reads HTML, urllib.parse resolves links, and urllib.robotparser checks the site’s crawler instructions. This method sees only what the server sends; it will not reliably find addresses rendered later by JavaScript or deliberately obfuscated in the browser.

What the Python workflow actually does

Email scraping is a four-stage process, not a single regular expression:

  1. Retrieve: request one URL that you are authorized to access and inspect its status, content type and character encoding.
  2. Parse: process the HTML as a document so that links and text are handled separately from markup.
  3. Extract: collect mailto: targets and plausible email-shaped strings from visible text as candidates.
  4. Review: normalize, deduplicate and manually or programmatically validate candidates before storing or using them.

Python’s documentation covers urllib.request, urllib.parse and HTML parsing in the standard library; it also identifies Requests as a higher-level HTTP client alternative (Python documentation). Neither approach guarantees that an address exists in the response or that it is current.

Check permission, robots.txt and rate limits first

Use this technique only on pages you may access and for a purpose consistent with the site’s rules and applicable law. Before fetching the page, read its robots.txt. Python’s RobotFileParser can retrieve the file and answer whether a user agent may fetch a URL (robotparser documentation).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.robotparser import RobotFileParser
from urllib.parse import urlparse

url = "https://example.com/contact"
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"

rp = RobotFileParser(robots_url)
rp.read()
user_agent = "itechguides-email-checker/1.0"
if not rp.can_fetch(user_agent, url):
    raise PermissionError("robots.txt disallows this fetch")

Robots rules are crawler instructions, not authentication, access control or blanket legal permission. RFC 9309 standardizes the Robots Exclusion Protocol (RFC 9309). Also follow terms of service, authentication requirements, reasonable request rates and any explicit denial. A single request with a clear user agent and a delay between permitted requests is safer than an uncontrolled crawler.

Standard-library example: fetch one page and extract candidates

The following script intentionally handles one URL. It accepts only an HTTP success response, checks that the response resembles HTML, decodes using the server’s declared charset when available, captures mailto: links, and searches visible text. Matches are candidates, not verified mailboxes.

#!/usr/bin/env python3
import re
from html.parser import HTMLParser
from urllib.parse import unquote, urljoin, urlparse
from urllib.request import Request, urlopen

EMAIL_RE = re.compile(
    r"(?i)\b[a-z0-9.!#$%&'*+/=?^_`{|}~-]+"
    r"@[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?"
    r"(?:\.[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?)+\b"
)

class EmailParser(HTMLParser):
    def __init__(self, page_url):
        super().__init__(convert_charrefs=True)
        self.page_url = page_url
        self.text_parts = []
        self.mailto_candidates = []
        self._skip_depth = 0

    def handle_starttag(self, tag, attrs):
        attrs = dict(attrs)
        if tag.lower() in {"script", "style", "noscript", "template"}:
            self._skip_depth += 1
        if tag.lower() == "a" and attrs.get("href", "").lower().startswith("mailto:"):
            raw = attrs["href"][len("mailto:"):]
            address = unquote(raw).split("?", 1)[0].strip()
            self.mailto_candidates.append(address)

    def handle_endtag(self, tag):
        if tag.lower() in {"script", "style", "noscript", "template"} and self._skip_depth:
            self._skip_depth -= 1

    def handle_data(self, data):
        if not self._skip_depth and data.strip():
            self.text_parts.append(data)

def fetch_html(url):
    request = Request(url, headers={"User-Agent": "ite chguides-email-checker/1.0"})
    with urlopen(request, timeout=20) as response:
        content_type = response.headers.get_content_type()
        if content_type not in {"text/html", "application/xhtml+xml"}:
            raise ValueError(f"Expected HTML, got {content_type}")
        raw = response.read()
        charset = response.headers.get_content_charset() or "utf-8"
        return raw.decode(charset, errors="replace"), response.geturl()

def extract_candidates(url):
    html, final_url = fetch_html(url)
    parser = EmailParser(final_url)
    parser.feed(html)
    candidates = set(parser.mailto_candidates)
    candidates.update(EMAIL_RE.findall(" ".join(parser.text_parts)))
    normalized = {candidate.strip().lower() for candidate in candidates if candidate.strip()}
    return final_url, sorted(normalized)

if __name__ == "__main__":
    page = "https://example.com/contact"
    final_url, emails = extract_candidates(page)
    print(f"Fetched: {final_url}")
    for email in emails:
        print(email)

Change the sample URL to a page you are allowed to fetch. The parser ignores script, style, noscript and template contents because those regions commonly contain code or hidden templates rather than displayed contact text. It follows neither links nor redirects itself; urlopen reports the final URL so you can see where the request ended. The regular expression requires a dot-separated domain, which avoids many obvious false positives but also misses valid or unusual address forms.

Fix the sample user-agent typo before running

In the code above, use the literal value ite chguides-email-checker/1.0 only if you intentionally want that exact identifier. A normal identifier without a space is better:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
"User-Agent": "itechguides-email-checker/1.0"

Identify your client honestly; do not impersonate a browser or evade a site’s controls.

Why extraction returns false positives and misses addresses

A match proves only that text resembles an email address. It does not prove that the mailbox exists, belongs to the organization, accepts mail or is intended for solicitation.

  • False positives: documentation examples, code samples, image filenames and text such as name@example.com can match.
  • False negatives: a site may write “name [at] example [dot] com,” split text across elements, expose an address only after a click, or load it with JavaScript.
  • Non-HTML responses: a PDF, JSON API response or login page is not handled by this HTML parser.
  • Encoding problems: an incorrect charset can turn visible text into replacement characters. The script uses the response declaration and replacement decoding to avoid crashing, but review unusual output.
  • Stale content: a cached page can contain an address that has since been retired.

Keep provenance with every candidate: the source URL, retrieval time, extraction method (mailto or visible text) and any review status. Do not silently turn a candidate list into a mailing list.

Handling relative links and mailto parameters

A mailto: link can include a display name, subject or body query string. The example removes everything after the first question mark and URL-decodes the address portion. If you need to preserve the complete link for an authorized workflow, parse it explicitly and treat the query fields as untrusted data:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import parse_qs, unquote, urlsplit

link = "mailto:person%40example.org?subject=Hello"
parts = urlsplit(link)
address = unquote(parts.path)
options = parse_qs(parts.query)
print(address, options.get("subject", [""])[0])

Do not automatically send mail, subscribe people, or disclose extracted query data. If a page contains ordinary links to a contact page, resolving them with urljoin(base_url, href) is technically possible, but each additional request needs its own permission and rate-limit decision.

Requests and an HTML parser: when a higher-level client helps

The standard library minimizes dependencies and gives precise control over requests, decoding and timeouts. Requests generally makes sessions, headers, redirects and error handling more convenient, but it does not make JavaScript execute and does not solve consent, authorization or legal questions. Pair either HTTP client with an HTML parser; do not parse HTML with a single giant regular expression.

Consideration urllib plus html.parser Requests plus an HTML parser
Dependencies Included with Python Additional packages to install and maintain
Control Explicit request and response handling Convenient sessions, headers and exceptions
HTML coverage Parses returned HTML only Also parses returned HTML only
JavaScript-rendered content Not executed Not executed
Best fit Small, auditable scripts Applications already using Requests

Because current third-party parser versions and performance vary, choose based on your project’s dependency policy rather than assuming one library finds more addresses.

JavaScript-rendered pages: know when a browser is required

View source and the HTTP response are different from the DOM after JavaScript runs. If an address appears only after a script fetches an API, opens a modal or completes a challenge, the script above cannot see it. A browser-automation tool can render permitted content, but it adds a browser runtime, synchronization problems, greater resource use and additional responsibilities around consent banners and access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

First inspect the response and browser network panel to determine whether the address is actually delivered by an authorized API. Prefer that documented API when available. Do not bypass CAPTCHAs, bot checks, authentication barriers or technical blocks. If the page is intentionally withholding contact data, respect that choice.

Validation, storage and data minimization

Useful validation is proportional to the purpose. At minimum, normalize case for comparison, remove surrounding whitespace and reject malformed domains. DNS or SMTP checks can be intrusive and are not proof of consent, so do not perform them by default. Keep only the fields you need, restrict access to stored results, set a retention period and delete candidates that are no longer necessary.

A public address is not blanket permission to collect, share or market to its owner. A UK-led joint regulator statement warns that scraping can affect personal information and identifies unwanted direct marketing or spam as a possible outcome (joint regulator statement). Its framing is not a universal rule for every country; review the law and site terms applicable to your location, the people involved and your intended use.

Commercial email and U.S. CAN-SPAM duties

If your purpose is commercial email to recipients in the United States, the FTC says CAN-SPAM applies to commercial messages, including business-to-business messages (FTC compliance guide). The guide describes duties including truthful sender and header information, non-deceptive subject lines, clear identification of advertising, a valid postal address, a working opt-out method, honoring opt-outs within 10 business days and oversight of vendors sending on your behalf. It also notes criminal prohibitions related to harvesting email addresses and dictionary attacks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extracting a visible address does not make a marketing campaign compliant. Rules outside the United States differ, and privacy, electronic-marketing, database, contract and computer-access laws can all matter. Obtain jurisdiction-specific advice before contacting people at scale.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

HTTP 403 or 429

Cause: the site denied the client or rate-limited requests. Fix: stop, verify permission and terms, slow down, and use an authorized API or contact the site owner. Do not rotate identities to evade a block.

“Expected HTML, got application/json”

Cause: the URL is an API endpoint, redirect target or error response. Fix: inspect the status, final URL and content type. Parse JSON with its documented schema instead of feeding it to the HTML parser.

No addresses found

Cause: the page may contain no address, use JavaScript, obfuscate text or place contact details in an image or PDF. Fix: compare the raw response with the rendered page, look for a documented contact endpoint and ask whether collection is permitted. Do not defeat an intentional obfuscation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

UnicodeDecodeError or garbled text

Cause: a missing or incorrect charset declaration. Fix: inspect the response headers and HTML metadata, then choose a verified encoding. The example’s errors="replace" keeps the process running but can hide characters that need manual review.

Too many irrelevant matches

Cause: code samples, navigation and hidden content resemble addresses. Fix: keep the parser’s script/style exclusions, limit extraction to relevant containers you are authorized to process, retain the source context and review candidates.

The script hangs

Cause: a server is slow or never completes a response. Fix: retain a finite timeout, fetch one page at a time and record failures. A timeout is a failed fetch, not a reason to hammer the server.

Or skip the browser setup

If your real problem is inspecting a JavaScript-rendered page before deciding whether an address is displayed, ScreenshotNeo can return a clean screenshot through one request. It captures the rendered page; it does not turn pixels into a verified email list, so use it as a visual inspection aid rather than as permission to collect contact data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the API documentation at screenshotneo.com/docs/ for the complete options. A minimal cURL request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The equivalent Python call is:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Operational checklist

  • Confirm the URL, purpose and authorization.
  • Read robots.txt and the site’s terms.
  • Use a descriptive user agent and a finite timeout.
  • Check status, content type, final URL and encoding.
  • Parse HTML structurally; collect mailto: links and visible-text candidates separately.
  • Deduplicate while preserving source URL and retrieval time.
  • Review false positives and dynamic or obfuscated content.
  • Minimize retention and secure any stored personal data.
  • Obtain legal review before commercial outreach, especially across jurisdictions.

Frequently Asked Questions

Can Python scrape every email visible in a browser?

No. A basic HTTP parser sees the server response, not content inserted later by JavaScript, browser-only state, images or deliberate obfuscation.

Does robots.txt make email scraping legal?

No. It communicates crawler preferences. It is not authentication, access control or a universal legal permission; terms, authorization and applicable law still govern.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a regex match a valid email address?

It is only a candidate. Regex can both over-match examples and miss unusual or obfuscated addresses, and it cannot prove that a mailbox exists or that contacting its owner is permitted.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.