You can extract email candidates from a permitted, mostly static page in Python by fetching its HTTP response, parsing the returned HTML, collecting mailto: links and visible text, then deduplicating and validating the results. The standard library is enough for a conservative one-page workflow: urllib.request retrieves the response, html.parser reads HTML, urllib.parse resolves links, and urllib.robotparser checks the site’s crawler instructions. This method sees only what the server sends; it will not reliably find addresses rendered later by JavaScript or deliberately obfuscated in the browser.
What the Python workflow actually does
Email scraping is a four-stage process, not a single regular expression:
- Retrieve: request one URL that you are authorized to access and inspect its status, content type and character encoding.
- Parse: process the HTML as a document so that links and text are handled separately from markup.
- Extract: collect
mailto:targets and plausible email-shaped strings from visible text as candidates. - Review: normalize, deduplicate and manually or programmatically validate candidates before storing or using them.
Python’s documentation covers urllib.request, urllib.parse and HTML parsing in the standard library; it also identifies Requests as a higher-level HTTP client alternative (Python documentation). Neither approach guarantees that an address exists in the response or that it is current.
Check permission, robots.txt and rate limits first
Use this technique only on pages you may access and for a purpose consistent with the site’s rules and applicable law. Before fetching the page, read its robots.txt. Python’s RobotFileParser can retrieve the file and answer whether a user agent may fetch a URL (robotparser documentation).
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
from urllib.robotparser import RobotFileParser
from urllib.parse import urlparse
url = "https://example.com/contact"
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
rp = RobotFileParser(robots_url)
rp.read()
user_agent = "itechguides-email-checker/1.0"
if not rp.can_fetch(user_agent, url):
raise PermissionError("robots.txt disallows this fetch")
Robots rules are crawler instructions, not authentication, access control or blanket legal permission. RFC 9309 standardizes the Robots Exclusion Protocol (RFC 9309). Also follow terms of service, authentication requirements, reasonable request rates and any explicit denial. A single request with a clear user agent and a delay between permitted requests is safer than an uncontrolled crawler.
Standard-library example: fetch one page and extract candidates
The following script intentionally handles one URL. It accepts only an HTTP success response, checks that the response resembles HTML, decodes using the server’s declared charset when available, captures mailto: links, and searches visible text. Matches are candidates, not verified mailboxes.
#!/usr/bin/env python3
import re
from html.parser import HTMLParser
from urllib.parse import unquote, urljoin, urlparse
from urllib.request import Request, urlopen
EMAIL_RE = re.compile(
r"(?i)\b[a-z0-9.!#$%&'*+/=?^_`{|}~-]+"
r"@[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?"
r"(?:\.[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?)+\b"
)
class EmailParser(HTMLParser):
def __init__(self, page_url):
super().__init__(convert_charrefs=True)
self.page_url = page_url
self.text_parts = []
self.mailto_candidates = []
self._skip_depth = 0
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag.lower() in {"script", "style", "noscript", "template"}:
self._skip_depth += 1
if tag.lower() == "a" and attrs.get("href", "").lower().startswith("mailto:"):
raw = attrs["href"][len("mailto:"):]
address = unquote(raw).split("?", 1)[0].strip()
self.mailto_candidates.append(address)
def handle_endtag(self, tag):
if tag.lower() in {"script", "style", "noscript", "template"} and self._skip_depth:
self._skip_depth -= 1
def handle_data(self, data):
if not self._skip_depth and data.strip():
self.text_parts.append(data)
def fetch_html(url):
request = Request(url, headers={"User-Agent": "ite chguides-email-checker/1.0"})
with urlopen(request, timeout=20) as response:
content_type = response.headers.get_content_type()
if content_type not in {"text/html", "application/xhtml+xml"}:
raise ValueError(f"Expected HTML, got {content_type}")
raw = response.read()
charset = response.headers.get_content_charset() or "utf-8"
return raw.decode(charset, errors="replace"), response.geturl()
def extract_candidates(url):
html, final_url = fetch_html(url)
parser = EmailParser(final_url)
parser.feed(html)
candidates = set(parser.mailto_candidates)
candidates.update(EMAIL_RE.findall(" ".join(parser.text_parts)))
normalized = {candidate.strip().lower() for candidate in candidates if candidate.strip()}
return final_url, sorted(normalized)
if __name__ == "__main__":
page = "https://example.com/contact"
final_url, emails = extract_candidates(page)
print(f"Fetched: {final_url}")
for email in emails:
print(email)
Change the sample URL to a page you are allowed to fetch. The parser ignores script, style, noscript and template contents because those regions commonly contain code or hidden templates rather than displayed contact text. It follows neither links nor redirects itself; urlopen reports the final URL so you can see where the request ended. The regular expression requires a dot-separated domain, which avoids many obvious false positives but also misses valid or unusual address forms.
Fix the sample user-agent typo before running
In the code above, use the literal value ite chguides-email-checker/1.0 only if you intentionally want that exact identifier. A normal identifier without a space is better:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall"User-Agent": "itechguides-email-checker/1.0"
Identify your client honestly; do not impersonate a browser or evade a site’s controls.
Why extraction returns false positives and misses addresses
A match proves only that text resembles an email address. It does not prove that the mailbox exists, belongs to the organization, accepts mail or is intended for solicitation.
- False positives: documentation examples, code samples, image filenames and text such as
name@example.comcan match. - False negatives: a site may write “name [at] example [dot] com,” split text across elements, expose an address only after a click, or load it with JavaScript.
- Non-HTML responses: a PDF, JSON API response or login page is not handled by this HTML parser.
- Encoding problems: an incorrect charset can turn visible text into replacement characters. The script uses the response declaration and replacement decoding to avoid crashing, but review unusual output.
- Stale content: a cached page can contain an address that has since been retired.
Keep provenance with every candidate: the source URL, retrieval time, extraction method (mailto or visible text) and any review status. Do not silently turn a candidate list into a mailing list.
Handling relative links and mailto parameters
A mailto: link can include a display name, subject or body query string. The example removes everything after the first question mark and URL-decodes the address portion. If you need to preserve the complete link for an authorized workflow, parse it explicitly and treat the query fields as untrusted data:
from urllib.parse import parse_qs, unquote, urlsplit
link = "mailto:person%40example.org?subject=Hello"
parts = urlsplit(link)
address = unquote(parts.path)
options = parse_qs(parts.query)
print(address, options.get("subject", [""])[0])
Do not automatically send mail, subscribe people, or disclose extracted query data. If a page contains ordinary links to a contact page, resolving them with urljoin(base_url, href) is technically possible, but each additional request needs its own permission and rate-limit decision.
Requests and an HTML parser: when a higher-level client helps
The standard library minimizes dependencies and gives precise control over requests, decoding and timeouts. Requests generally makes sessions, headers, redirects and error handling more convenient, but it does not make JavaScript execute and does not solve consent, authorization or legal questions. Pair either HTTP client with an HTML parser; do not parse HTML with a single giant regular expression.
Rank #3
| Consideration | urllib plus html.parser |
Requests plus an HTML parser |
|---|---|---|
| Dependencies | Included with Python | Additional packages to install and maintain |
| Control | Explicit request and response handling | Convenient sessions, headers and exceptions |
| HTML coverage | Parses returned HTML only | Also parses returned HTML only |
| JavaScript-rendered content | Not executed | Not executed |
| Best fit | Small, auditable scripts | Applications already using Requests |
Because current third-party parser versions and performance vary, choose based on your project’s dependency policy rather than assuming one library finds more addresses.
JavaScript-rendered pages: know when a browser is required
View source and the HTTP response are different from the DOM after JavaScript runs. If an address appears only after a script fetches an API, opens a modal or completes a challenge, the script above cannot see it. A browser-automation tool can render permitted content, but it adds a browser runtime, synchronization problems, greater resource use and additional responsibilities around consent banners and access controls.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →First inspect the response and browser network panel to determine whether the address is actually delivered by an authorized API. Prefer that documented API when available. Do not bypass CAPTCHAs, bot checks, authentication barriers or technical blocks. If the page is intentionally withholding contact data, respect that choice.
Validation, storage and data minimization
Useful validation is proportional to the purpose. At minimum, normalize case for comparison, remove surrounding whitespace and reject malformed domains. DNS or SMTP checks can be intrusive and are not proof of consent, so do not perform them by default. Keep only the fields you need, restrict access to stored results, set a retention period and delete candidates that are no longer necessary.
A public address is not blanket permission to collect, share or market to its owner. A UK-led joint regulator statement warns that scraping can affect personal information and identifies unwanted direct marketing or spam as a possible outcome (joint regulator statement). Its framing is not a universal rule for every country; review the law and site terms applicable to your location, the people involved and your intended use.
Commercial email and U.S. CAN-SPAM duties
If your purpose is commercial email to recipients in the United States, the FTC says CAN-SPAM applies to commercial messages, including business-to-business messages (FTC compliance guide). The guide describes duties including truthful sender and header information, non-deceptive subject lines, clear identification of advertising, a valid postal address, a working opt-out method, honoring opt-outs within 10 business days and oversight of vendors sending on your behalf. It also notes criminal prohibitions related to harvesting email addresses and dictionary attacks.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Extracting a visible address does not make a marketing campaign compliant. Rules outside the United States differ, and privacy, electronic-marketing, database, contract and computer-access laws can all matter. Obtain jurisdiction-specific advice before contacting people at scale.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
HTTP 403 or 429
Cause: the site denied the client or rate-limited requests. Fix: stop, verify permission and terms, slow down, and use an authorized API or contact the site owner. Do not rotate identities to evade a block.
“Expected HTML, got application/json”
Cause: the URL is an API endpoint, redirect target or error response. Fix: inspect the status, final URL and content type. Parse JSON with its documented schema instead of feeding it to the HTML parser.
No addresses found
Cause: the page may contain no address, use JavaScript, obfuscate text or place contact details in an image or PDF. Fix: compare the raw response with the rendered page, look for a documented contact endpoint and ask whether collection is permitted. Do not defeat an intentional obfuscation.
UnicodeDecodeError or garbled text
Cause: a missing or incorrect charset declaration. Fix: inspect the response headers and HTML metadata, then choose a verified encoding. The example’s errors="replace" keeps the process running but can hide characters that need manual review.
Best Value
Too many irrelevant matches
Cause: code samples, navigation and hidden content resemble addresses. Fix: keep the parser’s script/style exclusions, limit extraction to relevant containers you are authorized to process, retain the source context and review candidates.
The script hangs
Cause: a server is slow or never completes a response. Fix: retain a finite timeout, fetch one page at a time and record failures. A timeout is a failed fetch, not a reason to hammer the server.
Or skip the browser setup
If your real problem is inspecting a JavaScript-rendered page before deciding whether an address is displayed, ScreenshotNeo can return a clean screenshot through one request. It captures the rendered page; it does not turn pixels into a verified email list, so use it as a visual inspection aid rather than as permission to collect contact data.
Recommended Free Tools
Use the API documentation at screenshotneo.com/docs/ for the complete options. A minimal cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The equivalent Python call is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Operational checklist
- Confirm the URL, purpose and authorization.
- Read
robots.txtand the site’s terms. - Use a descriptive user agent and a finite timeout.
- Check status, content type, final URL and encoding.
- Parse HTML structurally; collect
mailto:links and visible-text candidates separately. - Deduplicate while preserving source URL and retrieval time.
- Review false positives and dynamic or obfuscated content.
- Minimize retention and secure any stored personal data.
- Obtain legal review before commercial outreach, especially across jurisdictions.
Frequently Asked Questions
Can Python scrape every email visible in a browser?
No. A basic HTTP parser sees the server response, not content inserted later by JavaScript, browser-only state, images or deliberate obfuscation.
Does robots.txt make email scraping legal?
No. It communicates crawler preferences. It is not authentication, access control or a universal legal permission; terms, authorization and applicable law still govern.
Free tools Windows power users keep installed
One-click scans. No signup required.
Is a regex match a valid email address?
It is only a candidate. Regex can both over-match examples and miss unusual or obfuscated addresses, and it cannot prove that a mailbox exists or that contacting its owner is permitted.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

