Free tools Windows power users keep installed
One-click scans. No signup required.
To find links in an HTML document, parse it with BeautifulSoup, select every <a> element, and read each element’s href safely:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]
This returns the anchor href values exactly as written, including relative paths such as /about. If you need usable absolute URLs, resolve those values against the page URL with Python’s urllib.parse.urljoin. The basic recipe finds links represented by anchor tags; URLs in images, scripts, forms, canonical tags, or other attributes require separate searches.
Install BeautifulSoup and choose a parser
The package is distributed as beautifulsoup4. Install it in the Python environment that will run your script:
python -m pip install beautifulsoup4
BeautifulSoup can use Python’s built-in html.parser, or optional parsers such as lxml and html5lib. Different parsers can build different trees from malformed HTML. Name the parser explicitly when you need repeatable results on different machines.
#1 Best Overall
html.parser: included with Python, so it needs no extra parser package.lxml: listed first among the commonly supported choices when it is available.html5lib: follows HTML5 parsing behavior more closely.
For a portable first script, use html.parser. If you change parsers, install the dependency and change the constructor at the same time.
Extract every anchor href from an HTML string
This complete example handles anchors with and without an href attribute:
from bs4 import BeautifulSoup
html = """
<nav>
<a href="/about">About</a>
<a href="team.html">Team</a>
<a>This anchor has no href</a>
<a href="https://example.com/news">News</a>
</nav>
"""
soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]
for href in links:
print(href)
The output is:
/about
team.html
None
https://example.com/news
find_all("a") returns all matching anchor tags in document order. get("href") returns the attribute value or None when the attribute is missing. That makes it safer than a["href"] when the input may contain anchors used only for scripts, buttons, or incomplete markup.
Exclude anchors without href values
If your output should contain only actual attribute values, filter out missing values:
links = [
a.get("href")
for a in soup.find_all("a")
if a.get("href")
]
This also excludes an empty string. If an empty attribute has meaning in your application, test explicitly for is not None instead.
Rank #2
Keep link text with each URL
When auditing navigation or accessibility, retain the visible text alongside the URL:
records = []
for a in soup.find_all("a"):
records.append({
"href": a.get("href"),
"text": a.get_text(" ", strip=True),
})
for record in records:
print(record["text"], "->", record["href"])
Fetch a page, then parse its HTML
Downloading a page and parsing HTML are separate operations. The parser cannot discover content that was never placed in the string you give it. This example uses Python’s standard-library HTTP client, then passes the response text to BeautifulSoup:
from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
page_url = "https://example.com/"
request = Request(page_url, headers={"User-Agent": "link-audit/1.0"})
with urlopen(request, timeout=30) as response:
html = response.read()
soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a") if a.get("href")]
for href in links:
print(href)
In production, handle network errors, status codes, content encoding, redirects, and robots or access policies appropriate to the site you are retrieving. A successful HTTP response still may not be the page you expected, so inspect the returned HTML when the result is empty.
Convert relative href values to absolute URLs
Pages commonly use relative references such as /docs, guide/start.html, ../pricing, fragments such as #install, or a scheme-relative URL such as //cdn.example.com/file. Resolve them against the URL of the document:
from urllib.parse import urljoin
from bs4 import BeautifulSoup
page_url = "https://example.com/products/index.html"
html = """
<a href="/about">About</a>
<a href="details.html">Details</a>
<a href="../contact">Contact</a>
<a href="https://other.example/item">Other site</a>
"""
soup = BeautifulSoup(html, "html.parser")
absolute_links = [
urljoin(page_url, href)
for a in soup.find_all("a")
if (href := a.get("href"))
]
for url in absolute_links:
print(url)
urljoin combines the base page URL and each reference. An absolute or scheme-relative reference can supply a different host or scheme, so do not assume the result remains on the original site.
Restrict results to one host
If you are crawling only one site, parse the resulting URLs and apply an allow-list before following them:
from urllib.parse import urljoin, urlparse
allowed_host = urlparse(page_url).netloc
same_site = [
url for url in absolute_links
if urlparse(url).netloc == allowed_host
]
For security-sensitive workflows, validate schemes and hosts before making requests. Never treat an untrusted href as safe merely because it was joined to a trusted base: an absolute input can override the base host.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Choose exactly what counts as a link
The basic expression means “every anchor element.” Real pages contain several kinds of URL-bearing markup, each requiring its own selector and attribute.
Only links inside a section
main = soup.select_one("main")
section_links = [] if main is None else [
a.get("href") for a in main.find_all("a") if a.get("href")
]
Filter by attributes or URL pattern
documentation_links = [
a.get("href")
for a in soup.find_all("a", href=True)
if "/docs/" in a["href"]
]
Using href=True asks BeautifulSoup for anchors that have the attribute, while get remains useful when you want one extraction expression that tolerates missing attributes.
Find other URL-bearing elements
These are not anchor links, but they may matter to a crawler or asset inventory:
images = [img.get("src") for img in soup.find_all("img") if img.get("src")]
scripts = [script.get("src") for script in soup.find_all("script") if script.get("src")]
stylesheets = [
link.get("href")
for link in soup.find_all("link", rel="stylesheet")
if link.get("href")
]
canonical = [
link.get("href")
for link in soup.find_all("link", rel="canonical")
if link.get("href")
]
Forms usually carry a destination in action, while images use src or responsive-image attributes. Search each tag and attribute deliberately rather than calling every URL a hyperlink.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Normalize, classify, and deduplicate results
Extraction and policy decisions should be separate. First collect values; then decide whether to keep fragments, email links, scripts, duplicates, or external hosts.
from urllib.parse import urljoin, urldefrag, urlparse
raw = [a.get("href") for a in soup.find_all("a") if a.get("href")]
absolute = [urljoin(page_url, href) for href in raw]
web_urls = []
for url in absolute:
without_fragment, _fragment = urldefrag(url)
parsed = urlparse(without_fragment)
if parsed.scheme in {"http", "https"}:
web_urls.append(without_fragment)
unique_urls = list(dict.fromkeys(web_urls))
This example removes fragment identifiers and preserves first-seen order while deduplicating. Do not remove fragments if they identify meaningful sections that your report must retain. Likewise, keep mailto:, tel:, or custom schemes when your application needs them; classify them rather than silently discarding them.
Why an apparently valid page returns no links
The input is not the page HTML
Print a short prefix, the response status, and the final URL after fetching. Login pages, bot checks, error documents, and consent interstitials can contain little or no useful navigation.
There are no anchor href attributes
An anchor can exist without href. Count both tags and values:
Recommended Free Tools
Best Value
anchors = soup.find_all("a")
print("anchors:", len(anchors))
print("hrefs:", sum(a.get("href") is not None for a in anchors))
The page builds links with JavaScript
A static response parse sees only the HTML delivered by the server. Links inserted after scripts run will not appear unless you use a browser-capable rendering step or obtain the generated data through another endpoint.
The parser shaped malformed markup differently
Try a specified parser and compare the relevant region. If malformed HTML is central to the input, install and name the parser whose behavior fits your requirement instead of relying on environment defaults.
Your selector is too narrow
Check the whole document with soup.find_all("a") before restricting the search to main, a class, or a CSS selector. Confirm that the selector matches the actual returned markup, not what you see after client-side rendering.
Performance and reliability practices
- Parse once and reuse the soup tree when producing several reports.
- Stream or batch large input files rather than constructing unnecessary copies of the HTML.
- Use a timeout for network requests and record the URL, status, and parser used for each document.
- Resolve relative URLs only when you have the correct final page URL, especially after redirects.
- Deduplicate after normalization, but preserve raw values if an audit must reproduce the source exactly.
- Respect access controls and site policies; extracting a link does not grant permission to crawl its destination.
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than parsing its anchor markup, ScreenshotNeo provides a single HTTP call. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallcURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for the capture options. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently Asked Questions
Does BeautifulSoup crawl links automatically?
No. It parses the HTML you provide and extracts matching elements. Following discovered URLs requires a separate fetching and crawling design.
Should I use find_all(‘a’, href=True) or get(‘href’)?
Use href=True when selecting only anchors that have the attribute; use get(‘href’) when you want missing attributes handled without an exception.
Why are links I can see in a browser missing from the response?
They may be inserted by JavaScript after the initial HTML loads, or the fetched document may be a login, consent, bot-check, or error page instead.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

