Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To find links in an HTML document, parse it with BeautifulSoup, select every <a> element, and read each element’s href safely:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]

This returns the anchor href values exactly as written, including relative paths such as /about. If you need usable absolute URLs, resolve those values against the page URL with Python’s urllib.parse.urljoin. The basic recipe finds links represented by anchor tags; URLs in images, scripts, forms, canonical tags, or other attributes require separate searches.

Install BeautifulSoup and choose a parser

The package is distributed as beautifulsoup4. Install it in the Python environment that will run your script:

python -m pip install beautifulsoup4

BeautifulSoup can use Python’s built-in html.parser, or optional parsers such as lxml and html5lib. Different parsers can build different trees from malformed HTML. Name the parser explicitly when you need repeatable results on different machines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • html.parser: included with Python, so it needs no extra parser package.
  • lxml: listed first among the commonly supported choices when it is available.
  • html5lib: follows HTML5 parsing behavior more closely.

For a portable first script, use html.parser. If you change parsers, install the dependency and change the constructor at the same time.

Extract every anchor href from an HTML string

This complete example handles anchors with and without an href attribute:

from bs4 import BeautifulSoup

html = """
<nav>
  <a href="/about">About</a>
  <a href="team.html">Team</a>
  <a>This anchor has no href</a>
  <a href="https://example.com/news">News</a>
</nav>
"""

soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]

for href in links:
    print(href)

The output is:

/about
team.html
None
https://example.com/news

find_all("a") returns all matching anchor tags in document order. get("href") returns the attribute value or None when the attribute is missing. That makes it safer than a["href"] when the input may contain anchors used only for scripts, buttons, or incomplete markup.

Exclude anchors without href values

If your output should contain only actual attribute values, filter out missing values:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
links = [
    a.get("href")
    for a in soup.find_all("a")
    if a.get("href")
]

This also excludes an empty string. If an empty attribute has meaning in your application, test explicitly for is not None instead.

Keep link text with each URL

When auditing navigation or accessibility, retain the visible text alongside the URL:

records = []
for a in soup.find_all("a"):
    records.append({
        "href": a.get("href"),
        "text": a.get_text(" ", strip=True),
    })

for record in records:
    print(record["text"], "->", record["href"])

Fetch a page, then parse its HTML

Downloading a page and parsing HTML are separate operations. The parser cannot discover content that was never placed in the string you give it. This example uses Python’s standard-library HTTP client, then passes the response text to BeautifulSoup:

from urllib.request import Request, urlopen
from bs4 import BeautifulSoup

page_url = "https://example.com/"
request = Request(page_url, headers={"User-Agent": "link-audit/1.0"})

with urlopen(request, timeout=30) as response:
    html = response.read()

soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a") if a.get("href")]

for href in links:
    print(href)

In production, handle network errors, status codes, content encoding, redirects, and robots or access policies appropriate to the site you are retrieving. A successful HTTP response still may not be the page you expected, so inspect the returned HTML when the result is empty.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert relative href values to absolute URLs

Pages commonly use relative references such as /docs, guide/start.html, ../pricing, fragments such as #install, or a scheme-relative URL such as //cdn.example.com/file. Resolve them against the URL of the document:

from urllib.parse import urljoin
from bs4 import BeautifulSoup

page_url = "https://example.com/products/index.html"
html = """
<a href="/about">About</a>
<a href="details.html">Details</a>
<a href="../contact">Contact</a>
<a href="https://other.example/item">Other site</a>
"""

soup = BeautifulSoup(html, "html.parser")
absolute_links = [
    urljoin(page_url, href)
    for a in soup.find_all("a")
    if (href := a.get("href"))
]

for url in absolute_links:
    print(url)

urljoin combines the base page URL and each reference. An absolute or scheme-relative reference can supply a different host or scheme, so do not assume the result remains on the original site.

Restrict results to one host

If you are crawling only one site, parse the resulting URLs and apply an allow-list before following them:

from urllib.parse import urljoin, urlparse

allowed_host = urlparse(page_url).netloc
same_site = [
    url for url in absolute_links
    if urlparse(url).netloc == allowed_host
]

For security-sensitive workflows, validate schemes and hosts before making requests. Never treat an untrusted href as safe merely because it was joined to a trusted base: an absolute input can override the base host.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose exactly what counts as a link

The basic expression means “every anchor element.” Real pages contain several kinds of URL-bearing markup, each requiring its own selector and attribute.

Only links inside a section

main = soup.select_one("main")
section_links = [] if main is None else [
    a.get("href") for a in main.find_all("a") if a.get("href")
]

Filter by attributes or URL pattern

documentation_links = [
    a.get("href")
    for a in soup.find_all("a", href=True)
    if "/docs/" in a["href"]
]

Using href=True asks BeautifulSoup for anchors that have the attribute, while get remains useful when you want one extraction expression that tolerates missing attributes.

Find other URL-bearing elements

These are not anchor links, but they may matter to a crawler or asset inventory:

images = [img.get("src") for img in soup.find_all("img") if img.get("src")]
scripts = [script.get("src") for script in soup.find_all("script") if script.get("src")]
stylesheets = [
    link.get("href")
    for link in soup.find_all("link", rel="stylesheet")
    if link.get("href")
]
canonical = [
    link.get("href")
    for link in soup.find_all("link", rel="canonical")
    if link.get("href")
]

Forms usually carry a destination in action, while images use src or responsive-image attributes. Search each tag and attribute deliberately rather than calling every URL a hyperlink.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize, classify, and deduplicate results

Extraction and policy decisions should be separate. First collect values; then decide whether to keep fragments, email links, scripts, duplicates, or external hosts.

from urllib.parse import urljoin, urldefrag, urlparse

raw = [a.get("href") for a in soup.find_all("a") if a.get("href")]
absolute = [urljoin(page_url, href) for href in raw]

web_urls = []
for url in absolute:
    without_fragment, _fragment = urldefrag(url)
    parsed = urlparse(without_fragment)
    if parsed.scheme in {"http", "https"}:
        web_urls.append(without_fragment)

unique_urls = list(dict.fromkeys(web_urls))

This example removes fragment identifiers and preserves first-seen order while deduplicating. Do not remove fragments if they identify meaningful sections that your report must retain. Likewise, keep mailto:, tel:, or custom schemes when your application needs them; classify them rather than silently discarding them.

Why an apparently valid page returns no links

The input is not the page HTML

Print a short prefix, the response status, and the final URL after fetching. Login pages, bot checks, error documents, and consent interstitials can contain little or no useful navigation.

There are no anchor href attributes

An anchor can exist without href. Count both tags and values:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
anchors = soup.find_all("a")
print("anchors:", len(anchors))
print("hrefs:", sum(a.get("href") is not None for a in anchors))

The page builds links with JavaScript

A static response parse sees only the HTML delivered by the server. Links inserted after scripts run will not appear unless you use a browser-capable rendering step or obtain the generated data through another endpoint.

The parser shaped malformed markup differently

Try a specified parser and compare the relevant region. If malformed HTML is central to the input, install and name the parser whose behavior fits your requirement instead of relying on environment defaults.

Your selector is too narrow

Check the whole document with soup.find_all("a") before restricting the search to main, a class, or a CSS selector. Confirm that the selector matches the actual returned markup, not what you see after client-side rendering.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and reliability practices

  • Parse once and reuse the soup tree when producing several reports.
  • Stream or batch large input files rather than constructing unnecessary copies of the HTML.
  • Use a timeout for network requests and record the URL, status, and parser used for each document.
  • Resolve relative URLs only when you have the correct final page URL, especially after redirects.
  • Deduplicate after normalization, but preserve raw values if an audit must reproduce the source exactly.
  • Respect access controls and site policies; extracting a link does not grant permission to crawl its destination.

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than parsing its anchor markup, ScreenshotNeo provides a single HTTP call. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the capture options. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Frequently Asked Questions

Does BeautifulSoup crawl links automatically?

No. It parses the HTML you provide and extracts matching elements. Following discovered URLs requires a separate fetching and crawling design.

Should I use find_all(‘a’, href=True) or get(‘href’)?

Use href=True when selecting only anchors that have the attribute; use get(‘href’) when you want missing attributes handled without an exception.

Why are links I can see in a browser missing from the response?

They may be inserted by JavaScript after the initial HTML loads, or the fetched document may be a login, consent, bot-check, or error page instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.