Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Direct answer: fetch the page with an HTTP client, keep the final URL and response details, parse the returned HTML with a standards-aware parser, then query title, meta, and link-related elements. Resolve relative URLs against the document base or final response URL. If the data is added only after JavaScript runs, a plain HTTP request cannot see it; use an authorized browser-rendered document or an official data interface instead.

What the workflow actually extracts

“HTML,” “metadata,” and “links” are related but different outputs. The response body is the markup delivered by the server (or by a rendering step). The <title> element supplies the document title. <meta> elements carry named, property-based, pragma, or character-encoding metadata. Links are relationships represented by elements such as a, area, form, and link; the right selector depends on whether you need navigation targets, stylesheets, canonical URLs, forms, or every relationship.

Keep these values with every extraction:

  • Requested URL and final response URL after redirects.
  • HTTP status, headers, content type, and response body.
  • The document’s base URL, if present, and the URL used for resolution.
  • Original attribute text as well as a normalized absolute URL when exact source markup matters.

Check the content type before parsing. A successful status code does not make a response HTML; an error page, JSON document, PDF, or login redirect may be returned with status 200.

A reliable extraction procedure

  1. Fetch. Send an appropriate request, follow redirects according to your policy, and retain status, headers, body, and final URL.
  2. Verify. Confirm that the response is HTML (and handle missing or misleading content-type headers defensively).
  3. Parse. Give the body to an HTML parser rather than using regular expressions. Choose a parser that matches your malformed-markup tolerance and dependency constraints.
  4. Extract. Read title text, metadata attributes, and the specific link elements relevant to your task. Treat absent fields as absent, not as errors.
  5. Resolve. Convert relative references with the document’s base element when one exists; otherwise use the final response URL. Preserve the raw value too.
  6. Validate. Test redirects, missing metadata, duplicate links, malformed markup, empty attributes, non-HTML responses, and pages whose visible content is rendered later.

Python: fetch and parse HTML, metadata, and links

This example uses Requests and Beautiful Soup. Install them with python -m pip install requests beautifulsoup4.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup


def extract_page(url: str) -> dict:
    response = requests.get(
        url,
        headers={"User-Agent": "MetadataExtractor/1.0"},
        timeout=30,
        allow_redirects=True,
    )
    content_type = response.headers.get("content-type", "").lower()
    if "html" not in content_type and "xhtml" not in content_type:
        raise ValueError(f"Expected HTML, got {content_type or 'unknown content type'}")

    soup = BeautifulSoup(response.text, "html.parser")
    base_tag = soup.find("base", href=True)
    resolution_base = urljoin(response.url, base_tag["href"]) if base_tag else response.url

    title_tag = soup.find("title")
    title = title_tag.get_text(" ", strip=True) if title_tag else None

    metadata = []
    for tag in soup.find_all("meta"):
        key = tag.get("name") or tag.get("property") or tag.get("http-equiv") or "charset"
        value = tag.get("content") if key != "charset" else tag.get("charset")
        if value is not None:
            metadata.append({"key": key, "value": value})

    links = []
    for tag in soup.find_all(["a", "area", "form", "link"]):
        attribute = "action" if tag.name == "form" else "href"
        raw = tag.get(attribute)
        if raw is None:
            continue
        links.append({
            "element": tag.name,
            "raw": raw,
            "absolute": urljoin(resolution_base, raw),
            "text": tag.get_text(" ", strip=True) if tag.name != "link" else None,
        })

    return {
        "requested_url": url,
        "final_url": response.url,
        "status": response.status_code,
        "title": title,
        "metadata": metadata,
        "links": links,
    }


if __name__ == "__main__":
    import json
    print(json.dumps(extract_page("https://example.com"), indent=2))

The code keeps duplicate links because duplicates can be meaningful in source analysis. Deduplicate later by normalized URL only when your downstream task calls for it. Fragments (#section), query strings, trailing slashes, and URL case rules should not be discarded without an explicit policy.

Choosing a Beautiful Soup parser

Beautiful Soup can use Python’s built-in parser, lxml, or html5lib. They differ in dependency requirements, error recovery, and behavior on broken markup. There is no universal current speed winner for every workload: select for correctness and deployment constraints, then measure on representative pages.

Metadata details that prevent common mistakes

Title is not a meta tag

Read title separately. Do not assume an absent title can be replaced by og:title or another social field; those values serve different consumers.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Preserve metadata keys

Useful keys may appear as name, property, http-equiv, or charset. Keep the key and value, and allow multiple entries. A page may contain both standard description metadata and Open Graph or other property-based fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Head is the primary location, not an absolute boundary

Page metadata is normally in head, which can also contain link, script, style, base, noscript, and template. Real-world malformed documents can move or repair nodes during parsing, so validate against the parser’s tree rather than assuming perfect source formatting.

Collecting links deliberately

Selecting only visible anchors misses relationships. Use:

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
  • a[href] and area[href] for navigational targets.
  • form[action] for submission endpoints.
  • link[href] for canonical, alternate, stylesheet, icon, preload, and other resource relationships; inspect rel as well as href.

Resolve references after considering the document’s base element. Keep non-HTTP schemes such as mailto:, tel:, and javascript: identifiable rather than silently treating them as web pages. Ignore unrelated URL-bearing attributes unless your specification explicitly includes them.

When static fetching is not enough

A parser sees only the markup supplied to it. If a product list, article body, or metadata is inserted by client-side JavaScript, it will be absent from a normal response. First check whether the site exposes an authorized data interface. Otherwise, obtain a rendered document with a browser-capable workflow and run the same extraction against that resulting DOM. Rendering can fail because of authentication, bot checks, timing, cross-origin behavior, or scripts that depend on user interaction; do not assume that a rendering library succeeds on every site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to diagnose a rendering mismatch

  1. Save the raw response and inspect whether the missing text exists there.
  2. Compare the raw response URL with the browser’s final URL and frame context.
  3. Wait for a meaningful selector or network-idle condition rather than an arbitrary short delay.
  4. Check whether consent dialogs, login gates, or bot challenges prevent the application from reaching its normal state.
  5. Extract from the rendered DOM only after the required content is present.

Performance, reliability, and responsible access

  • Reuse HTTP connections and impose timeouts; never let one origin block an entire batch indefinitely.
  • Cache responses when freshness permits, and record cache age with the result.
  • Limit concurrency per origin, honor applicable access rules and terms, and identify your client honestly.
  • Retry transient network failures with bounded exponential backoff, not permanent 4xx responses or repeated bot challenges.
  • Record parser errors, status changes, and extraction counts so a template change is visible.

Crawler directives describe how a crawler discovers and processes instructions; they do not by themselves settle permission to copy, republish, or automate a site. Treat access, licensing, privacy, and applicable law as separate questions.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Its capture request can produce a PNG, JPEG, WebP, or PDF, while the extraction workflow above remains the right choice when you need the actual HTML tree and attributes.

For a visual, rendered check without managing a browser, call the API (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting checklist

Symptom Likely cause Fix
“Expected HTML” error Redirect to login, JSON, PDF, or another non-HTML response Inspect final URL, status, and content type; handle that response explicitly.
Relative links become wrong domains Resolution used the requested URL instead of final URL or ignored base Resolve against the document base, falling back to final response URL.
Title or description is empty Element is absent, empty, or added by JavaScript Return null/empty intentionally; inspect rendered output only when required.
Links are missing Selector covers only a, or URLs are in forms/link elements Include a, area, form, and link according to your goal.
Parser output differs from source Malformed HTML was repaired Compare parser choices and retain raw HTML when exact source fidelity matters.
Rendered page never settles Consent, authentication, bot protection, or script failure Use an authorized session, wait for a meaningful condition, or use an official interface.

FAQ

Should I use regular expressions for HTML?

No. HTML nesting, entities, malformed markup, and attributes require a parser that builds a document tree.

Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Should duplicate URLs be removed?

Only if the consumer wants a set. Keep duplicates for auditing, placement analysis, or preserving source order.

Can robots.txt tell me whether extraction is legal?

No. It describes crawler access behavior. Permission, terms, copyright, privacy, and other obligations require separate evaluation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.