What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Website metadata is distributed across the HTML <head>, response headers, and structured-data blocks. To extract it reliably, record the URL and response details, parse the original HTML, collect standard and social fields without overwriting duplicates, inspect JSON-LD/Microdata/RDFa separately, and render the page when JavaScript changes the live DOM. The result describes what the server or browser exposes—not necessarily the title or description a search engine will display.
What counts as website metadata?
The HTML document’s <head> is the primary place for page metadata. It commonly contains:
- Document title: the
<title>element. - Name/content meta elements: description, robots directives, viewport settings, and other declarations.
- Social metadata: Open Graph properties such as
og:title,og:description, andog:image, plus Twitter/X card fields when present. - Link relations: canonical, alternate-language, author, and other machine-readable relationships.
- Structured data: JSON-LD scripts, Microdata, and RDFa. These describe entities and properties rather than acting as ordinary name/value meta tags.
HTTP response headers are a separate layer. For example, X-Robots-Tag can apply crawler directives to PDFs, images, and other non-HTML resources.
Do not treat meta keywords as a dependable SEO field; major search engines generally ignore it. Also keep the original location and spelling of each value because duplicate or conflicting declarations can reveal implementation problems.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Manual extraction in a browser
Inspect the original response source
- Open the target page in a browser.
- Use View Page Source (often available by right-clicking the page or with
view-source:before the URL). - Search for
<title>,name="description",name="robots", andproperty="og:. - Search for
application/ld+jsonto locate JSON-LD blocks and forrel="canonical"to locate the canonical link. - Record the exact text, attributes, and order. If a field appears more than once, keep every occurrence.
Source inspection tells you what the initial HTTP response contained. It may not match what a visitor sees after scripts run.
Inspect the live DOM
Open developer tools with Inspect or Ctrl+Shift+I/Cmd+Option+I, choose the Elements panel, expand <head>, and repeat the searches. This DOM can include metadata inserted or changed by JavaScript. Compare it with the source when you need to know whether a value was server-rendered or added later.
Check headers and request context
In developer tools, open Network, reload the page, select the document request, and inspect Headers. Record the final URL after redirects, status code, content type, encoding, and relevant directives such as X-Robots-Tag. A robots directive cannot guide a crawler that was unable to access the response in the first place.
A reliable extraction workflow
- Define the target. Decide whether you need server HTML, post-JavaScript DOM, HTTP headers, or all three.
- Fetch and document context. Save the requested URL, fetch time, final URL, status, content type, and response headers.
- Parse without flattening. Collect the title, meta elements, links, social fields, and structured-data blocks into separate records.
- Normalize carefully. Trim surrounding whitespace and decode HTML entities, but retain raw values and source locations. Never silently invent a missing value.
- Detect duplicates and conflicts. Return arrays for repeated names or properties rather than allowing the last value to overwrite earlier declarations.
- Render when necessary. Use a browser renderer if scripts inject or modify metadata, then compare rendered values with response-source values.
- Interpret, do not overclaim. An extracted title or description is an input to search and social systems, not a guarantee of the displayed result.
Python extractor for HTML metadata
The following script fetches a URL, follows redirects, records context, and emits separate collections for ordinary tags, links, Open Graph/Twitter fields, JSON-LD, and headers. It preserves duplicates. Install the two dependencies with python -m pip install requests beautifulsoup4.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
import json
import sys
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
url = sys.argv[1]
response = requests.get(
url,
headers={"User-Agent": "metadata-inspector/1.0"},
timeout=30,
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
result = {
"requested_url": url,
"final_url": response.url,
"fetched_at": datetime.now(timezone.utc).isoformat(),
"status": response.status_code,
"content_type": response.headers.get("Content-Type"),
"headers": dict(response.headers),
"title": soup.title.get_text(" ", strip=True) if soup.title else None,
"meta": [],
"links": [],
"json_ld": [],
}
for tag in soup.find_all("meta"):
item = {"attributes": dict(tag.attrs), "content": tag.get("content")}
result["meta"].append(item)
for tag in soup.find_all("link"):
href = tag.get("href")
result["links"].append({
"attributes": dict(tag.attrs),
"absolute_href": urljoin(response.url, href) if href else None,
})
for tag in soup.find_all("script", attrs={"type": "application/ld+json"}):
raw = tag.string or tag.get_text()
try:
result["json_ld"].append({"raw": raw, "parsed": json.loads(raw)})
except json.JSONDecodeError:
result["json_ld"].append({"raw": raw, "parsed": None, "parse_error": True})
print(json.dumps(result, indent=2, ensure_ascii=False))
This is a response-source extractor. It does not execute JavaScript, so a page that creates its title, description, or JSON-LD only after load requires a rendering step and a second pass over the resulting DOM.
What to extract and how to interpret it
Title and description
The <title> is a page declaration. A description meta element is another input that Google may use for a snippet, but Google can choose visible page text instead. Google’s title link is generated automatically from multiple signals and may differ from the title element.
Robots directives
Read both HTML meta name="robots" values and the X-Robots-Tag response header. The header is especially important for files that are not HTML. Treat these as crawler instructions, not evidence that a crawler has already followed them, and verify that the resource was accessible.
Open Graph and Twitter/X cards
Store each property/name and its content exactly. Common Open Graph values include og:title, og:description, and og:image. Multiple images or locale variants are valid, so an array is safer than a single scalar field.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Structured data
Keep JSON-LD, Microdata, and RDFa distinct from ordinary meta tags. JSON-LD may contain nested entities, arrays, and an @graph; do not flatten it into unrelated name/content pairs. Google supports all three formats and generally recommends JSON-LD when it is practical to implement and maintain. Valid markup alone does not guarantee a rich result; eligibility depends on the documentation for the particular search feature.
Canonical and other links
Capture the relation, destination, and any alternate-language attributes. Resolve relative URLs against the final response URL while retaining the original href.
Response source versus rendered DOM
| Question | Use | Limitation |
|---|---|---|
| What did the server initially return? | HTTP fetch and view-source | Misses metadata inserted by scripts |
| What exists after scripts run? | Browser developer tools or a renderer | Timing, consent dialogs, and failed scripts can change the result |
| What crawler directive applies to a PDF or image? | HTTP response headers | Not represented by the HTML document’s head |
| What might appear in search? | Extracted declarations plus the page and site context | Search engines select and rewrite titles and snippets |
When auditing a site, save both snapshots and label them clearly. A difference is useful evidence: it can show that server-side rendering is incomplete, a script is failing, or a framework is replacing metadata during navigation.
Encoding, malformed markup, and edge cases
- Prefer UTF-8. An HTML5 character-encoding declaration should be UTF-8 and appear entirely within the first 1,024 bytes.
- Handle redirects by recording both requested and final URLs; canonical values should be evaluated against the final document.
- Reject or flag non-HTML content instead of parsing a PDF error page as HTML.
- Expect missing, empty, duplicated, malformed, or contradictory fields. Report them explicitly.
- Do not assume a missing description, image, canonical, or structured-data block has a specific SEO consequence.
- Watch for access controls, bot checks, rate limits, consent walls, compressed responses, and incorrect character decoding.
- Do not treat invalid elements in
<head>as harmless. Invalid markup can interfere with how metadata is processed, and later elements may be ignored.
Common failures and fixes
The script returns a 403 or CAPTCHA page
Check the status and content type before parsing. Use an authorized, rate-limited request where permitted, or inspect the page in a normal browser. Never label a challenge page as the target’s metadata.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
The title is missing in fetched HTML
Compare response source with the live DOM. If the DOM has a title, metadata is likely client-generated; use a renderer and wait for the relevant selector or page state.
JSON-LD will not parse
Preserve the raw block, report the parse error, and inspect trailing commas, HTML entities, multiple top-level objects, or an unexpected script type. Do not discard the block silently.
Relative image or canonical URLs look broken
Resolve them against the final response URL, not the originally requested URL, while retaining the raw attribute for auditability.
Google shows different text
This is expected behavior. Compare the extracted declarations with visible page copy and remember that search systems generate title links and snippets from several signals.
Recommended Free Tools
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Or skip the browser setup
ScreenshotNeo can capture a rendered page while you inspect its visible result. Its API accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed, while bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Every response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
Use the one-call request below; the ScreenshotNeo documentation lists the options for full-page capture, waiting, custom headers, cookies, JavaScript, CSS, device settings, PDFs, caching, async jobs, and more.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Cost, performance, and reliability considerations
- For one page, view-source is fastest and costs nothing. For many URLs, reuse connections, set timeouts, limit concurrency, and cache results with a recorded fetch time.
- Separate fetch failures from parse failures. A valid HTTP response can still contain an error page, partial HTML, or a challenge.
- Rendering is slower and less deterministic than parsing source. Wait for a meaningful selector or network-idle condition, and record the browser state and timing.
- Keep raw HTML, headers, and parser output when audits must be reproducible. Redact credentials and personal data from stored headers.
- Respect site terms, robots policies, authentication boundaries, and rate limits. Extraction should not bypass access controls.
Metadata audit checklist
- Requested URL, final URL, timestamp, status, content type, and encoding recorded.
- Title, standard meta fields, canonical/alternate links, Open Graph, and Twitter/X fields captured.
- JSON-LD, Microdata, and RDFa identified separately.
X-Robots-Tagand HTML robots directives checked.- Duplicates, missing values, malformed markup, redirects, and encoding issues reported.
- Source HTML compared with rendered DOM when scripts can alter metadata.
- Search-display conclusions qualified as possibilities rather than guarantees.
Frequently Asked Questions
Can website metadata be extracted from a URL without downloading the page?
No. The extractor must receive the HTML response or a rendered document; a URL alone does not contain the page’s current metadata.
Should duplicate meta descriptions be merged?
No. Preserve each occurrence and its location, then flag the duplication for review. Merging can hide which declaration a consumer used.
Is structured data required for a page title or description?
No. Titles and descriptions are ordinary HTML metadata. Structured data is a separate representation for entities and properties.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

