Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsTo extract metadata from a website, fetch its HTML, inspect the page’s <head>, and parse each metadata layer separately: title and standard meta tags, link relations such as canonical and language alternates, social tags, robots directives, and JSON-LD structured data. If the initial HTML does not contain information visible in the browser, inspect the rendered DOM or use a JavaScript-capable renderer.
What website metadata includes
Metadata is not one field or one format. Different tags serve search engines, social networks, browsers, and software that describes a page’s entities. Extracting one layer does not mean the others are present or correct.
| Layer | What to look for | What it tells you |
|---|---|---|
| Core HTML metadata | <title>, <meta name="description">, charset and viewport |
Page title, summary, and document or display information |
| Link relations | <link rel="canonical">, language alternates |
Preferred page URL and related localized versions |
| Social metadata | Open Graph properties and Twitter Card tags | Titles, descriptions, images, and URLs used for social previews |
| Robots directives | meta name="robots", Googlebot-specific tags, and X-Robots-Tag |
Crawl, indexing, or search-result presentation instructions |
| Structured data | <script type="application/ld+json"> |
Machine-readable entities and their relationships, often using Schema.org vocabulary |
Google describes meta tags as HTML tags that provide additional information to search engines and other clients. They are not interchangeable with structured data: a description tag is not a JSON-LD entity, and a robots directive is not descriptive metadata.
How to extract metadata from one URL
1. Fetch the page and preserve the response
Record the URL you requested, the final URL after redirects, the HTTP status, content type, retrieval time, and the raw HTML. Preserving the response makes it possible to distinguish a site change from a parsing error later. For a server-rendered page, the initial response is often all you need.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
For a quick command-line check, use curl -L to follow redirects and save the response:
curl -L -D headers.txt "https://example.com/" -o page.html
Review headers.txt for status and response headers, including any X-Robots-Tag. This command saves the response body; it does not run JavaScript or extract fields by itself.
2. Inspect the document head
Look inside the page’s <head> for the title, meta elements, links, and structured-data scripts. Google’s documentation lists title, meta, link, script, style, base, noscript, and template as valid head elements. Invalid markup in the head can cause later metadata to be ignored, so a browser’s rendered view alone may conceal a source-HTML problem.
For a manual check, open the page in a browser and use View Page Source, then search for <title, description, canonical, og:, twitter:, robots, and application/ld+json. Use the browser’s developer tools to inspect the live DOM when the source and visible page differ.
3. Extract each namespace without conflating them
- Core fields: read the title element and relevant meta tags such as description, viewport, charset, and robots.
- Links: capture the canonical URL and alternate-language links, including their relationship attributes.
- Social previews: collect Open Graph values such as
og:title,og:description,og:type,og:url, andog:image, plus the page’s Twitter Card fields. - Structured data: find every JSON-LD script, not just the first. Parse each object or array and retain its
@context,@type,@id, URLs, and nested entities. - Robots: check both HTML directives and response headers. Record these separately from descriptive fields.
Schema.org provides machine-readable definitions and a JSON-LD context for its vocabulary. A successful JSON parse only confirms syntax; it does not establish that a type or property is semantically appropriate.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
4. Compare the raw response with the rendered page
If the browser shows a title or social field that is absent from the saved HTML, a client-side application may add or alter it after JavaScript runs. Compare the initial response with the browser’s live DOM. For automated extraction, choose a tool that renders JavaScript when needed; a raw HTTP parser cannot see metadata that is created only after page scripts execute.
Extract metadata with Python
This example uses Requests and Beautiful Soup to extract common tags and parse every JSON-LD script. Install the dependencies with python -m pip install requests beautifulsoup4, then save the script as extract_metadata.py.
import json
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
retrieved_at = datetime.now(timezone.utc).isoformat()
response = requests.get(url, timeout=30, headers={"User-Agent": "MetadataExtractor/1.0"})
soup = BeautifulSoup(response.text, "html.parser")
# Only collect tags in the head; malformed pages may not have one.
head = soup.head or soup
def meta_values(attribute, key):
return [tag.get("content") for tag in head.find_all("meta", attrs={attribute: key}) if tag.get("content")]
def links(rel):
result = []
for tag in head.find_all("link", rel=True):
rels = tag.get("rel", [])
if rel in rels and tag.get("href"):
result.append({
"href": urljoin(response.url, tag["href"]),
"hreflang": tag.get("hreflang"),
"type": tag.get("type"),
})
return result
metadata = {
"requested_url": url,
"final_url": response.url,
"status": response.status_code,
"content_type": response.headers.get("Content-Type"),
"retrieved_at": retrieved_at,
"http_x_robots_tag": response.headers.get("X-Robots-Tag"),
"title": head.title.get_text(" ", strip=True) if head.title else None,
"description": meta_values("name", "description"),
"robots": meta_values("name", "robots"),
"googlebot": meta_values("name", "googlebot"),
"canonical": links("canonical"),
"alternates": links("alternate"),
"open_graph": {},
"twitter": {},
"json_ld": [],
}
for tag in head.find_all("meta"):
key = tag.get("property") or tag.get("name")
content = tag.get("content")
if not key or content is None:
continue
if key.startswith("og:"):
metadata["open_graph"].setdefault(key, []).append(content)
elif key.startswith("twitter:"):
metadata["twitter"].setdefault(key, []).append(content)
for tag in head.find_all("script", attrs={"type": "application/ld+json"}):
raw = tag.string or tag.get_text()
try:
metadata["json_ld"].append(json.loads(raw))
except (json.JSONDecodeError, TypeError) as error:
metadata["json_ld"].append({"parse_error": str(error), "raw": raw})
print(json.dumps(metadata, ensure_ascii=False, indent=2))
The script keeps duplicate values instead of silently choosing one. It resolves relative canonical and alternate URLs against the final response URL. For an audit, preserve the raw HTML too, inspect the response headers, and decide how to handle malformed markup or multiple conflicting values rather than treating the first match as automatically authoritative.
Extract metadata with JavaScript or another parser
In a browser console, inspect the current DOM with a small snippet. This reads the DOM as it exists when you run it; if the site changes metadata asynchronously, wait for that update first.
const head = document.head;
const getMeta = (key, attr = "name") =>
[...head.querySelectorAll(`meta[${attr}="${key}"]`)].map(el => el.content);
const result = {
url: location.href,
title: head.querySelector("title")?.textContent?.trim() ?? null,
description: getMeta("description"),
robots: getMeta("robots"),
googlebot: getMeta("googlebot"),
canonical: [...head.querySelectorAll('link[rel~="canonical"]')].map(el => el.href),
openGraph: [...head.querySelectorAll('meta[property^="og:"]')]
.map(el => [el.getAttribute("property"), el.content]),
twitter: [...head.querySelectorAll('meta[name^="twitter:"]')]
.map(el => [el.name, el.content]),
jsonLd: [...head.querySelectorAll('script[type="application/ld+json"]')]
.map(el => {
try { return JSON.parse(el.textContent); }
catch (error) { return { parseError: String(error), raw: el.textContent }; }
})
};
console.log(result);
For a non-rendered HTML file, use an HTML parser rather than regular expressions. HTML permits variations in attribute order, whitespace, quoting, and nesting; a parser can locate elements and attributes more reliably. Regardless of language, retain duplicate tags and parse errors so they can be investigated instead of discarded.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Extract metadata from JavaScript-heavy pages
Use a rendered-DOM approach when the initial HTML lacks fields that appear after the page loads. A renderer can execute scripts before extraction; for example, OpenGraph.io documents a full_render option and proxy options for its metadata endpoint. See its Open Graph metadata extraction API documentation for those capabilities. Rendering can add latency and may encounter consent flows, bot checks, or other page-specific behavior, so retain the original response details where possible.
Do not assume that what appears visually is also in the metadata. A page can render a heading while its title tag remains missing, or have a social image in HTML that differs from its visible hero image. Compare the exact fields you need in both the initial response and rendered DOM.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Validate results before using them
- Check the final and canonical URLs: confirm that redirects and canonical declarations make sense, and that relative URLs resolve to the intended origin.
- Check duplicates and conflicts: multiple descriptions, canonical links, or social properties can produce ambiguous results. Preserve all values and flag conflicts for review.
- Validate JSON-LD syntax and meaning: parse each block, then check that its types and properties fit the Schema.org vocabulary and reflect content a visitor can see on the page.
- Keep robots separate:
noindex,nofollow, andnosnippetare crawler or presentation instructions, not descriptive data. Google notes that a crawler must be allowed to fetch a page or resource to discover robots directives. - Record evidence: retain requested and final URLs, status, content type, timestamp, headers, and source HTML for reproducible audits.
Google’s guidance on the robots meta tag explains these controls. For head markup and allowed elements, see Google’s valid page metadata documentation. Schema vocabulary is available from Schema.org.
Check canonical and robots tags across many URLs
For a URL inventory, process pages in batches and emit one structured record per requested URL. Keep failures as records with their status or error instead of dropping them; otherwise, an incomplete crawl can look like a clean audit. Distinguish a missing tag from a fetch failure, a blocked response, and a page that only supplies metadata after rendering.
Choose the extraction method by requirement:
- One-off inspection: browser source view, developer tools, or
curlplus a parser. - Large raw-HTML inventory: an HTTP client and parser with concurrency limits, retries for transient errors, and preserved response evidence.
- JavaScript-rendered extraction: a browser or hosted metadata service with rendering support.
- Search or structured-data audit: add explicit checks for conflicting tags, JSON syntax, URL validity, and consistency with visible page content.
Troubleshooting missing or incorrect metadata
The title or description is missing from the fetched HTML
First verify that the response is HTML and inspect its final URL and status. If the field exists in the browser’s live DOM but not in the response, it is likely being added or changed by JavaScript; use a rendered inspection. If it is absent from both, the page may not provide that field.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
The page has metadata but your parser finds none
Confirm that the parser is looking at the final response body, not a redirect message or error page. Check whether malformed markup moved elements outside the expected head subtree, and inspect the raw response to compare. A parser that assumes one attribute order or exact capitalization may also miss valid markup.
Free tools Windows power users keep installed
One-click scans. No signup required.
JSON-LD fails to parse
Capture the raw script text and parse error. A block can contain invalid JSON, be empty, or use syntax that is not JSON. Parse every JSON-LD script independently so one bad block does not conceal valid blocks elsewhere on the page.
Canonical or social URLs are relative or unexpected
Resolve relative references against the effective document URL, then compare the result with the declared canonical and Open Graph URL. Preserve the original value as well as the resolved one; a wrong base URL or redirect assumption can otherwise obscure the source of the discrepancy.
Robots instructions seem absent
Check the response’s X-Robots-Tag header as well as HTML meta tags. A crawler that cannot fetch the page or resource cannot discover directives contained there, so access restrictions and directive absence are different conditions.
Or skip the browser setup
If you need a rendered screenshot alongside an inspection, ScreenshotNeo takes a website URL in one API request and returns an image or PDF. It is a screenshot API, not a metadata parser: use the extraction methods above to collect tags and JSON-LD. Its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
Example cURL request, adapted to the page you want to capture (see the ScreenshotNeo API documentation):
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo’s free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
Does a meta description control whether Google indexes a page?
No. It is descriptive metadata; indexing directives such as noindex are separate controls.
Is JSON-LD the same as Open Graph metadata?
No. JSON-LD expresses structured entities and properties, while Open Graph supplies social-preview metadata.
Can I extract metadata from a URL without opening a browser?
Yes, when the metadata is present in the raw HTML response; use an HTTP client and HTML parser. A rendered browser is needed when scripts add the fields after load.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

