Free tools Windows power users keep installed
One-click scans. No signup required.
To extract structured data reliably, use a staged pipeline: fetch the page, parse JSON-LD, traverse Microdata and RDFa, preserve graph structure and provenance, then validate the combined result. Use an HTTP client when the markup is in the server response; use a browser renderer when JavaScript adds it after load.
Choose the right extraction path first
“Structured data” is not one format. A page may publish Schema.org entities as JSON-LD in a <script type="application/ld+json"> block, annotate visible HTML with Microdata, describe relationships with RDFa, or use several formats at once. Your first decision is whether the required markup exists in the initial HTTP response.
| Situation | Acquisition method | Trade-off |
|---|---|---|
| Markup is present in the server response | HTTP client plus HTML parser | Fast, inexpensive and reproducible; no JavaScript execution |
| JavaScript injects JSON-LD or attributes after load | Browser automation, then inspect the rendered DOM | Covers client-generated data but consumes more time and resources |
| Both static and rendered representations matter | Run the HTTP pass, then a rendered fallback when required | Best coverage; requires duplicate detection and provenance rules |
Google documents that JSON-LD generated by JavaScript can be processed when it is available in the rendered DOM. A static download alone therefore cannot prove that a page has no structured data.
Define a lossless internal record
Do not immediately flatten everything into a dictionary of fields. JSON-LD is a serialization of an RDF dataset, and its @graph can contain multiple connected entities. Microdata and RDFa likewise express subject–predicate–object relationships. Keep a common envelope while retaining the original representation:
#1 Best Overall
{
"source_url": "https://example.com/article",
"format": "jsonld|microdata|rdfa",
"type": "Article",
"id": "https://example.com/article#post",
"properties": {},
"raw": {},
"source_element": "optional selector or fragment"
}
- source_url: the URL requested or rendered.
- format: which syntax produced the record.
- type and id: values from
@type/@id,itemtype/itemid, or RDFatypeof/about. - properties: a normalized view for application code.
- raw: the untouched JSON object, element attributes, or RDFa values.
- source_element: enough location information to audit a value later.
Preserving raw fragments prevents a normalization mistake from becoming impossible to diagnose. Delay array flattening, graph expansion and type coercion until the consumer’s schema requires them.
Extract JSON-LD with Python
JSON-LD is usually the simplest and most complete layer to parse first. The following script fetches the initial HTML, selects every JSON-LD script, preserves arrays and graph objects, and records malformed blocks instead of silently dropping them.
import json
import requests
from bs4 import BeautifulSoup
url = "https://example.com/article"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
records = []
for node in soup.select('script[type="application/ld+json"]'):
raw = node.string or node.get_text()
try:
value = json.loads(raw)
records.append({
"source_url": url,
"format": "jsonld",
"type": value.get("@type") if isinstance(value, dict) else None,
"id": value.get("@id") if isinstance(value, dict) else None,
"properties": value,
"raw": raw
})
except json.JSONDecodeError as exc:
records.append({
"source_url": url,
"format": "jsonld",
"parse_error": str(exc),
"raw": raw
})
result = {"url": url, "jsonld": records}
print(json.dumps(result, ensure_ascii=False, indent=2))
Install the two dependencies with python -m pip install requests beautifulsoup4. The parser accepts a top-level object, a top-level array, or an object containing @graph; leave each shape intact. Some publishers put comments, trailing commas or multiple JSON values in a script. Those are invalid JSON: retain the raw text and report the error rather than applying an unsafe “repair” that changes meaning.
Handling graph and context fields
- Keep
@context; it defines how terms map to vocabularies. - Keep
@graphas an array of nodes until you deliberately build an ID index. - Keep
@idvalues as identifiers, not display labels. - Keep arrays even when they contain one item; cardinality can matter.
- Resolve relative URLs only in a later normalization stage, using the document URL as the base.
Add Microdata extraction
Microdata is attached to HTML elements rather than a JSON script. Look for itemscope, itemtype, itemprop and, when present, itemid. A scope starts an item; descendants carrying itemprop become its properties. A nested itemscope is itself the property value and should be parsed recursively.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Use the element’s value-bearing attribute instead of always taking visible text:
meta→contentaudio,embed,iframe,img,source,track,video→srca,area,link→hrefobject→datadata,meter→value- all other elements → their text content
When a property has multiple values, store an array. Follow nested scopes rather than flattening them into strings. Preserve the containing element’s selector or a compact HTML fragment in source_element so a consumer can explain where a value originated.
Add RDFa extraction
RDFa expresses relationships through attributes such as about (subject), typeof (type), property (predicate), resource, href and src (object identifiers). Build triples or subject-centered records:
- Establish the current subject from
about; inherit the parent subject when it is omitted. - Read
typeofas one or more types. - For each
property, choose the object fromcontent,datetime,resource,href,src, or text according to the element. - Resolve relative IRIs against the page URL.
- Attach nested statements to the appropriate subject and retain the original attributes.
A subject–predicate–object model is safer than a flat map because it can represent repeated predicates and links between entities. The W3C RDFa API defines queries by type, subject and property; matching that conceptual model makes later exports predictable.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
Use a rendered DOM when JavaScript supplies the data
Fetch the page first. If the response contains no relevant scripts or attributes, load the URL in a real browser (for example, an approved automation environment), wait for the application to finish, and capture the final DOM. Also inspect network responses when a widget retrieves structured payloads without inserting them into the DOM.
In browser-side JavaScript, the JSON-LD portion can be collected after rendering with this pattern:
const response = await fetch(url);
const html = await response.text();
const doc = new DOMParser().parseFromString(html, "text/html");
const blocks = [...doc.querySelectorAll('script[type="application/ld+json"]')]
.map(node => {
try { return JSON.parse(node.textContent); }
catch (error) { return { _parse_error: true, raw: node.textContent }; }
});
For automation, wait for a meaningful selector, a settled application state or a bounded delay. Record whether each record came from the original response or rendered DOM. Run duplicate detection by stable identifiers such as @id, item IDs and resolved RDF subjects; identical data from two passes should not create two entities.
Normalize without losing meaning
After all format-specific passes, map records into your application schema. Keep one record per source representation until you have an explicit precedence policy. For example, you might prefer JSON-LD for a field when its value is valid, use Microdata as corroboration, and flag conflicts instead of silently choosing one.
Recommended normalization checks
- Resolve relative URLs against the document’s base URL.
- Normalize scalar-versus-array handling without deleting multiplicity.
- Index graph nodes by
@idand connect references rather than copying nested objects repeatedly. - Retain unknown properties for forward compatibility.
- Attach a provenance object containing format, source element, raw value and extraction time.
- Detect duplicate entities and conflicting values, then expose the conflict to downstream code.
Validate the combined result
During development, submit the source URL or extracted markup to Schema.org’s Markup Validator. It can extract JSON-LD, RDFa and Microdata, combine them, summarize the graph and identify syntax mistakes. Validate the combined representations, not only the JSON-LD script, because a page can contain valid syntax with contradictory entities.
What validation catches
- Malformed JSON-LD that your parser recorded as an error.
- Missing or incorrectly nested Microdata scopes.
- RDFa subjects or properties that cannot be resolved.
- Conflicting representations of the same entity.
- Unexpected graph shape after normalization.
Common failures and precise fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| No records found | You searched only JSON-LD, or data is injected later | Add Microdata and RDFa passes; retry with a rendered DOM |
| One entity appears many times | Top-level objects and @graph nodes were flattened together |
Index stable IDs and deduplicate after preserving provenance |
| Nested properties are missing | Nested item scopes or graph references were converted to strings | Recurse through scopes and retain node references |
| URLs are unusable | Relative values were treated as absolute | Resolve against the page URL during normalization |
| Parser crashes on one script | Malformed JSON-LD | Catch the decode error, retain raw text and report the location |
| Static and rendered values disagree | Different representations or client personalization | Keep both records, apply a documented precedence rule and flag the conflict |
| Browser pass is empty | Captured before hydration or blocked by a bot check | Wait for a known selector, inspect console/network errors and verify access manually |
Performance, reliability and operating costs
Use the cheapest reliable path per URL. An HTTP parser is generally faster and easier to reproduce than a browser. Reserve rendering for pages that demonstrably require JavaScript, and cache the raw response and rendered snapshot when your freshness policy permits.
- Set connection and total timeouts; never allow an unbounded fetch.
- Limit response size and reject unexpected content types before parsing.
- Use bounded retries with backoff for transient network errors, not for deterministic parse failures.
- Rate-limit requests and respect the site’s access rules.
- Store status code, final URL, content type, extraction path and timestamps with each result.
- Test against pages containing arrays,
@graph, nested items, malformed blocks and JavaScript-only markup.
Browser rendering costs more CPU and introduces timing variability. A two-stage design—HTTP first, browser fallback only when needed—usually gives the best balance of coverage and throughput.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo can load a page in a managed browser when you need the post-render view before inspecting structured data. Its screenshot endpoint is also useful for preserving visual evidence alongside the extracted record.
Best Value
One GET request returns an image or PDF; use the response headers to distinguish a clean capture from a bot check, blank page, timeout, failed load or cache hit. Those unsuccessful cases are not billed.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the complete parameter list and rendering options in the ScreenshotNeo documentation. The same request from Python is:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie banners, newsletter popups and chat widgets before the shot. Bot checks, blank pages and failed loads are never billed, and an MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Practical decision checklist
- Fetch the URL and record status, content type and final URL.
- Parse every JSON-LD block, preserving
@context,@graph, arrays and raw text. - Traverse Microdata scopes and value-bearing attributes.
- Traverse RDFa subjects, types, properties and resource attributes.
- If data is absent, render the page and repeat the DOM passes.
- Normalize into your internal envelope without flattening relationships.
- Deduplicate by stable IDs and flag conflicting values.
- Validate the combined output with Schema.org’s Markup Validator.
- Persist provenance so every field can be traced to a representation and source element.
Frequently Asked Questions
Should I save the extracted JSON or the original HTML?
Save both when possible. HTML and raw JSON-LD preserve the evidence needed to reprocess a page after your normalization rules change.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesCan a page expose more than one structured-data format?
Yes. Treat JSON-LD, Microdata and RDFa as separate representations, then deduplicate and reconcile them using stable identifiers and an explicit precedence policy.
When is browser rendering unavoidable?
It is required when the initial HTTP response lacks the data and JavaScript inserts it after load, or when a widget delivers the payload only through runtime network calls.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

