Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To parse JSON while web scraping, first determine whether the server returned JSON itself or HTML containing a JSON payload. For a direct JSON response, check the HTTP status and then decode the body. For JSON embedded in a page, parse the HTML, locate the specific script element, and decode its contents. Treat status, decoding, and data-shape validation as separate steps: a JSON decoder can succeed even when the HTTP request failed, and valid JSON can still have fields your scraper does not expect.

How do I parse JSON in web scraping?

A reliable scraper handles several distinct layers rather than passing every response straight to a JSON parser:

  1. Fetch: retain the response status, headers, and body.
  2. Check HTTP success: reject or handle unsuccessful statuses before treating the payload as a successful result.
  3. Identify the representation: determine whether the body is JSON, HTML, or something else.
  4. Decode and validate: parse JSON, then check that the resulting object has the expected types and fields.
  5. Extract and normalize: map the source structure into the data your application needs, while retaining enough context to debug changes.

Requests provides response.json() for decoding a JSON response. Its documentation cautions: “It should be noted that the success of the call to r.json() does not indicate the success of the response.” A server may return valid JSON describing an error with an unsuccessful HTTP status. Check response.raise_for_status() or handle the status explicitly, separately from JSON decoding. See the Requests Quickstart documentation.

Direct JSON response: runnable Python example

Use this path when the endpoint returns JSON rather than a web page. Replace the example URL with an endpoint you are allowed to access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

url = "https://example.com/api/items"

try:
    response = requests.get(url, timeout=30)
    response.raise_for_status()
    data = response.json()
except requests.exceptions.HTTPError as exc:
    print(f"HTTP error: {exc}; status={response.status_code}")
except requests.exceptions.JSONDecodeError as exc:
    print(f"Response was not valid JSON: {exc}")
except requests.exceptions.RequestException as exc:
    print(f"Request failed: {exc}")
else:
    if not isinstance(data, dict):
        raise ValueError(f"Expected a JSON object, got {type(data).__name__}")
    items = data.get("items")
    if not isinstance(items, list):
        raise ValueError("Expected 'items' to be a list")
    print(items)

Timeouts and exception types can vary by client library and version; consult the Requests documentation for the version installed in your project. Avoid logging credentials or sensitive response contents when reporting errors.

How do I extract JSON from a website that returns HTML?

A page can contain data in a <script> element even though its outer response is HTML. In that case, do not try to decode the entire HTML document as JSON. Parse the markup first, select the element that carries the payload, and decode that element’s text.

The following example uses Requests and Beautiful Soup to find a JSON-LD script. Install the dependencies with python -m pip install requests beautifulsoup4. The target must actually include a script with the specified type; handle absence as an extraction failure rather than silently returning an empty result.

import json
import requests
from bs4 import BeautifulSoup

url = "https://example.com/article"
response = requests.get(url, timeout=30)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
script = soup.find("script", type="application/ld+json")
if script is None or not script.string and not script.get_text(strip=True):
    raise ValueError("No JSON-LD script found")

payload = script.string or script.get_text()
try:
    data = json.loads(payload)
except json.JSONDecodeError as exc:
    raise ValueError(f"JSON-LD script is not valid JSON: {exc}") from exc

print(data)

Beautiful Soup recommends specifying a parser because different parsers can construct different trees from malformed markup, and available parsers may differ across installations. The example explicitly chooses Python’s built-in html.parser; use the same parser consistently in environments where extraction output must be reproducible. See the Beautiful Soup documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JSON-LD is structured data, not just arbitrary script text

The World Wide Web Consortium (W3C) describes JSON-LD 1.1 as “a JSON-based format to serialize Linked Data.” Ordinary JSON decoding gives you its JSON syntax as Python data, but it does not by itself perform every JSON-LD processing step. If your use case depends on linked-data semantics—such as interpreting contexts, identifiers, or graph relationships—use a JSON-LD-aware processor and the relevant processing rules. See the W3C JSON-LD 1.1 Recommendation and JSON-LD 1.1 Processing Algorithms and API.

Also, not every script element is JSON-LD. Select by its declared type and by the page structure you have verified; a generic script may contain executable JavaScript rather than a standalone JSON document. Do not extract a value by searching all visible page text for braces and assuming the first match is the data.

How can I tell whether the data is in an API response, HTML, or dynamic content?

Inspect the actual response body and headers before choosing a parser. A response’s content type is a useful clue, not a substitute for examining the body: endpoints can be misconfigured, and error pages may be HTML even when an application expected JSON.

  • JSON body: decode the response directly and validate the top-level type and expected fields.
  • HTML with embedded data: parse the markup and identify the exact element or script containing the payload.
  • Data appears after page load: the initial HTML may not include it. Look for the data request that supplies it, an embedded script path, or a rendered-page approach. Scrapy’s documentation distinguishes ordinary response parsing from cases involving dynamically loaded content; see Scrapy: Dynamic content.

If the site exposes both an API-like route and a page with embedded data, compare whether each includes the fields you need, whether its representation is documented and stable, whether data is loaded dynamically, and what access conditions apply. There is no universally best route: prefer the route that is adequate, reliable, and permitted for your use.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should I handle encoding, parsing, and data-shape errors?

Keep response decoding separate from JSON parsing

Requests exposes decoded text as response.text and raw bytes as response.content. Its text decoding uses the response encoding, which may need deliberate inspection or adjustment for unusual responses. When encoding matters, inspect the response headers and the library’s encoding behavior rather than assuming every page uses the same character set. For HTML extraction with Beautiful Soup, pass decoded text or bytes intentionally and use a consistent parser.

Validate the structure after decoding

JSON syntax alone cannot guarantee that the payload matches your scraper’s assumptions. Check the top-level type, required keys, and value types before using them. For instance, a JSON array is valid JSON but not interchangeable with a JSON object; an object may be valid while omitting the items field your code expects. Treat missing keys, unexpected null values, and selector mismatches as explicit extraction errors with enough non-sensitive context to diagnose them.

Preserve a reproducible example when debugging

Keep the original payload or a safe, representative fixture separate from normalized application data. This makes it easier to distinguish a source-page change from a decoder bug or a change in your own transformation logic. Strip secrets and personal data before saving or sharing captured responses.

Why does my JSON parser fail on a scraped response?

Symptom Likely cause What to check or do
JSON decode error at the start of the body The response is empty, HTML, or another non-JSON format. Check status, headers, and a safe preview of the body. Confirm whether you need an HTML parser rather than a JSON decoder.
JSON decode error near the end The payload may be truncated, malformed, or not a complete JSON document. Check the response length and whether the server or network ended the response early. Capture a safe fixture and reproduce the decode locally.
Decoder succeeds but the scraper reports failure The status may be unsuccessful, or the decoded value may have an unexpected structure. Check HTTP status independently, then validate top-level type, keys, and field types.
JSON-LD selector returns nothing The page lacks that script, uses a different structure, or data is loaded dynamically. Inspect the returned HTML, verify the script’s type and selector, and investigate the relevant data request or rendered content.
HTML extraction differs across machines Different parser implementations can recover malformed markup differently. Specify and consistently install a parser; compare the parsed tree against a captured fixture.
Text is garbled or characters are wrong The response encoding may not be what your code assumes. Inspect response encoding information and the HTTP client’s decoding behavior; use raw bytes when you need to control decoding deliberately.

Requests documents that empty or invalid response bodies can cause JSON decoding failures and that decoding success does not prove the HTTP request succeeded. Those are different failure layers: diagnose request/status problems before changing selectors or JSON parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How can I scrape JSON responsibly and reliably?

Before automating collection, check the target site’s terms and access conditions and any requirements that apply to your use and jurisdiction. A site’s robots.txt is relevant to crawler behavior, but it is not a replacement for those checks. Google’s documentation explains how robots.txt can manage crawling traffic and warns that it should not be used to hide pages from search results; that guidance describes Google’s crawler, not a universal rule for every scraper. See Google Search Central: Introduction to robots.txt.

For operational reliability, set a reasonable timeout, avoid unnecessary repeated requests, retain status and relevant headers for diagnostics, and make extraction assumptions explicit. A cache can reduce repeat fetching where your use and the target’s rules permit it. If page content is inconsistent, compare the saved response with the current response before concluding that a parser is at fault.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you need a screenshot or rendered-page capture rather than the underlying JSON payload, ScreenshotNeo is a website screenshot API and MCP server. Its one-GET API can return PNG, JPEG, WebP, or PDF output. For a basic capture, use cURL like this; replace the URL with the page you need and supply your API key:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie and consent banners are accepted like a visitor and removed, along with supported newsletter popups and chat widgets, before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

FAQ

Is JSON-LD the same as an API response?

No. JSON-LD is a JSON-based linked-data format that may be embedded in HTML; an API response is a response route and representation, not a particular data standard.

Should I parse all script tags on a page?

No. Identify the script type and the specific element that contains the data you need. A script may contain executable JavaScript or unrelated content rather than a JSON payload.

Does a valid JSON response mean the scrape succeeded?

No. A decoder can parse a valid error object returned with an unsuccessful HTTP status. Check the status and validate the data shape as separate steps.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.