Define the record before choosing an extractor. Automatic extraction works when you specify the fields you need, identify what kind of source you have, choose a method suited to that source, and validate every returned value against the original text. A JSON object that matches your schema can still contain a false date, a misidentified person, or an invented value.
What “structured extraction” means
Unstructured text is prose, email, a report, a support ticket, or another source without a fixed row-and-column format. Structured information is a record with named, typed fields that software can store and query. For example, a contract paragraph might become:
{
"parties": ["Northwind Ltd", "Alpine Components"],
"effective_date": "2026-01-15",
"renewal_term_months": 12,
"governing_law": "New York",
"source_quote": "..."
}
The extraction task is not merely “turn text into JSON.” It is a specification, interpretation, and verification pipeline. Decide what counts as one record, whether a field can be absent, and how to represent uncertainty before selecting a model or API.
Step 1: Design a schema first
Define record boundaries
State what one record represents: an invoice, a person, an adverse event, a job posting, or one event mentioned in a paragraph. If a document contains several records, specify whether the output is an array and how duplicates are handled.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Classify every field
- Required: the record is incomplete without it.
- Optional: return a value only when the source supports one.
- Repeated: use an array, even when only one item appears.
- Nullable or unknown: use
null(or a documented unknown value), never a guess. - Evidence: preserve a source span or quote for fields that need auditability.
Specify types and allowed values
Use an ISO date format such as YYYY-MM-DD, numbers rather than numeric strings, and enumerations for controlled values. Define rules for conflicting mentions, approximate dates, units, negation, and pronouns. A schema can constrain shape; it cannot establish that the content is true.
Step 2: Identify the input and its failure modes
| Input | Upstream work | Typical risk |
|---|---|---|
| Clean digital text | Normalize encoding and remove irrelevant boilerplate | Context, negation, and references are misread |
| Scanned page or image PDF | OCR, orientation correction, and quality checks | Character errors change names, numbers, or dates |
| Forms | Layout-aware key/value detection | Values are assigned to the wrong label |
| Tables | Cell and row reconstruction | Columns shift or merged cells disappear |
| Web pages | Capture or parse the rendered page, then remove navigation and consent UI | Dynamic content, popups, and bot checks obscure the data |
For scans, forms, and tables, OCR and layout analysis are a distinct upstream step from mapping content into your custom semantic schema. Do not send an OCR transcript to a language model without checking confidence and reading order.
Step 3: Choose an extraction approach
Schema-constrained LLM output
Use a general-purpose model when fields depend on context, relationships, or domain-specific wording. OpenAI’s documentation says, “You can define structured fields to extract from unstructured input data, such as research papers.” Its Structured Outputs feature constrains the response to a supplied schema; function calling is instead intended to connect a model to application functions in a pipeline that fetches text, converts it, and saves data. See the OpenAI Structured Outputs guide and Function Calling documentation. Check the current model’s supported JSON Schema subset.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Constrained output solves parsing and shape. It does not prove that a value is supported by the source. Require the model to return null for absent evidence and, where appropriate, a quote and character offsets.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Named-entity analysis
Entity APIs are a good fit when you need predefined classes such as people, organizations, locations, dates, or products. Google Cloud Natural Language’s entity analysis returns recognized entities and associated information; review its basics and analyzeEntities API reference. Compare supported entity types, language and domain fit, offsets, precision, and recall on your corpus. An entity API will not automatically implement a custom business schema such as “contract termination reason.”
Document-analysis and OCR services
Use a layout-aware service for scans, forms, and tables. AWS Textract’s AnalyzeDocument operations cover detected text, forms, tables, query responses, and signatures, with form keys linked to values. See Textract analysis and its response objects. You will usually map its blocks, cells, or key/value pairs into your own schema and then validate the mapping.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
| Approach | Best fit | Evaluate |
|---|---|---|
| Schema-constrained LLM | Custom fields and contextual prose | Field accuracy, absent/ambiguous evidence, latency, cost, privacy, schema support |
| Named-entity API | Supported entity classes | Entity types, language/domain fit, precision/recall, offsets, integration |
| Document/OCR service | Scans, forms, and tables | OCR and layout accuracy, representation, customization, throughput, cost, data handling |
A practical LLM extraction pipeline
- Ingest: preserve the original file, page number, and source identifier.
- Preprocess: OCR scans, reconstruct reading order, remove repeated headers, and retain page or character locations.
- Chunk: split long documents at logical boundaries while carrying section headings and record identifiers into each chunk.
- Extract: submit the schema, explicit null and ambiguity rules, and the relevant text.
- Validate: check JSON/schema validity, required fields, types, enumerations, dates, ranges, and cross-field rules.
- Ground: verify important values against source spans; route unsupported or conflicting values to review.
- Merge: combine chunk results deterministically, resolving duplicate mentions according to a written policy.
- Store: retain raw input, normalized output, model/version, prompt or schema version, evidence, and validation status.
Prompt and output rules
Tell the extractor exactly what each field means. “Customer” might mean the purchaser, account owner, or end user. Instruct it to return null when the text does not support a value, distinguish explicit statements from inferences, and include evidence for high-risk fields. Keep temperature and model settings consistent during evaluation.
Python example with a schema
The exact SDK syntax varies by provider and model; the important parts are the schema, null policy, and evidence fields. The following pattern uses an API that supports JSON Schema-constrained output:
from typing import Optional, List
from pydantic import BaseModel
class Evidence(BaseModel):
quote: str
start: Optional[int] = None
end: Optional[int] = None
class Case(BaseModel):
case_id: Optional[str] = None
people: List[str] = []
event_date: Optional[str] = None
severity: Optional[str] = None
evidence: List[Evidence] = []
text = open("input.txt", encoding="utf-8").read()
# Send `text` and Case's JSON schema to your chosen structured-output API.
# In the instruction: use null for missing values, never invent, and quote support.
# Parse the response with Case.model_validate_json(response_text).
In production, avoid mutable default lists in ordinary Python classes, enforce date parsing after the model response, and reject records that fail validation rather than silently coercing them.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Validation: make extraction accountable
Schema and type checks
- Reject malformed JSON and unknown fields if your contract forbids them.
- Check required fields, numeric ranges, date formats, enum values, and array cardinality.
- Normalize units and time zones explicitly; retain the original value alongside the normalized one.
Source-support checks
For each material value, locate the quoted span in the source and verify that the span actually entails the value. A date mentioned in a “previous version” paragraph may not be the document’s effective date. Check negation (“no prior diagnosis”), uncertainty (“可能”), and attribution (“the vendor claims”).
Business rules
Cross-field rules catch errors that JSON validation cannot: an end date cannot precede a start date; a currency amount needs a currency; a percentage must be within its permitted range; and a “renewal” event should have a renewal term or an explicit reason it is unknown. Send violations to a review queue with the source and model output.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Evaluation before production
Create a representative, manually labeled set covering short and long documents, spelling variation, tables, scans, missing fields, and contradictory statements. Compare field-level precision and recall, schema-validity rate, error categories, latency, cost, privacy constraints, and integration effort. Measure per-field performance, not only whether the whole record “looks right.” Keep a fixed test set and version your schema, prompts, OCR settings, and model.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Vendor benchmark claims require narrow attribution. OpenAI’s August 6, 2024 launch announcement reported 100% on its complex JSON-schema-following evaluation for gpt-4o-2024-08-06, versus less than 40% for gpt-4-0613. That is an OpenAI-reported schema-following result, not an independent benchmark or a claim of factual accuracy on arbitrary text; see the announcement.
Reliability, privacy, and cost decisions
- Reliability: use retries with idempotency keys, timeouts, bounded concurrency, and dead-letter handling. Cache extraction by a content hash and schema version.
- Long documents: chunk with overlap, preserve headings, and run a second pass to reconcile entities and totals.
- Human review: require review for low OCR confidence, missing evidence, conflicting values, or high-impact decisions.
- Privacy: classify personal or regulated data, confirm retention and regional processing terms, minimize prompts, and restrict logs.
- Cost: estimate OCR, model tokens, retries, storage, and review labor. Batch low-risk work, but do not trade away evidence for a lower unit price.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Valid JSON, wrong facts | Schema controls shape, not truth | Require evidence, add source-support checks, and review ambiguous fields |
| Many missing values | OCR loss, poor chunking, or overly strict definitions | Inspect OCR and reading order, add context, and distinguish null from extraction failure |
| Tables shifted | Plain-text extraction discarded layout | Use layout-aware table extraction and test merged cells |
| Dates disagree | Multiple dates or unclear semantics | Name each date field precisely and apply a precedence rule with evidence |
| Schema errors after a model update | Unsupported schema feature or changed model behavior | Pin versions, use the provider’s supported subset, and rerun contract tests |
| Timeouts and rate limits | Large chunks or excessive concurrency | Reduce chunk size, back off with jitter, and queue work asynchronously |
Or skip the browser setup
If your source is a rendered web page, first obtaining a clean screenshot can simplify OCR and layout capture. ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether it was billed. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
One call returns PNG, JPEG, WebP, or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all options, including full-page lazy-image loading, CSS-selector element capture, device presets, custom CSS/JavaScript, wait conditions, headers and cookies, blocking requests, PDF page ranges, signed links, async webhooks, bulk capture of up to 100 URLs per call, and caching TTLs. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently asked questions
Can structured output be treated as proof?
No. It enforces the response shape. Verify each important value against the source and your business rules.
Should I use an entity API or an LLM?
Use an entity API for its supported, predefined classes; use a schema-constrained LLM for contextual custom fields. Evaluate both on labeled examples from your corpus.
What changes for scanned PDFs?
Add OCR and layout reconstruction first, measure OCR quality, and preserve page coordinates or quotes so semantic extraction can be audited.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

