Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Use an LLM to draft varied API test payloads, but treat every response as untrusted until it passes JSON parsing, schema validation, business-rule checks, and—when calls depend on one another—execution against the API. The reliable approach is contract-first: give the model the current API specification and clear field meanings, constrain its output where possible, then validate it independently.
What makes LLM-generated API test data useful—and risky?
LLMs can produce many plausible variations quickly, including values that fit a domain better than generic placeholders. That can help exercise edge cases and reduce the effort of hand-writing fixtures. But plausibility is not correctness: a response may be malformed JSON, violate the request schema, contradict a business rule, or refer to a resource that does not exist in the API’s current state.
Keep the quality checks separate. Syntax validation asks whether the response parses as JSON. Schema validation checks fields, types, allowed values, and extra properties. Business validation checks meaning and relationships. Runtime testing checks whether the API accepts the payload in the state created by preceding calls.
How to generate realistic JSON test data with an LLM
1. Start with the API contract
Provide the current OpenAPI description and the request body’s JSON Schema when available. Include the endpoint’s purpose, parameter meanings, required and optional fields, allowed values, and relevant business rules. A field named status, for example, is not self-explanatory if the API permits only certain transitions or uses domain-specific states.
#1 Best Overall
Microsoft’s guidance recommends relevant, well-structured reference material, clear API path and parameter descriptions, and business policies that help the model interpret an API specification correctly. Its cited synthetic-data generation guidance is marked preview, so availability and behavior may change: Microsoft Foundry synthetic data generation.
2. Specify the output shape and field meanings
Tell the model exactly what to return: for example, an array of request-body objects, with no prose or Markdown surrounding the JSON. Explain each field’s meaning, valid formats, and constraints rather than relying on property names alone. State whether unknown properties are forbidden, how many records to generate, and which cases to include, such as valid boundary values or intentionally invalid inputs.
Google Cloud’s synthetic-data API supports required output field specifications, optional per-field guidance, optional few-shot examples, and a task description. Its documentation specifically favors explicit guidance when a field name could be ambiguous. The stateless API reference documents a maximum of 50 examples in a single request; that is a service limit, not a recommendation to generate that many at once: Google Cloud synthetic-data API reference.
3. Add only useful examples, then review a small batch
Examples can clarify formats, tone, and domain conventions. Choose representative examples that illustrate the target contract without making the output a near-copy of one fixture. Google says examples can improve generated synthetic-data quality and relevance. Generate a small batch first, inspect it, and refine the reference material or generation settings before scaling. Microsoft recommends this iterative approach.
Rank #3
4. Constrain generation where supported
Prefer structured output or constrained decoding tied to a response schema when the model service supports it. A prompt that says “return JSON only” is weaker: Google Cloud documents JSON mode without a response schema as a strong hint, not a guarantee of valid JSON. Its guidance recommends using both JSON response mode and a response schema, or validating client-side and retrying when a schema cannot be predefined. Supported schema fields are a subset, and complex schemas can fail validation or exceed service limits. Check the service’s current supported-schema documentation before relying on a constraint: Google Cloud controlled generation.
How to verify generated payloads before using them
Parse and validate outside the model
Do not make the model the final judge of its own output. Parse the response with a JSON parser, then validate it against the same contract your API expects. Check required fields, data types, formats, enumerations, bounds, and whether extra properties are allowed. Reject or quarantine failures rather than silently coercing them into passing fixtures.
OWASP’s LLM Verification Standard control 5.5 says JSON responses should be valid JSON and undergo schema validation for expected fields and unwanted extra properties. It also recommends structured output or constrained decoding as defense in depth where supported: OWASP LLM Verification Standard.
Test business meaning and API state separately
A schema can confirm that a quantity is numeric and that a status is an allowed string; it cannot by itself establish that the quantity is available or that the status is valid after the preceding operation. Add explicit tests for cross-field consistency, resource relationships, business policies, and valid state transitions. Where possible, execute the generated requests against a controlled test environment and inspect the API’s responses.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to generate data for multi-call API workflows
When one operation creates a resource consumed by a later operation, treat dependencies inferred by an LLM as hypotheses, not facts. Validate them by executing the calls: use concrete responses from producer operations to populate consumer inputs, then refine resource pools and input constraints based on runtime behavior.
The APIPilot preprint reports 92.3% operation coverage, up to 58.6% code coverage, and an 88.1% workflow execution success rate in an evaluation of 16 REST API services. These are the authors’ results for that evaluation, not a forecast or guarantee for another API: APIPilot preprint.
Can synthetic test data replace production data?
Synthetic values can reduce reliance on actual captured values, but the label “synthetic” is not a universal privacy guarantee. Katalon documents a synthetic mode that derives values from captured patterns without using the actual captured values, contrasting it with raw and raw-with-mocked-PII modes. That product description illustrates one approach; it does not establish that every synthetic-data workflow is anonymous, risk-free, or legally compliant: Katalon data masking documentation.
Recommended Free Tools
Do not send sensitive production data to a model unless your organization has approved the service and its data handling. Prefer contract examples and purpose-made synthetic values, and assess privacy, retention, access, and legal requirements through your organization’s controls.
A practical generation-and-checking loop
- Assemble the contract: select the current OpenAPI operation, request schema, field descriptions, and applicable business rules.
- Define test cases: specify how many payloads to draft and which valid, boundary, or invalid scenarios they should cover.
- Request bounded JSON: use schema-constrained output if available; otherwise request a precise shape and treat the response as untrusted.
- Validate locally: parse the JSON and check schema requirements, unexpected fields, and business invariants in code.
- Execute workflow cases: for dependent calls, use actual test-environment responses to supply later inputs and assess runtime behavior.
- Review before scaling: inspect a small batch, correct contract or prompt gaps, and only then increase generation volume.
For repeatable tests, preserve the prompt or generation configuration, schema version, model/service settings, and validation results alongside the fixtures. This makes it easier to identify whether a later failure came from a contract change, a generation change, or application behavior.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

