Mock the lowest layer that answers the question your test is asking. For application workflows, use an in-memory scripted model that yields deterministic text and tool or handoff results. For request serialization, authentication, provider defaults, server-sent-event (SSE) framing, retries, cancellation, or network failures, keep your real OpenAI adapter and intercept its HTTP transport. Use an explicit event fixture only when event ordering, partial rendering, or malformed streams is the behavior under test.
Start by choosing the test boundary
A streamed-response test can accidentally test the wrong thing. A fixture that returns one final string cannot reveal a renderer that mishandles deltas; a hand-written event list can make an ordinary workflow test brittle. Decide whether the contract is your application workflow, the normalized SDK stream, or the provider wire protocol.
| Boundary | Best fixture | What it proves | Main trade-off |
|---|---|---|---|
| Application workflow | Scripted in-memory model | Accumulated text, tool calls, handoffs, retries, state transitions | Does not prove provider framing or HTTP behavior |
| Normalized stream | Explicit SDK stream events | Exact ordering, partial rendering, cancellation and completion handling | More maintenance when event types change |
| Provider/HTTP | Controlled HTTP server or transport interceptor | Request JSON, headers, SSE framing, status codes, disconnects and timeouts | Fixtures track provider and SDK versions |
| Browser/proxy | Fixtures on both sides of the conversion | Correct translation from upstream SSE to your browser format, such as NDJSON | Requires two contracts and two sets of failure tests |
The OpenAI Agents SDK guidance makes the same distinction: use a scripted model for normal workflow tests, and reserve a stream-event helper for cases where the exact normalized event sequence is part of the behavior under test.
Know which stream you are imitating
Responses API
Responses streaming uses server-sent events with semantic lifecycle events. Common events include response.created, response.output_text.delta, response.completed, and error. A consumer normally appends text from each delta and treats the completion event as the terminal success signal.
#1 Best Overall
Chat Completions
Chat Completions streaming delivers incremental chunks. Each chunk has a delta object that may contain a role token, a content token, or nothing. The first chunk can establish the role, content can arrive over several chunks, and a final chunk may contain no content. Do not substitute a Chat Completions chunk fixture for a Responses event fixture: their shapes and termination conventions differ.
Preserve the representation at every hop
The original HTTP response may be SSE, while an internal adapter exposes an async iterable of parsed events and your browser endpoint emits newline-delimited JSON (NDJSON). These are not interchangeable. The Node SDK’s ResponseStream.fromReadableStream() expects newline-separated JSON, not the original SSE wire format. A proxy or fixture must therefore emit the representation expected by the code under test.
Stub the model for workflow tests
Keep this fixture deliberately boring: one deterministic response, optional tool or handoff metadata, and an ordinary stream produced by your application adapter. The following JavaScript example is a complete in-memory async generator; replace the adapter call in your application with this scripted source.
class ScriptedModel {
constructor(events) { this.events = events; }
async *stream() {
for (const event of this.events) yield event;
}
}
async function collectText(model) {
let text = '';
let completed = false;
for await (const event of model.stream()) {
if (event.type === 'text.delta') text += event.text;
if (event.type === 'response.completed') completed = true;
}
if (!completed) throw new Error('stream did not complete');
return text;
}
const model = new ScriptedModel([
{ type: 'text.delta', text: 'Hello' },
{ type: 'text.delta', text: ', test!' },
{ type: 'response.completed' }
]);
collectText(model).then(result => {
if (result !== 'Hello, test!') throw new Error('unexpected text');
});
Assert the final accumulated message, tool arguments, handoff target, retry count, and state transitions at this boundary. You are testing your business logic, not whether an upstream server chose a particular frame order.
Recommended Free Tools
Python equivalent
class ScriptedModel:
def __init__(self, events):
self.events = events
async def stream(self):
for event in self.events:
yield event
async def collect_text(model):
text = []
completed = False
async for event in model.stream():
if event["type"] == "text.delta":
text.append(event["text"])
elif event["type"] == "response.completed":
completed = True
if not completed:
raise RuntimeError("stream did not complete")
return "".join(text)
# In an async test:
# result = await collect_text(ScriptedModel([
# {"type": "text.delta", "text": "Hello"},
# {"type": "text.delta", "text": ", test!"},
# {"type": "response.completed"},
# ]))
# assert result == "Hello, test!"
Use an explicit event sequence when timing and ordering matter
Switch from a scripted final answer to explicit normalized events when you need to prove that the UI renders partial output, stops after cancellation, closes resources, or handles an error after some text has already appeared. Include a terminal completion event for the success case and assert that the consumer closes its stream.
- Emit several small deltas rather than one large message to exercise incremental rendering.
- Place role metadata before content when your consumer expects that order.
- Test an empty delta and a completion event with no text.
- Deliver an
errorevent after one or more deltas and verify the UI reports failure without pretending the answer completed. - Cancel while a delayed event is pending and verify that no later event mutates application state.
Raw Responses streams in the Node SDK are single-consumer. If two independent consumers are required, use stream.tee() and test each branch separately; do not iterate the same raw stream twice.
Rank #2
Mock the HTTP layer for wire-level confidence
Keep the real model adapter and intercept its HTTP request when you need to verify serialized JSON, authorization headers, provider defaults, SSE parsing, retries, or transport failures. The fixture should match the request, return the exact media type expected by the endpoint, write one event per frame, then send the endpoint’s terminal event and close.
A minimal queued SSE server in Node.js
import http from 'node:http';
const queue = [
{ event: 'response.created', data: { id: 'test-1' } },
{ event: 'response.output_text.delta', data: { delta: 'Hello' } },
{ event: 'response.output_text.delta', data: { delta: ' from a fixture.' } },
{ event: 'response.completed', data: {} }
];
const server = http.createServer((req, res) => {
if (req.url !== '/v1/responses' || req.method !== 'POST') {
res.writeHead(404).end();
return;
}
res.writeHead(200, {
'content-type': 'text/event-stream',
'cache-control': 'no-cache',
connection: 'keep-alive'
});
for (const item of queue) {
res.write(`event: ${item.event}\n`);
res.write(`data: ${JSON.stringify(item.data)}\n\n`);
}
res.end();
});
server.listen(8787, () => console.log('fixture on http://127.0.0.1:8787'));
Point your adapter’s base URL or injected transport at this server. Assert the outgoing method, path, body, authorization header, and streaming flag before consuming the response. Keep the queue per test so one test cannot consume another test’s events.
Free tools Windows power users keep installed
One-click scans. No signup required.
Failure fixtures worth keeping
- Non-200 response: return an error status and a JSON body; assert the adapter surfaces a useful exception and does not enter the token loop.
- Mid-stream error: send valid deltas, then an
errorevent; assert partial text is marked incomplete. - Truncated body: omit the terminal event and close the socket; assert a disconnect or incomplete-stream error.
- Malformed JSON: send one invalid
data:payload; assert the parser fails according to your contract rather than silently dropping it. - Slow delivery: delay between frames to exercise read timeouts, cancellation and back-pressure.
- Duplicate or out-of-order events: use these only if your client is required to defend against them; assert the documented behavior.
Test a browser or proxy conversion separately
If your server forwards provider output to a browser, write two tests. First, feed the upstream SSE fixture into the adapter and assert the internal event sequence. Second, feed that sequence into the browser endpoint and assert the exact NDJSON (or other application format) sent to the client. This catches a common defect: passing SSE text to code that expects newline-separated JSON. Include the terminal marker your browser client uses and close the response after it is written.
Assertions that make streaming tests reliable
- Assert each visible delta when progressive rendering is a requirement, not only the final concatenated string.
- Assert that completion is observed exactly once and that resources are closed afterward.
- Assert no text arrives after cancellation or an error.
- Use fake timers or controlled promises for delays; never depend on real network timing for ordinary unit tests.
- Give every fixture a request matcher so an accidental endpoint, method, or body change fails immediately.
- Run the same fixture repeatedly and reset the queue between tests.
Common failures and fixes
The test passes but the UI never renders partial text
Your fixture probably returns only a final message. Emit multiple deltas and assert the intermediate render calls.
The parser reports invalid JSON
Check the boundary format. SSE requires event and data lines separated by a blank line; an NDJSON reader expects one JSON object per line. Emit the format consumed by that layer.
The test waits forever
The fixture omitted the terminal completion event, never closed the response, or left a delayed promise unresolved. Add a completion assertion, close the server response, and make pending timers cancellable.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #3
Only the first consumer receives events
Raw streams are single-consumer. Use stream.tee() for two independent readers or place a replayable event buffer in front of them.
A Chat Completions fixture fails against a Responses consumer
Use the matching event model. Chat Completions uses chunks with delta; Responses uses named semantic events such as response.output_text.delta.
Retries duplicate visible text
Test retry behavior at the adapter boundary with a failed first response and a clean second response. Define whether the application clears partial text, resumes, or displays an explicit retry state; then assert that policy.
Performance, maintenance and cost considerations
In-process scripted models are fastest and easiest to run in every pull request. Controlled HTTP fixtures are slower but catch serialization and parser regressions. Keep event queues short, name each failure scenario, and avoid arbitrary sleeps except in tests explicitly covering timing. SDK-normalized fixtures usually survive provider wire changes better; wire fixtures provide higher fidelity but must be reviewed when API event names or SDK helpers change.
Use a small contract suite at the HTTP boundary and many focused workflow tests above it. That combination limits maintenance while still detecting broken framing, authentication, retries and disconnect handling.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your separate task is capturing a rendered test page rather than validating a model stream, ScreenshotNeo provides a one-request website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Use the API as documented at https://screenshotneo.com/docs/:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account to try it.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
FAQ
Should I mock the model, SDK, or HTTP request?
Mock the model for workflow behavior, the normalized stream for exact event behavior, and HTTP for wire, authentication, retry or network behavior.
Can one fixture cover Responses and Chat Completions?
No. Their streamed shapes differ, so maintain a fixture that matches each endpoint your code supports.
How do I simulate a disconnect?
Send valid frames, omit the terminal event, close the response, and assert that the consumer reports an incomplete stream and releases resources.
When should I use a delayed fixture?
Only when timing, cancellation, timeout, back-pressure or progressive rendering is the behavior under test; use immediate events elsewhere.
Frequently Asked Questions
What is the smallest useful streaming fixture?
A request-matched response that emits one or more deltas, a terminal completion event, and then closes. Add failures as separate named scenarios.
Are SSE and NDJSON interchangeable in tests?
No. SSE has event and data framing; NDJSON is one JSON object per line. Emit the representation expected by the layer under test.
How can I test cancellation deterministically?
Pause the fixture on a controllable promise, cancel the consumer, release the promise, and assert that no subsequent event changes state.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

