Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To stream an AI answer from FastAPI through Next.js, have FastAPI return a StreamingResponse built from an async generator, have a Next.js App Router Route Handler fetch that endpoint and return its body unchanged, parse the stream in the browser with an explicit message format, and turn off buffering in every proxy between the two. Each of those steps has a specific failure mode, and most “the whole answer arrives at once” reports trace back to one of them.

This guide follows the bytes across the full chain: model client, FastAPI, Next.js, reverse proxy, host, and browser. It covers four details that cause most of the trouble: choosing the right Next.js layer, keeping the stream intact through the proxy path, finding the buffers that sit between the hops, and handling headers, errors, and cancellation.

Which Next.js layer should own the stream

The phrase “Next.js proxy” causes confusion because Next.js has a feature called proxy.ts that is unrelated to forwarding a streamed backend response. In current Next.js releases that file runs before routing and is meant for routing decisions and request or response adjustments; earlier releases used middleware.ts for the same role. The backend stream belongs in a Route Handler, such as app/api/chat/route.ts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Concern Route Handler (app/api/chat/route.ts) proxy.ts
When it runs When a request matches its route and method Before routing, for matching requests
Typical job Validate the chat request, call FastAPI, return the response body Redirects, rewrites, and light header or cookie changes
Waiting on a backend Appropriate: it awaits the backend and returns the result Next.js states Proxy is not intended for slow data fetching
Returning a streamed body Yes, through a standard Web Response that wraps a ReadableStream Not the place for the chat response

Put simply, the Route Handler is the HTTP endpoint your browser talks to, and proxy.ts is a pre-route hook you may or may not need at all. The Route Handler conventions page (updated April 30, 2026) documents Web Request and Response use and returning a ReadableStream, and the Backend for Frontend guide (updated March 25, 2026) shows Route Handlers validating a request before proxying it to another backend.

Step 1: FastAPI yields encoded pieces, not a finished answer

FastAPI’s StreamingResponse accepts an async generator or a regular generator and sends the body as the iterator yields. Its documentation says yielded chunks are sent as they are and are not converted to JSON, so the application owns the encoding. That means a generator that yields Python dictionaries will not produce JSON lines on its own. Serialize each event to bytes inside the generator.

n# main.pynimport jsonnfrom fastapi import FastAPInfrom fastapi.responses import StreamingResponsenfrom pydantic import BaseModelnfrom model_client import stream_completion  # your async model clientnnapp = FastAPI()nnclass ChatRequest(BaseModel):n    prompt: strnndef encode(event: dict) -> bytes:n    return (json.dumps(event, ensure_ascii=False) + '\n').encode('utf-8')nnasync def answer_events(prompt: str):n    try:n        async for piece in stream_completion(prompt):n            yield encode({'type': 'token', 'text': piece})n        yield encode({'type': 'done'})n    except Exception:n        # The 200 status and headers are already sent by this point.n        yield encode({'type': 'error', 'message': 'generation failed'})nn@app.post('/chat/stream')nasync def chat_stream(body: ChatRequest):n    return StreamingResponse(n        answer_events(body.prompt),n        media_type='application/x-ndjson',n        headers={'Cache-Control': 'no-cache'},n    )n

Two details matter here. First, the generator must await the model client’s reads. FastAPI’s documentation notes that a coroutine only sees cancellation at an await point, so an async for over a real async client is the right shape. A generator that loops over synchronous work with no awaits will keep running after the client has left. Second, the except Exception block does not catch asyncio.CancelledError, which is a BaseException in modern Python, so cancellation still propagates. Add a finally block that closes the upstream model client if your SDK requires it.

Step 2: Next.js forwards the body without reading it

The Route Handler should validate the request, call FastAPI, check the upstream status, and hand the upstream ReadableStream straight to a new Response. It should not call text() or json() on the upstream body, and it should not collect chunks into a string, because either one turns a stream into a finished answer before the browser sees anything.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
n// app/api/chat/route.tsnconst FASTAPI_URL = process.env.FASTAPI_URL; // server-only, e.g. http://127.0.0.1:8000nnfunction readPrompt(body: unknown): string | null {n  if (typeof body === 'object' && body !== null && 'prompt' in body) {n    const prompt = (body as { prompt: unknown }).prompt;n    if (typeof prompt === 'string' && prompt.trim().length > 0 && prompt.length <= 4000) {n      return prompt;n    }n  }n  return null;n}nnexport async function POST(request: Request) {n  if (!FASTAPI_URL) {n    return Response.json({ error: 'FASTAPI_URL is not set' }, { status: 500 });n  }nn  let body: unknown;n  try {n    body = await request.json();n  } catch {n    return Response.json({ error: 'Body must be JSON' }, { status: 400 });n  }nn  const prompt = readPrompt(body);n  if (prompt === null) {n    return Response.json({ error: 'prompt must be a non-empty string' }, { status: 400 });n  }nn  let upstream: Response;n  try {n    upstream = await fetch(`${FASTAPI_URL}/chat/stream`, {n      method: 'POST',n      headers: {n        'Content-Type': 'application/json',n        Accept: 'application/x-ndjson',n      },n      body: JSON.stringify({ prompt }),n      signal: request.signal,n      cache: 'no-store',n    });n  } catch {n    return Response.json({ error: 'Upstream unreachable' }, { status: 502 });n  }nn  if (!upstream.ok || upstream.body === null) {n    return Response.json({ error: 'Upstream failed' }, { status: 502 });n  }nn  return new Response(upstream.body, {n    status: 200,n    headers: {n      'Content-Type': 'application/x-ndjson; charset=utf-8',n      'Cache-Control': 'no-cache, no-transform',n      'X-Accel-Buffering': 'no',n    },n  });n}n

The signal: request.signal line ties the upstream request to the browser’s connection, which is the basis for the cancellation behavior discussed below. The 4000-character limit is an example value; set it to suit your model’s input budget.

Step 3: Choose a framing format and parse it deliberately

A browser reader receives arbitrary byte chunks. One network chunk may contain half a message, three messages, or a single token split across two chunks. Your protocol has to say where one message ends. The three common options differ in how much they ask of the client.

Framing What the client does Strength Limitation
Plain text chunks Appends decoded text to the message Simplest to build No room for end markers, errors, or metadata
Server-Sent Events Parses event: and data: fields Standard event names and retry semantics The browser EventSource API only issues GET requests, so a POST chat call needs a manual parser
Newline-delimited JSON (NDJSON) Splits on newlines and calls JSON.parse per line Typed events, explicit done and error markers, room for metadata Each line must be complete, so partial lines must be buffered

The examples in this guide use NDJSON with three event types: token, done, and error. The browser parser treats a stream that closes without a done event as incomplete.

n// lib/stream-chat.tsnexport async function streamAnswer(n  prompt: string,n  onToken: (text: string) => void,n): Promise<void> {n  const res = await fetch('/api/chat', {n    method: 'POST',n    headers: { 'Content-Type': 'application/json' },n    body: JSON.stringify({ prompt }),n  });n  if (!res.ok || res.body === null) {n    throw new Error(`Chat request failed with HTTP ${res.status}`);n  }nn  const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();n  let buffer = '';n  let finished = false;nn  const handleLine = (line: string) => {n    if (line.trim() === '') return;n    const event = JSON.parse(line);n    if (event.type === 'token') onToken(event.text);n    else if (event.type === 'done') finished = true;n    else if (event.type === 'error') throw new Error(event.message);n  };nn  while (true) {n    const { value, done } = await reader.read();n    if (done) break;n    buffer += value;n    let newline = buffer.indexOf('\n');n    while (newline !== -1) {n      handleLine(buffer.slice(0, newline));n      buffer = buffer.slice(newline + 1);n      newline = buffer.indexOf('\n');n    }n  }n  handleLine(buffer);n  if (!finished) throw new Error('Stream ended before the done event');n}n

Call streamAnswer from a React component and append each token to state. Appending to state on every token is usually fine for chat-length output; if rendering becomes the bottleneck, batch updates with requestAnimationFrame.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Headers and status codes

  • Set Content-Type explicitly. The Route Handler above sets application/x-ndjson; charset=utf-8 rather than copying whatever the upstream returned.
  • Set Cache-Control: no-cache, no-transform. This asks shared caches and intermediaries not to store or rewrite a live answer.
  • Set X-Accel-Buffering: no on the Next.js response. The Next.js self-hosting guide (updated October 1, 2026) describes this header as the way to ask nginx not to buffer the response. It is ignored by proxies that do not implement it, so it is not a substitute for the configuration in the buffering section below.
  • Forward only the request headers the backend needs. Forward Content-Type and Accept, plus any authentication header your backend requires, which you should derive on the server. Do not copy the browser’s headers wholesale. The Next.js NextResponse documentation (updated March 25, 2026) warns that indiscriminate header forwarding can leak secrets and that inappropriate response headers can break framework behavior, including streaming.
  • Leave connection-management headers to the runtime. Connection, Transfer-Encoding, and Content-Length describe the hop, not the content, and the server or runtime manages them. Do not set them by hand to make streaming work.

Errors before and after the first byte

An HTTP status code can change only until the response starts. Once the 200 headers and first body bytes have been sent, the only way to report a failure to the browser is inside the message stream.

When the failure happens What the browser sees Where it is handled
Invalid request body HTTP 400 with a JSON error Next.js Route Handler, before fetching FastAPI
FastAPI unreachable, or returns a non-2xx status HTTP 502 with a JSON error Next.js Route Handler, before returning the stream
Model fails after the first token HTTP 200, then an error line, then the stream closes FastAPI generator emits the error event; the client throws
Connection drops mid-answer The reader ends or throws without a done event Client treats the answer as incomplete

A failure that happens before the first token can still return a real status code. To get that, pull the first event from the model before constructing the StreamingResponse. If that first read raises, FastAPI can return a 502 instead of a 200 with an error line. The trade-off is that the first token’s latency is now paid before any headers are sent, which is usually acceptable.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Cancellation and host timeouts

A chat stream can be abandoned in two ways: the user closes the tab, or the host stops the function. Each needs a different check.

  • Browser to Next.js. Confirm that request.signal fires when the browser disconnects in your Next.js version and runtime. If it does, the abort propagates through the fetch call to FastAPI, and the next await in the generator raises, which lets the model client close. If it does not fire, the upstream request keeps generating after nobody is listening. Test this with a deliberate tab close and watch the FastAPI logs.
  • Next.js to FastAPI. When the abort reaches FastAPI, the generator should exit at its next await and run its cleanup. A generator that is busy in synchronous code will not notice, which is why the await points in Step 1 matter.
  • Host timeouts. The Next.js Backend for Frontend guide (updated March 25, 2026) notes that on some function-style hosts, long-running handlers can be terminated when a timeout is reached. A chat answer that takes longer than the platform’s maximum duration will be cut off regardless of the code. Check the maximum duration for your provider’s current plan in its live documentation, because these limits change independently of the framework docs.

Finding the invisible buffer

If the code is correct and the browser still receives everything at once, the buffer is in a hop you did not write. Check the hops in order and stop at the first one that delays chunks.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. The generator. Log a timestamp each time the generator yields. If the timestamps are spread out, FastAPI is producing progressively.
  2. FastAPI directly. Run curl -N -X POST http://127.0.0.1:8000/chat/stream -H 'Content-Type: application/json' -d '{"prompt":"hello"}'. The -N flag disables curl’s own buffering. Lines should appear over time.
  3. Next.js. Run the same request against http://localhost:3000/api/chat. If FastAPI streams but Next.js does not, the problem is in the Route Handler, usually a body read somewhere that collects chunks.
  4. The reverse proxy. The Next.js self-hosting guide says that if you use nginx or a similar proxy, you must configure it to disable buffering. For nginx, the relevant location block looks like this:
nlocation /api/ {n    proxy_pass http://127.0.0.1:3000;n    proxy_http_version 1.1;n    proxy_buffering off;n}n
  1. The load balancer, CDN, or host platform. The same Next.js guide warns that load balancers and reverse proxies may buffer chunked output, and that some load-balancer integrations buffer by default. Look for response buffering, compression, or caching settings on the platform, and confirm them in its current documentation.
  2. The browser. In DevTools, the Network panel’s entry for /api/chat should stay pending while tokens arrive. If it shows the full body only at completion, one of the hops above is buffering.

Work through the hops in this order even when the cause seems obvious. A buffer at the edge is the most common cause, and it is easy to blame the application code for it.

Verifying the stream in your own stack

To confirm progressive delivery in a real environment, add a deliberate delay inside the generator loop, such as await asyncio.sleep(0.5) after each yield. Then record the time each token reaches the browser, either by logging inside the onToken callback or by reading the Network timing. Compare the spacing of the browser timestamps with the spacing of the generator timestamps. Matching gaps mean the stream is intact end to end. A single burst at the end means a hop is buffering.

Run this check locally and again on the deployed stack. Local success does not prove that your production proxy, CDN, or host passes chunks through, because each of those can be configured differently. Pin the versions of FastAPI, Starlette, Next.js, and your Node runtime that you test, since the handling of response bodies and cancellation can change between releases.

n

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.