Short answer: a 2026 benchmark reported that HAProxy 3.4.4 delivered small, regularly emitted SSE frames in bursts, with a 206 ms time to first token under one specific test condition. That is not a universal HAProxy delay or an official product statistic. The benchmark did not establish the exact flush trigger; a separate personal post by an HAProxy employee attributes the behavior to kernel batching, but the benchmark had not tested the suggested option http-no-delay setting.
What does the 206 ms figure mean?
Remdore’s 2026 first-person benchmark reports about 206 ms to first token for HAProxy 3.4.4 when its synthetic emitter sent roughly 60-byte SSE frames every 50 ms. It also reports a 0.0 ms frame gap and 5.12 frames per client read for that condition. Those measurements describe the tested setup, not a general property of HAProxy or a guaranteed delay in other deployments. Read the benchmark report.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
HAProxy In-Depth: Definitive Reference for Developers and Engineers | $9.95 | Buy on Amazon |
The benchmark checked that the synthetic emitter met its intended pacing before accepting a run, measured a direct path in the same cell, and used raw sockets on the client to avoid client-library buffering. The author says the published comparisons used a 2-vCPU Ubuntu 24.04 droplet after macOS with Docker Desktop showed scheduling and network artifacts. These details strengthen the report’s account of its own controlled setup, but it is still a first-person experiment, not an independently audited reproduction.
Why might streamed tokens arrive in bursts?
Server-sent events (SSE) can deliver small chunks frequently. A proxy and the network stack sit between the application writing those chunks and the client reading them. If writes are combined before delivery, the client can receive several frames in one read rather than one frame at a time. That pattern is consistent with the benchmark’s frames-per-read result, but the exact mechanism and trigger were not resolved there.
Ron Northcutt, identified in his post as an HAProxy employee sharing his personal view rather than an official statement, describes the behavior this way: “HAProxy isn’t buffering tokens. By default, it asks the kernel to batch response body writes into full packets.” Read Northcutt’s post. This is an attributed explanation, not a finding independently demonstrated by the benchmark.
How frame size and timing changed the reported result
The same benchmark found substantially different behavior when it changed the workload. Its reported results were:
| Test condition | Reported HAProxy result |
|---|---|
| Roughly 60-byte frames emitted every 50 ms | About 206 ms to first token; 0.0 ms reported frame gap; 5.12 frames per read |
| Roughly 1.1 KB frames | About 53 ms to first token; 1.46 frames per read |
| Roughly 60-byte frames emitted every 5 ms | 13.67 frames per read |
| Roughly 60-byte frames emitted every 200 ms | 2.05 frames per read |
These are results from the report’s particular experiment, not general performance figures. The larger-frame result is consistent with an accumulation or flush effect, but the author declined to name a byte threshold because the inferred values did not reconcile. The varying frames-per-read ratios also show why a comparison needs to hold frame size and emission interval constant.
Would option http-no-delay fix it?
Northcutt’s personal post points readers toward option http-no-delay. However, Remdore’s benchmark lists testing that option and its cost as future work, so the cited result does not demonstrate that it fixes the delay or establish what trade-off it brings. Treat it as a setting to investigate in a controlled test, not as a proven remedy.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →The report also says that X-Accel-Buffering: no did not materially change its HAProxy result. It describes that header as an nginx convention, not a HAProxy configuration control; adding it should not be treated as a demonstrated fix for this HAProxy behavior.
What the nginx comparison does—and does not—show
The benchmark found no measurable effect from nginx’s proxy_buffering off in its small-payload conditions when the client read promptly. In a separate test with about 328 KB of payload and a client sleeping 200 ms between reads, it reports 53 ms to first token with buffering on and 3 ms with buffering off. These results concern different payload and client-read conditions; they do not establish that disabling nginx buffering always changes token latency, nor do they isolate a universal difference between nginx and HAProxy.
How much should you infer from the real-model check?
Remdore also reports 15 calls to a DigitalOcean serverless inference endpoint, all returning HTTP 200. In that limited comparison, it measured 5.43 frames per read through HAProxy and 1.04 through nginx. The author cautions that model timing varies, so this small comparison does not by itself establish causation or show that HAProxy caused a particular user-visible delay.
How to test your own streaming path
For a useful comparison, keep the workload and observation method the same across proxy configurations. The benchmark’s own results indicate that changing frame size, emission interval, or client read behavior can change what the client sees.
- Record the proxy version and relevant configuration for each run.
- Keep SSE frame size and emission interval constant; report both.
- Use the same client and read behavior for every path, and note whether reads are prompt or delayed.
- Measure first-token time, inter-frame gaps, and frames received per client read rather than relying on a single latency number.
- Compare against a direct path under the same conditions, and repeat runs when upstream model timing varies.
- If testing
option http-no-delay, change only that setting between otherwise matched runs and measure both its effect and any cost relevant to your workload.
The exact flush trigger, the effect and cost of option http-no-delay, and explicit upstream keepalive configuration for Caddy and Traefik were listed as unresolved work in the benchmark. Its findings should therefore be used to frame a local test, not as a substitute for one.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

