Streaming Responses with SSE
How token streaming works end to end, why server-sent events are the default over WebSockets, and the infrastructure details (proxies, timeouts, cancellation, resumption) that break in production.
Why stream
A 600-token answer at 60 tokens/s takes 10 seconds. Without streaming, the user stares at a spinner for 10 seconds. With streaming, text appears after about 0.5 s and flows faster than they can read. Streaming is the difference between an AI feature that feels slow and one that feels responsive (see TTFT and TPOT).
Server-sent events
SSE is a simple standard: an HTTP response with Content-Type: text/event-stream that stays open while the server writes events:
event: token
data: {"text": "The refund"}
event: token
data: {"text": " window is"}
event: done
data: {"usage": {"input_tokens": 3120, "output_tokens": 212}}| SSE | WebSockets | Long polling | |
|---|---|---|---|
| Direction | Server → client | Both ways | Server → client (per request) |
| Protocol | Plain HTTP | Upgrade to WS | Plain HTTP |
| Proxy / CDN friendliness | Good, if buffering is disabled | Needs WS support | Good |
| Reconnect | Built into the browser EventSource | Manual | Manual |
| Best for | LLM text output | Voice, collaborative editing, bidirectional agents | Legacy fallback |
Most LLM APIs stream with SSE, and most applications relay it the same way. Pick WebSockets or WebRTC for real-time voice, where audio flows both ways continuously.
The end-to-end path
Every hop must flush each chunk immediately:
- Reverse proxies (Nginx and others) buffer responses by default. Disable proxy buffering for stream routes.
- Load balancers and CDNs: raise idle timeouts well above your longest response (60 s defaults kill long generations), and make sure streaming is supported on that route.
- Serverless platforms: check they support streamed responses and long enough execution times.
- Compression middleware can buffer chunks until a block fills up. Disable or configure it for SSE.
Cancellation
If the user closes the tab or presses Stop, propagate the cancel upstream: abort the provider request or free the inference slot. Otherwise the model keeps generating tokens nobody reads, and you pay for them. On self-hosted engines, cancelled requests also release KV cache for other users.
Resumption
Networks drop, especially on mobile:
- Give each response an ID and number the events, so the client can reconnect and ask to resume from the last event it received. SSE's
Last-Event-IDsupports this. - To resume, the backend must keep generating or buffer the output server-side, decoupled from the client connection. Store the partial output in a fast store while the generation runs.
- For long agent tasks, run generation as a background job and let clients subscribe to its progress. A dropped connection then doesn't kill the task.
Streaming structured output and tools
- Tool calls stream too. Arguments arrive as partial JSON. Buffer them until complete before executing.
- Partial JSON can be parsed incrementally to render progressive UIs, such as a form filling in field by field.
- Guardrails on streamed output check chunks as they go, or use a short hold-back buffer. Full-answer moderation contradicts streaming, so decide which risks justify delaying output.
Key takeaways
- Streaming sends tokens as they are generated, which cuts perceived latency from total time to time to first token.
- Server-sent events (SSE) are one-way HTTP streams and the standard for LLM output. Use WebSockets when you need two-way real-time traffic, such as voice.
- Every hop must support streaming, including load balancers, proxies, CDNs and serverless platforms. Buffering anywhere silently breaks it.
- Handle client disconnects by cancelling the upstream generation, and plan for resuming interrupted streams.