GenAI System Design
6. Integrating LLMs into Existing Systems

Streaming Responses with SSE

How token streaming works end to end, why server-sent events are the default over WebSockets, and the infrastructure details (proxies, timeouts, cancellation, resumption) that break in production.

Lesson 4 of 6 9 min

Why stream

A 600-token answer at 60 tokens/s takes 10 seconds. Without streaming, the user stares at a spinner for 10 seconds. With streaming, text appears after about 0.5 s and flows faster than they can read. Streaming is the difference between an AI feature that feels slow and one that feels responsive (see TTFT and TPOT).

Server-sent events

SSE is a simple standard: an HTTP response with Content-Type: text/event-stream that stays open while the server writes events:

event: token data: {"text": "The refund"} event: token data: {"text": " window is"} event: done data: {"usage": {"input_tokens": 3120, "output_tokens": 212}}
SSEWebSocketsLong polling
DirectionServer → clientBoth waysServer → client (per request)
ProtocolPlain HTTPUpgrade to WSPlain HTTP
Proxy / CDN friendlinessGood, if buffering is disabledNeeds WS supportGood
ReconnectBuilt into the browser EventSourceManualManual
Best forLLM text outputVoice, collaborative editing, bidirectional agentsLegacy fallback

Most LLM APIs stream with SSE, and most applications relay it the same way. Pick WebSockets or WebRTC for real-time voice, where audio flows both ways continuously.

The end-to-end path

Model / provider
SSE
Your gateway
pass through, count tokens
App backend
add citations, guardrails
Load balancer / CDN
no buffering!
Browser
render incrementally

Every hop must flush each chunk immediately:

  • Reverse proxies (Nginx and others) buffer responses by default. Disable proxy buffering for stream routes.
  • Load balancers and CDNs: raise idle timeouts well above your longest response (60 s defaults kill long generations), and make sure streaming is supported on that route.
  • Serverless platforms: check they support streamed responses and long enough execution times.
  • Compression middleware can buffer chunks until a block fills up. Disable or configure it for SSE.

Cancellation

If the user closes the tab or presses Stop, propagate the cancel upstream: abort the provider request or free the inference slot. Otherwise the model keeps generating tokens nobody reads, and you pay for them. On self-hosted engines, cancelled requests also release KV cache for other users.

Resumption

Networks drop, especially on mobile:

  • Give each response an ID and number the events, so the client can reconnect and ask to resume from the last event it received. SSE's Last-Event-ID supports this.
  • To resume, the backend must keep generating or buffer the output server-side, decoupled from the client connection. Store the partial output in a fast store while the generation runs.
  • For long agent tasks, run generation as a background job and let clients subscribe to its progress. A dropped connection then doesn't kill the task.

Streaming structured output and tools

  • Tool calls stream too. Arguments arrive as partial JSON. Buffer them until complete before executing.
  • Partial JSON can be parsed incrementally to render progressive UIs, such as a form filling in field by field.
  • Guardrails on streamed output check chunks as they go, or use a short hold-back buffer. Full-answer moderation contradicts streaming, so decide which risks justify delaying output.

Key takeaways

  • Streaming sends tokens as they are generated, which cuts perceived latency from total time to time to first token.
  • Server-sent events (SSE) are one-way HTTP streams and the standard for LLM output. Use WebSockets when you need two-way real-time traffic, such as voice.
  • Every hop must support streaming, including load balancers, proxies, CDNs and serverless platforms. Buffering anywhere silently breaks it.
  • Handle client disconnects by cancelling the upstream generation, and plan for resuming interrupted streams.

Go deeper

Finished reading? Mark it done to track your progress.