GenAI System Design
6. Integrating LLMs into Existing Systems

The LLM Gateway Pattern

Why every company using LLMs at scale ends up with a gateway, what it does (auth, routing, quotas, caching, logging, cost attribution), and how to design one.

Lesson 1 of 6 9 min

Why a gateway

Without one, every team calls providers directly with its own API keys, SDKs, retry logic and logging. Six months later nobody knows the monthly spend per product, a leaked key is shared by ten services, and switching providers means changing forty codebases.

An LLM gateway centralises this:

Apps and agents
one internal API
LLM gateway
auth, quotas, routing, cache, logs
Provider A
Provider B
Self-hosted models
vLLM clusters

It is the same idea as an API gateway, specialised for things classic gateways don't understand: tokens, streaming, model names and prompt content.

Core responsibilities

ResponsibilityWhat it does
Unified APIOne request format (often OpenAI-compatible) translated to each provider's API
Auth and key managementApps authenticate to the gateway. Provider keys live only in the gateway, rotated centrally
RoutingChoose model or provider per request: by alias, cost, latency, region, or a classifier (next lesson)
Fallbacks and retriesRetry transient errors, fail over to another provider or region
Quotas and rate limitsPer team, app or user budgets in tokens and dollars, not just requests (rate limiting)
CachingResponse caching, semantic caching, and help with provider prompt caching (caching)
ObservabilityLog model, tokens, latency (TTFT, total), cost, status and trace IDs per request
Cost attributionTag every request with team, product and feature, for chargeback and budgeting
GuardrailsPII redaction, prompt-injection screening, output moderation, data-residency rules

Design considerations

Latency. The gateway is on every request's critical path. Target single-digit milliseconds of overhead. Do heavy work, such as detailed logging or cost computation, asynchronously after the response starts.

Streaming. The gateway must pass SSE streams through token by token, never buffering the whole response. Count tokens from the stream's final usage event, or by tokenizing incrementally, to bill and enforce quotas.

Statelessness and scale. Keep gateway nodes stateless, with quota counters in Redis or similar. Deploy multi-zone. A gateway outage takes down every AI feature, so it needs the availability of your most critical service, and a bypass plan.

Model aliases. Apps request chat-default or summarise-cheap, not a raw model ID. The platform team can upgrade the model behind an alias after evals pass, without code changes in every app.

Data governance. Decide centrally what gets logged. Full prompts are valuable for debugging and evals, but they contain user data, so apply redaction, retention limits and access control.

Build or buy

Open-source and commercial gateways exist (LiteLLM, Portkey, Kong and Envoy AI gateway plugins, cloud provider gateways). Many companies start with one and extend it. The design questions are the same either way. In an interview, describe the responsibilities and the data flow, not a product name.

Key takeaways

  • An LLM gateway is a reverse proxy between your applications and model providers. It gives you one API, one place for policy, and one source of usage data.
  • Its core jobs are auth, routing and fallbacks, token-based quotas, caching, logging and tracing, cost attribution, and guardrails.
  • Keep the gateway thin and fast. It sits on the critical path of every LLM call and must stream without buffering.
  • The gateway is what makes providers swappable and costs visible.

Go deeper

Finished reading? Mark it done to track your progress.