The LLM Gateway Pattern
Why every company using LLMs at scale ends up with a gateway, what it does (auth, routing, quotas, caching, logging, cost attribution), and how to design one.
Why a gateway
Without one, every team calls providers directly with its own API keys, SDKs, retry logic and logging. Six months later nobody knows the monthly spend per product, a leaked key is shared by ten services, and switching providers means changing forty codebases.
An LLM gateway centralises this:
It is the same idea as an API gateway, specialised for things classic gateways don't understand: tokens, streaming, model names and prompt content.
Core responsibilities
| Responsibility | What it does |
|---|---|
| Unified API | One request format (often OpenAI-compatible) translated to each provider's API |
| Auth and key management | Apps authenticate to the gateway. Provider keys live only in the gateway, rotated centrally |
| Routing | Choose model or provider per request: by alias, cost, latency, region, or a classifier (next lesson) |
| Fallbacks and retries | Retry transient errors, fail over to another provider or region |
| Quotas and rate limits | Per team, app or user budgets in tokens and dollars, not just requests (rate limiting) |
| Caching | Response caching, semantic caching, and help with provider prompt caching (caching) |
| Observability | Log model, tokens, latency (TTFT, total), cost, status and trace IDs per request |
| Cost attribution | Tag every request with team, product and feature, for chargeback and budgeting |
| Guardrails | PII redaction, prompt-injection screening, output moderation, data-residency rules |
Design considerations
Latency. The gateway is on every request's critical path. Target single-digit milliseconds of overhead. Do heavy work, such as detailed logging or cost computation, asynchronously after the response starts.
Streaming. The gateway must pass SSE streams through token by token, never buffering the whole response. Count tokens from the stream's final usage event, or by tokenizing incrementally, to bill and enforce quotas.
Statelessness and scale. Keep gateway nodes stateless, with quota counters in Redis or similar. Deploy multi-zone. A gateway outage takes down every AI feature, so it needs the availability of your most critical service, and a bypass plan.
Model aliases. Apps request chat-default or summarise-cheap, not a raw model ID. The platform team can upgrade the model behind an alias after evals pass, without code changes in every app.
Data governance. Decide centrally what gets logged. Full prompts are valuable for debugging and evals, but they contain user data, so apply redaction, retention limits and access control.
Build or buy
Open-source and commercial gateways exist (LiteLLM, Portkey, Kong and Envoy AI gateway plugins, cloud provider gateways). Many companies start with one and extend it. The design questions are the same either way. In an interview, describe the responsibilities and the data flow, not a product name.
Key takeaways
- An LLM gateway is a reverse proxy between your applications and model providers. It gives you one API, one place for policy, and one source of usage data.
- Its core jobs are auth, routing and fallbacks, token-based quotas, caching, logging and tracing, cost attribution, and guardrails.
- Keep the gateway thin and fast. It sits on the critical path of every LLM call and must stream without buffering.
- The gateway is what makes providers swappable and costs visible.