Designing an LLM Gateway
A full walkthrough for a company-wide LLM platform serving 300 internal apps, covering a unified API, key management, token quotas across stateless nodes, routing and failover, streaming at scale, redaction, cost attribution, and multi-region availability.
Lesson 6 of 10 15 min
1. Requirements
Functional
- One OpenAI-compatible API for all models: hosted providers plus self-hosted clusters
- Per-app credentials. Provider keys never leave the platform
- Per-team and per-app quotas in tokens per minute and dollars per month
- Model aliases, routing, retries and failover
- Logging and tracing with PII redaction, cost attribution, and dashboards per team
- Guardrail hooks (PII, injection screening) configurable per app
Non-functional
- 5,000 engineers and ≈ 300 apps, ≈ 2K QPS at peak, growing 3× a year
- Under 10 ms of added p99 latency, excluding guardrail model calls
- 99.99% availability, multi-region
- Streaming support for tens of thousands of concurrent streams
2. Operating point
Reliability and latency overhead are the constraints. Governance (cost, data) is the value. The gateway itself is cheap compared with the model spend flowing through it.
3. Estimates
| Quantity | Value |
|---|---|
| Peak QPS | 2K now, 6K next year |
| Concurrent streams | 2K × ≈ 10 s ≈ 20K (60K next year) |
| Usage events | 2K per second, ≈ 170M a day |
| Log content (sampled 10%, ≈ 20 KB per request) | ≈ 350 GB a day before redaction and compression |
| Model spend under management | Often $1–10M+ a month, so a few ms of gateway overhead is irrelevant in cost terms, but not in latency |
4. Architecture
Apps / SDK
OpenAI-compatible
Global LB
nearest healthy region
Gateway nodes
stateless, async I/O
Policy
auth, quota, guardrails
Router
alias → provider/model
Providers + self-hosted
Side systems
- Config service: apps, teams, aliases, budgets and routing weights, pushed to nodes with hot reload (versioned, auditable).
- Quota store: Redis cluster per region.
- Usage pipeline: gateway emits events to Kafka, which feed a warehouse (cost attribution, dashboards), anomaly alerts and eval-data sampling.
- Tracing: OpenTelemetry spans exported asynchronously.
- Secrets manager for provider keys, rotated centrally.
5. Deep dives
Deep dive A: token quotas across stateless nodes
A central counter hit on every request adds latency and becomes a hotspot. The design:
- Estimate the request's tokens (input count plus
max_tokens). - Reserve atomically in Redis (a Lua script over a sliding-window counter per app and per team).
- On completion, reconcile with the actual usage from the provider's final usage event.
- Local leasing for high-QPS apps: a node leases a chunk of budget (for example 50K tokens) and spends it locally, which cuts Redis round trips at the cost of slight over-admission.
- Fail open or closed by tier: if Redis is down, critical apps fail open with local limits, and batch apps fail closed.
Monthly dollar budgets are computed from usage events (token × price) with alerts at 80% and 100%, and optional hard stops or downgrades.
Deep dive B: routing and failover
- Aliases (
chat-default,extract-cheap) map to ordered candidate lists with weights and constraints (region, data policy). - Health scoring from live error rates, 429s and TTFT per provider and region. Circuit breakers open on sustained failures, with half-open probing.
- Retries only before the first streamed token, with backoff and jitter, respecting
Retry-After. - Failover order: same model in another region, then an eval-approved equivalent, then degraded mode.
- Data policy constraints: apps tagged "EU data" can only route to EU endpoints, including when failing over.
Deep dive C: streaming at scale
- Async, event-driven nodes (Go, Rust or Envoy-based) handle tens of thousands of concurrent streams per node with low memory per connection.
- Pass through chunks immediately. Count tokens from the final usage event, or tokenize incrementally if the provider doesn't send one.
- Client disconnect cancels the upstream request.
- Idle and total timeouts configured per alias. Load balancers tuned for long-lived connections.
6. Security, privacy and governance
- Redaction pipeline: PII detection before logging, and optionally before sending to providers for apps that require it.
- Access control on logs by team, with retention limits.
- Audit log of config changes (who changed routing or budgets).
- Guardrail hooks as plugins in the policy stage, run in parallel where possible, with latency budgets.
7. Availability
- Multi-region active-active. Nodes are stateless, so any region can serve any app.
- Redis per region, and quotas tolerate brief inconsistency.
- Bypass plan: documented emergency direct-provider credentials for tier-0 apps, used only in a gateway-wide incident.
- Canary deploys, config validation, and fast rollback. A bad routing config is a company-wide outage.
8. Bottlenecks and follow-ups
| Likely question | Answer sketch |
|---|---|
| "One team's batch job causes 429s for everyone?" | Per-team token budgets, priority lanes, batch traffic through a queue or the provider batch API |
| "Chargeback?" | Usage events tagged with app, team and feature, priced per model version including cached tokens, into the warehouse |
| "Add semantic caching?" | An opt-in per app for low-risk intents, tenant-scoped keys, with hit-rate and quality monitoring |
| "Gateway adds 40 ms?" | Profile. Move logging and cost computation async, use local quota leases, co-locate with regions |
Key takeaways
- The gateway is tier-0 infrastructure. Every AI feature depends on it, so design for high availability, low overhead and a bypass plan.
- Enforce token budgets across stateless nodes with reservations in a shared store, plus local caching to keep latency low.
- Model aliases, health-based routing and circuit breakers give swappable, resilient providers.
- Usage events go to an async pipeline for cost attribution, analytics and eval data, never on the critical path.
Go deeper
Finished reading? Mark it done to track your progress.