GenAI System Design
6. Integrating LLMs into Existing Systems

Rate Limiting by Tokens

Why request-count limits fail for LLM traffic, how to meter tokens with token buckets or sliding windows, estimating tokens before a call, and fair sharing of provider quotas across tenants.

Lesson 3 of 6 10 min

Why requests are the wrong unit

One request might be a 50-token classification and another a 150K-token document analysis, a 3,000× difference in cost and provider capacity. A limit of "60 requests per minute per user" lets a single user send 60 huge requests and use up your entire provider quota.

Providers limit you on requests per minute (RPM) and tokens per minute (TPM), per model, for your whole organisation. Your internal limits must mirror that.

Rate limiting by requests vs by tokens

Two tenants share one provider account limited to 150K tokens per minute. Tenant B starts a bulk job at second 10.

Tenant A (chat, steady)
Tenant B (bulk document job)
served 429 from your gateway (that tenant only) 429 from the provider (hits everyone)
Provider tokens used in the last 60 s (limit 150K)
074.5K149K03059second
Tenant A requests served
27%
Tenant B requests served
17%

Limiting by request count lets B's few huge requests use up the shared provider quota, so A's small chat requests start failing: a noisy neighbour. Budgeting tokens per tenant at the gateway keeps B inside its share, and A is never affected.

Algorithms

Token bucket. Each tenant has a bucket with capacity C tokens, refilled at R tokens per second. A request is allowed if the bucket holds enough for it, and then the amount is subtracted. It allows short bursts up to C while enforcing an average rate R.

Sliding window counter. Sum the tokens used in the last 60 seconds and allow the request if sum + estimate ≤ limit. It is closer to how providers count, and easy to implement in Redis with sorted sets or bucketed counters.

Either works. What matters is what you count and where you store it: a shared, fast store such as Redis, updated atomically with Lua scripts or atomic increments, because gateway nodes are stateless.

Estimating before the call

You don't know the true token count until the response finishes. Handle it in two phases:

  1. Reserve before the call: count input tokens (tokenizer, or chars ÷ 4 plus a margin) plus max_tokens for the output.
  2. Reconcile after the call: refund the difference between the reservation and the actual usage reported by the provider.

Setting max_tokens sensibly matters twice: it bounds cost, and it keeps reservations from being wildly pessimistic.

Fair sharing across tenants

With a fixed provider quota shared by many tenants or teams:

  • Per-tenant TPM budgets, summing to at most the provider limit (or over-committed carefully, with a global check as well).
  • Priority tiers: interactive traffic gets guaranteed capacity, and batch or background jobs use what's left. A background job should never cause 429s for a user-facing chat.
  • Queue instead of reject for batch work. Hold requests until budget frees up, rather than failing them.
  • Weighted fair queuing when demand exceeds supply, so each tenant gets its share and no single tenant starves the others.

Long-window budgets

Short windows protect provider limits. Long windows protect your wallet:

  • Daily or monthly spend caps per tenant, team or API key, in dollars (tokens × price, with cached tokens priced accordingly).
  • Alerts at 50%, 80% and 100% of budget, then either a hard stop or a downgrade to a cheaper model.
  • Per-user caps on consumer products, to prevent abuse such as scripts farming your free tier.

Telling clients what happened

  • Return 429 with Retry-After, and headers showing remaining tokens and requests (as providers do).
  • Distinguish "you exceeded your quota" from "we are over capacity", because clients should react differently.
  • Give internal SDKs automatic backoff that honours these headers.

Key takeaways

  • LLM requests vary 1,000× in cost, so limit tokens (and dollars) per minute, not just requests.
  • Providers enforce organisation-wide tokens-per-minute limits. Without per-tenant budgets, one heavy user causes 429s for everyone.
  • Estimate input tokens before the call and reserve the maximum output, then reconcile with actual usage afterwards.
  • Combine short-window rate limits (TPM) with long-window budgets (daily or monthly spend) and priority tiers.

Go deeper

Finished reading? Mark it done to track your progress.