GenAI System Design
6. Integrating LLMs into Existing Systems

Model Routing and Fallbacks

How to send each request to the right model and survive provider failures, covering static and dynamic routing, cascades, retries with backoff, failover, and circuit breakers.

Lesson 2 of 6 9 min

Why route

Different requests need different models (see model tiers), and providers have outages, rate limits and latency spikes. Routing addresses both cost/quality and reliability.

Routing strategies

StrategyHow it decidesProsCons
Static (by alias or feature)summarise → small model, agent-plan → frontierSimple, predictableDoesn't adapt to request difficulty
Classifier routerA small model or heuristic labels difficulty or intent, then picks the modelAdapts per requestRouter errors, extra latency (≈ 50–300 ms)
CascadeTry the cheap model first and escalate if confidence is low or validation failsPays for big models only when neededHard cases pay for two calls. Needs a reliable confidence signal
Latency / cost-awarePick the provider or region with the best current latency or price for the same modelSmooths provider varianceNeeds live metrics
Request
Router
intent + difficulty
Small model
≈ 70% of traffic
Validate
schema, confidence, rules
Escalate to frontier
only on failure
A cascade: most requests stop at the small model.

Confidence signals for cascades: schema validation failures, the model saying it's unsure, low retrieval scores, a verifier model's score, or task-specific checks such as code that compiles and tests that pass.

Retries

Retry only errors that can succeed on a retry:

  • 429 rate limited: wait, preferably for the provider's Retry-After, then retry.
  • 5xx / overloaded / timeouts: retry with exponential backoff plus jitter, for example 0.5 s, 1 s, 2 s ± random.
  • 4xx validation errors (bad request, context too long, content policy): don't retry unchanged. Fix the request or fail.

Cap retries (2–3) and total time. Retry storms during a provider incident make the incident worse.

For streaming, a failure mid-stream is awkward: tokens are already on the user's screen. Options are to restart the answer (and tell the UI), or to continue from the partial output on another model (hard to make seamless). Retries are simplest before the first token.

Failover

When a provider or region is degraded:

  1. Detect: rising error rate, TTFT p95 breaches, or timeouts. Use health checks and passive metrics from real traffic.
  2. Open a circuit breaker for that provider: stop sending traffic for a cool-down period, then probe with a small share of requests.
  3. Route to a fallback: the same model in another region or cloud (best), an equivalent model from another provider, or a smaller model in degraded mode.
Fallback optionQuality riskNotes
Same model, other region or cloudNoneNeeds capacity or quota there
Different provider, similar tierMediumPrompts, tool calling and output formats may differ
Smaller model (degraded mode)HigherTell users if capabilities are reduced
Cached or static responseVariesFor FAQ-like traffic

Load balancing across keys and deployments

Large deployments spread load across multiple provider accounts, regions or provisioned-throughput deployments, weighted by remaining quota. Keep session stickiness where prompt caching matters, because the same conversation on the same deployment reuses its cached prefix.

Key takeaways

  • Route by task (a static alias), by content (a classifier), or by confidence (a cascade that escalates to bigger models when needed).
  • Retry only transient errors (429, 5xx, timeouts), with exponential backoff and jitter, and respect Retry-After headers.
  • Fail over to another provider, region or model when a provider is degraded, and use circuit breakers to stop hammering a failing one.
  • Fallback models may behave differently. Test prompts on them and watch quality after failover.

Go deeper

Finished reading? Mark it done to track your progress.