Model Routing and Fallbacks
How to send each request to the right model and survive provider failures, covering static and dynamic routing, cascades, retries with backoff, failover, and circuit breakers.
Why route
Different requests need different models (see model tiers), and providers have outages, rate limits and latency spikes. Routing addresses both cost/quality and reliability.
Routing strategies
| Strategy | How it decides | Pros | Cons |
|---|---|---|---|
| Static (by alias or feature) | summarise → small model, agent-plan → frontier | Simple, predictable | Doesn't adapt to request difficulty |
| Classifier router | A small model or heuristic labels difficulty or intent, then picks the model | Adapts per request | Router errors, extra latency (≈ 50–300 ms) |
| Cascade | Try the cheap model first and escalate if confidence is low or validation fails | Pays for big models only when needed | Hard cases pay for two calls. Needs a reliable confidence signal |
| Latency / cost-aware | Pick the provider or region with the best current latency or price for the same model | Smooths provider variance | Needs live metrics |
Confidence signals for cascades: schema validation failures, the model saying it's unsure, low retrieval scores, a verifier model's score, or task-specific checks such as code that compiles and tests that pass.
Retries
Retry only errors that can succeed on a retry:
- 429 rate limited: wait, preferably for the provider's
Retry-After, then retry. - 5xx / overloaded / timeouts: retry with exponential backoff plus jitter, for example 0.5 s, 1 s, 2 s ± random.
- 4xx validation errors (bad request, context too long, content policy): don't retry unchanged. Fix the request or fail.
Cap retries (2–3) and total time. Retry storms during a provider incident make the incident worse.
For streaming, a failure mid-stream is awkward: tokens are already on the user's screen. Options are to restart the answer (and tell the UI), or to continue from the partial output on another model (hard to make seamless). Retries are simplest before the first token.
Failover
When a provider or region is degraded:
- Detect: rising error rate, TTFT p95 breaches, or timeouts. Use health checks and passive metrics from real traffic.
- Open a circuit breaker for that provider: stop sending traffic for a cool-down period, then probe with a small share of requests.
- Route to a fallback: the same model in another region or cloud (best), an equivalent model from another provider, or a smaller model in degraded mode.
| Fallback option | Quality risk | Notes |
|---|---|---|
| Same model, other region or cloud | None | Needs capacity or quota there |
| Different provider, similar tier | Medium | Prompts, tool calling and output formats may differ |
| Smaller model (degraded mode) | Higher | Tell users if capabilities are reduced |
| Cached or static response | Varies | For FAQ-like traffic |
Load balancing across keys and deployments
Large deployments spread load across multiple provider accounts, regions or provisioned-throughput deployments, weighted by remaining quota. Keep session stickiness where prompt caching matters, because the same conversation on the same deployment reuses its cached prefix.
Key takeaways
- Route by task (a static alias), by content (a classifier), or by confidence (a cascade that escalates to bigger models when needed).
- Retry only transient errors (429, 5xx, timeouts), with exponential backoff and jitter, and respect Retry-After headers.
- Fail over to another provider, region or model when a provider is degraded, and use circuit breakers to stop hammering a failing one.
- Fallback models may behave differently. Test prompts on them and watch quality after failover.