The Quality, Latency, Cost Triangle
The central trade-off of every LLM system, the levers that move you along each edge, and how to reason about them with numbers.
The triangle
In classic systems, faster usually means cheaper, because fewer servers are doing less work. In LLM systems the three corners pull against each other:
- Quality rises with bigger models, more retrieved context, and reasoning (thinking) tokens.
- Latency rises with every one of those.
- Cost is roughly linear in tokens processed, times the per-token price of the model you chose.
Two kinds of latency
Streaming changes what "latency" means, so always name which one you mean:
- Time to first token (TTFT): how long before text starts appearing. It covers network, queueing, retrieval, and the model's prefill of the prompt. For chat, target under about 1 second.
- Total time: TTFT plus output tokens ÷ decode speed. A 500-token answer at 60 tokens/s adds over 8 seconds. This matters for agents, tool calls, and APIs that return JSON, where nothing is useful until the response is complete.
Latency budget for one streamed answer
Drag the sliders. Notice which part dominates the total, and which part the user actually waits for.
- Network60 ms (1%)
- Queueing50 ms (1%)
- Retrieval150 ms (3%)
- Time to first token400 ms (7%)
- Decode5.0 s (88%)
Decode usually dominates the total. Streaming hides it, which is why chat products optimise time to first token and batch jobs optimise throughput.
Cost is token arithmetic
Cost per request = input tokens × input price + output tokens × output price.
Take a mid-tier model at $2 input and $10 output per million tokens, and a request with 2,000 tokens in and 400 out:
- Input: 2,000 × $2 / 1M = $0.004
- Output: 400 × $10 / 1M = $0.004
- Total ≈ $0.008 per request
At 10 million requests a day that is $80K a day, about $2.4M a month. In a classic system, cost per request is rarely a design topic. In an LLM system it often decides the design.
The levers
Read + as "better" and − as "worse" for that corner.
| Lever | Quality | Latency | Cost | When to use |
|---|---|---|---|---|
| Bigger model | + | − | − | Hard reasoning, high cost of errors |
| Routing / cascade (small model first) | ≈ | + | ++ | Mixed-difficulty traffic |
| Retrieval instead of huge prompts | + | + | + | Private or fresh knowledge |
| Prompt (prefix) caching | = | + (TTFT) | + | Long shared system prompts or documents |
| Semantic cache of answers | risky | ++ | ++ | Many repeated, low-risk questions |
| Batch API / offline jobs | = | − (minutes) | + | Anything not interactive |
| Quantized self-hosted model | slightly − | + | + at scale | High, steady volume |
| Reasoning / thinking tokens | + | −− | −− | Planning, maths, complex code |
| Shorter output (format constraints) | = | + | + | Always |
Worked example: routing
Suppose 70% of your traffic is easy (greetings, simple lookups) and a small model at $0.25 / $2 per million handles it well:
- Small model per request: 2,000 × 0.25/1M + 400 × 2/1M = $0.0013
- Blended: 0.7 × $0.0013 + 0.3 × $0.008 = $0.0033, a 59% saving
The catch is that you need a router, which can be a classifier, a cheap LLM call, or a heuristic. You also need an eval set that shows the small model really is good enough on the traffic you send it. Mention both.
Key takeaways
- Every design decision in an LLM system trades quality, latency and cost against each other.
- Routing easy requests to small models often cuts cost by more than half with little quality loss.
- Prompt caching, batch APIs and shorter prompts cut cost without touching quality.
- Latency has two numbers, time to first token and total time. Say which one your product cares about.