GenAI System Design
0. How to Ace a GenAI System Design Interview

The Quality, Latency, Cost Triangle

The central trade-off of every LLM system, the levers that move you along each edge, and how to reason about them with numbers.

Lesson 3 of 6 9 min

The triangle

QualityLatencyCostBigger model, reasoning,more retrieved context+ quality, − latencyRoute easy queries to asmall model, escalate hard oneskeeps quality, cuts costCaching, batching, quantization, shorter prompts+ latency and cost, watch quality
Every lever moves you along one edge. Say which corner the product cares about most, then choose levers deliberately.

In classic systems, faster usually means cheaper, because fewer servers are doing less work. In LLM systems the three corners pull against each other:

  • Quality rises with bigger models, more retrieved context, and reasoning (thinking) tokens.
  • Latency rises with every one of those.
  • Cost is roughly linear in tokens processed, times the per-token price of the model you chose.

Two kinds of latency

Streaming changes what "latency" means, so always name which one you mean:

  • Time to first token (TTFT): how long before text starts appearing. It covers network, queueing, retrieval, and the model's prefill of the prompt. For chat, target under about 1 second.
  • Total time: TTFT plus output tokens ÷ decode speed. A 500-token answer at 60 tokens/s adds over 8 seconds. This matters for agents, tool calls, and APIs that return JSON, where nothing is useful until the response is complete.

Latency budget for one streamed answer

Drag the sliders. Notice which part dominates the total, and which part the user actually waits for.

First token appears after
660 ms
what users perceive as speed
Full answer done after
5.7 s
matters for agents and APIs
  • Network60 ms (1%)
  • Queueing50 ms (1%)
  • Retrieval150 ms (3%)
  • Time to first token400 ms (7%)
  • Decode5.0 s (88%)

Decode usually dominates the total. Streaming hides it, which is why chat products optimise time to first token and batch jobs optimise throughput.

Cost is token arithmetic

Cost per request = input tokens × input price + output tokens × output price.

Take a mid-tier model at $2 input and $10 output per million tokens, and a request with 2,000 tokens in and 400 out:

  • Input: 2,000 × $2 / 1M = $0.004
  • Output: 400 × $10 / 1M = $0.004
  • Total ≈ $0.008 per request

At 10 million requests a day that is $80K a day, about $2.4M a month. In a classic system, cost per request is rarely a design topic. In an LLM system it often decides the design.

The levers

Read + as "better" and − as "worse" for that corner.

LeverQualityLatencyCostWhen to use
Bigger model+−−Hard reasoning, high cost of errors
Routing / cascade (small model first)≈+++Mixed-difficulty traffic
Retrieval instead of huge prompts+++Private or fresh knowledge
Prompt (prefix) caching=+ (TTFT)+Long shared system prompts or documents
Semantic cache of answersrisky++++Many repeated, low-risk questions
Batch API / offline jobs=− (minutes)+Anything not interactive
Quantized self-hosted modelslightly −++ at scaleHigh, steady volume
Reasoning / thinking tokens+−−−−Planning, maths, complex code
Shorter output (format constraints)=++Always

Worked example: routing

Suppose 70% of your traffic is easy (greetings, simple lookups) and a small model at $0.25 / $2 per million handles it well:

  • Small model per request: 2,000 × 0.25/1M + 400 × 2/1M = $0.0013
  • Blended: 0.7 × $0.0013 + 0.3 × $0.008 = $0.0033, a 59% saving

The catch is that you need a router, which can be a classifier, a cheap LLM call, or a heuristic. You also need an eval set that shows the small model really is good enough on the traffic you send it. Mention both.

Key takeaways

  • Every design decision in an LLM system trades quality, latency and cost against each other.
  • Routing easy requests to small models often cuts cost by more than half with little quality loss.
  • Prompt caching, batch APIs and shorter prompts cut cost without touching quality.
  • Latency has two numbers, time to first token and total time. Say which one your product cares about.

Go deeper

Finished reading? Mark it done to track your progress.