The Six-Step Framework
A repeatable structure for any GenAI system design question, with a time budget for a 45-minute interview and a worked example.
Why a GenAI-specific framework
The classic framework (requirements, estimation, high-level design, deep dive) still works, but it leaves out two things interviewers now probe for. One is the quality bar: how good the answers must be, and how you measure that. The other is the operating point: where you sit between quality, latency and cost. Put both up front and the rest of the design follows from them.
Step 1: Clarify the task and the quality bar (≈ 5 min)
Ask the usual functional and scale questions, then add the GenAI ones:
- What exactly does the model produce? An answer, a summary, a classification, code, or an action taken through a tool?
- What does "good" mean, and who judges it? Consider factual accuracy against a source, tone, format, or task success such as a ticket being resolved.
- What is the cost of a wrong answer? A bad movie recommendation is cheap. A wrong medical or financial statement is not. This decides how much you invest in grounding, citations and human review.
- What data does the model need? Public knowledge, private documents, per-user data? Consider freshness and permissions too.
- Is it interactive or batch? A chat UI streams and cares about time to first token. A nightly job cares about throughput and price.
Step 2: Choose the operating point (≈ 3 min)
State which corner of the quality, latency, cost triangle matters most. For example: "This is a customer-facing chat, so first token under 1 second at p95 and cost under 1 cent per conversation turn. We accept a mid-sized model and make up quality with retrieval."
This one sentence justifies later choices such as model size, routing, caching and whether to self-host.
Step 3: Estimate (≈ 5 min)
Convert users into requests, requests into tokens, and tokens into GPUs or dollars. The back-of-envelope lesson walks through it. Aim for the right order of magnitude. The goal is to know whether you need 5 GPUs or 500.
Step 4: Architecture (≈ 12 min)
Draw two things:
- The online request path: client, gateway, orchestration, retrieval, model, guardrails, streaming back.
- The offline pipelines: document ingestion and embedding, eval runs, analytics, and fine-tuning if it is justified.
Name the model choice (hosted API or self-hosted open model) and explain it from step 2.
Step 5: Deep dive (≈ 12 min)
Pick the one or two components most likely to break at your scale, or ask the interviewer to choose. Common deep dives:
- Retrieval quality: chunking, hybrid search, reranking, freshness.
- Serving: GPU memory, KV cache, batching, autoscaling.
- Agent loop: tool design, step limits, state, retries.
- Cost: caching, routing to smaller models, batch APIs.
Step 6: Evaluate, safeguard, operate (≈ 8 min)
Many candidates run out of time before this step, and interviewers notice. Cover:
- Offline evals: a golden set of questions with reference answers, scored automatically (LLM-as-judge plus exact checks) on every prompt or model change.
- Online signals: thumbs up or down, task success, escalation rate, and A/B tests.
- Guardrails: prompt injection defences, PII filtering, output moderation, and permission checks on retrieved data.
- Operations: tracing each request's prompt, retrieved chunks, tokens and latency; fallbacks when a provider is down; spend alerts.
Worked example: "Design an assistant that drafts replies to support tickets"
| Step | What you say |
|---|---|
| Clarify | Agents see a drafted reply and edit it before sending. The quality bar is the share of drafts sent with minor edits. Knowledge comes from the help centre plus past resolved tickets, and permissions are per customer. |
| Operating point | This is not user-facing in real time, so a 5-second draft is fine. Quality first, then cost. |
| Estimate | 50K tickets a day ≈ 0.6 QPS average. 4K tokens in and 300 out per draft. On a $2 / $10 per million token model that is ≈ $0.011 per draft, about $550 a day. A hosted API is the easy choice. |
| Architecture | Ticket event goes onto a queue, then a worker retrieves relevant articles and similar tickets (hybrid search), then an LLM drafts with citations, then the draft is stored and shown in the agent UI. Offline: nightly ingestion of new articles and resolved tickets. |
| Deep dive | Retrieval over past tickets, with PII scrubbing and per-customer filters. |
| Evaluate | A golden set of 500 historical tickets with the replies agents actually sent. Track edit distance between draft and sent reply as the online metric. |
Key takeaways
- Use six steps: clarify the task and quality bar, pick the operating point, estimate, design the architecture, deep-dive, then evaluate and safeguard.
- Define what a good answer is, and who judges it, before designing anything.
- Spend most of the time on architecture and one or two deep dives, but never skip evals and guardrails.
- Tie every component back to a number or a requirement you stated earlier.