Common Mistakes and How to Avoid Them
The ten mistakes that sink GenAI system design interviews, what each one sounds like, and what to say instead.
1. Reaching for components before the problem
Sounds like: "We'll use a vector database, LangChain and an agent" in the first two minutes.
Instead: clarify what the model must produce and how good it must be. Many problems need no retrieval at all, such as classification or rewriting user text. Many "agent" problems are a fixed pipeline of two LLM calls.
2. No definition of quality
Sounds like: a complete architecture with no mention of how you would know it works.
Instead: name a metric, such as answer accuracy on a golden set, task success rate, or edit distance to what humans send. Say how you measure it offline before every change and online after launch.
3. Ignoring token economics
Sounds like: "We'll put the whole knowledge base in the context window, since the model supports 1M tokens."
Instead: do the arithmetic. 500K tokens per request at $2 per million is $1 per question, and prefill alone takes many seconds. Retrieve the few thousand tokens that matter. Long context is for occasional deep analysis, not every request.
4. Treating the LLM like a fast RPC
Sounds like: a synchronous call with a 30-second timeout inside a request handler, and no retries.
Instead: stream tokens to the client, usually over server-sent events. Set separate timeouts for first token and for the whole response. Retry on 429 and 5xx with backoff and jitter. Move long jobs to a queue, and keep a fallback model or provider for outages.
5. Sizing GPUs from weights alone
Sounds like: "A 70B model is 140 GB, so two H100s."
Instead: two H100s hold the weights with almost no room left. Add the KV cache for your concurrent users (≈ 320 KB per token for 70B), which usually means FP8 weights, more GPUs, or capped context. See GPU memory and the KV cache.
6. RAG that leaks data
Sounds like: one shared index over all company documents, queried for every user.
Instead: attach access-control metadata to every chunk, and filter at query time by the caller's identity. For strong isolation, use separate indexes per tenant. Retrieval that ignores permissions is a data breach waiting to happen.
7. No defence against prompt injection
Sounds like: the model reads web pages or emails and can also call tools that send email or change data.
Instead: treat all retrieved or user-supplied text as untrusted. Give tools least privilege, and require confirmation for side effects. Separate instructions from data, and filter outputs. The most dangerous mix is untrusted input, private data and the ability to act, all in one agent.
8. Fine-tuning as the first answer
Sounds like: "We'll fine-tune the model on our docs so it knows them."
Instead: use retrieval for knowledge, since it is fresh, citable and permission-aware. Fine-tune for behaviour: format, tone, a narrow task, or distilling a big model into a cheaper one. Try prompting and RAG first.
9. Forgetting the offline half
Sounds like: a detailed online path with nothing about how documents get into the index or stay current.
Instead: draw the ingestion pipeline: change detection, parsing, chunking, embedding, upserts and deletes, and re-embedding when you change embedding models. Stale or missing documents are the most common cause of bad RAG answers.
10. Unbounded agents
Sounds like: "The agent loops until the task is done."
Instead: cap steps, tokens and wall-clock time per task. Checkpoint state so a failure resumes rather than restarts. Log every tool call. Context grows with each step, so cost grows faster than linearly with step count.
Key takeaways
- Start from the task and its quality bar, not from a favourite component such as a vector database or an agent framework.
- Treat LLM calls as slow, expensive, rate-limited and fallible dependencies, and design timeouts, streaming, retries and fallbacks.
- Size GPU memory for the KV cache as well as the weights.
- Retrieval must respect permissions, and every design needs evals and prompt-injection defences.