Anatomy of a RAG System
The online and offline halves of retrieval-augmented generation, what each component does, where the latency and quality budget goes, and how to present RAG in an interview.
Why RAG exists
LLMs know what was in their training data, up to a cutoff date, and nothing about your company's documents, tickets or product catalogue. You have three ways to give a model your knowledge:
- Put it in the prompt, which is RAG, or long context for small corpora.
- Train it in with fine-tuning. This is poor for facts, as lesson 7 explains.
- Let it call tools that look things up, which is agentic retrieval.
RAG is the default because the knowledge stays fresh (update the index, not the model), citable (you know which chunks were used), and permission-aware (filter what each user can retrieve).
The two halves
Offline: ingestion
Online: query
What each component is responsible for
| Component | Job | Typical failure |
|---|---|---|
| Connectors and parsing | Turn source formats into clean text with structure | Tables flattened into nonsense, headers and footers repeated on every chunk |
| Chunking | Units that are self-contained and retrievable | Facts split from their context (next lesson) |
| Embedding and indexing | Searchable representation | Wrong model for the domain, stale vectors |
| Query rewriting | Turn a conversational turn into a good search query | "What about Android?" searched literally |
| Retrieval | High recall of relevant chunks | Missing exact-term matches without BM25 |
| Reranking | High precision in the final top-k | Irrelevant chunks distract the model |
| Generation | Answer faithfully from the context, with citations | Ignoring the context, or inventing citations |
The latency budget
For a typical RAG request:
| Step | Time |
|---|---|
| Query rewrite (small LLM) | 150–400 ms, optional |
| Embed query | 10–100 ms |
| Hybrid search | 10–40 ms |
| Rerank top 50 | 50–150 ms |
| LLM TTFT (4–8K-token prompt) | 300–800 ms |
| Generation | 2–6 s (streamed) |
Retrieval costs a few hundred milliseconds, and the LLM dominates. Retrieval quality, though, decides whether the answer is right.
The generation prompt
A solid RAG prompt:
- States the task and tells the model to answer only from the provided sources, and to say so when they don't contain the answer.
- Presents chunks with IDs and source metadata (title, date, URL) so the model can cite them.
- Puts static instructions first and the retrieved chunks and question last. This maximises prompt-cache hits.
- Asks for citations in a parseable format, so you can verify that each cited ID was actually provided and render links.
How to present RAG in an interview
- Draw both halves. Many candidates forget ingestion.
- State the scale: number of documents and chunks, update rate, QPS. Size the index with the vector sizing method.
- Name the retrieval design: hybrid, rerank, top-k and a context budget in tokens.
- Cover permissions, freshness, citations and evaluation. These separate a senior answer from a tutorial answer.
Key takeaways
- RAG retrieves relevant content at query time and puts it in the prompt, so the model answers from your data rather than only its training data.
- It has two halves: an offline ingestion pipeline (parse, chunk, embed, index) and an online query path (rewrite, retrieve, rerank, generate, cite).
- Most RAG quality problems are retrieval problems, such as bad chunks, missing documents or no hybrid search, not LLM problems.
- Design citations, permissions and freshness in from the start. They are hard to add later.