GenAI System Design
4. RAG Architecture

Anatomy of a RAG System

The online and offline halves of retrieval-augmented generation, what each component does, where the latency and quality budget goes, and how to present RAG in an interview.

Lesson 1 of 7 9 min

Why RAG exists

LLMs know what was in their training data, up to a cutoff date, and nothing about your company's documents, tickets or product catalogue. You have three ways to give a model your knowledge:

  1. Put it in the prompt, which is RAG, or long context for small corpora.
  2. Train it in with fine-tuning. This is poor for facts, as lesson 7 explains.
  3. Let it call tools that look things up, which is agentic retrieval.

RAG is the default because the knowledge stays fresh (update the index, not the model), citable (you know which chunks were used), and permission-aware (filter what each user can retrieve).

The two halves

Offline: ingestion

Sources
docs, wikis, tickets, DBs
Parse
PDF, HTML, tables → text
Chunk
structure-aware, with metadata
Embed
batched
Index
vector + keyword, with ACLs

Online: query

User question
Rewrite
standalone query, filters
Retrieve
hybrid: BM25 + vectors
Rerank
cross-encoder, top 5–10
Generate
LLM with context + citations
Post-process
verify citations, guardrails

What each component is responsible for

ComponentJobTypical failure
Connectors and parsingTurn source formats into clean text with structureTables flattened into nonsense, headers and footers repeated on every chunk
ChunkingUnits that are self-contained and retrievableFacts split from their context (next lesson)
Embedding and indexingSearchable representationWrong model for the domain, stale vectors
Query rewritingTurn a conversational turn into a good search query"What about Android?" searched literally
RetrievalHigh recall of relevant chunksMissing exact-term matches without BM25
RerankingHigh precision in the final top-kIrrelevant chunks distract the model
GenerationAnswer faithfully from the context, with citationsIgnoring the context, or inventing citations

The latency budget

For a typical RAG request:

StepTime
Query rewrite (small LLM)150–400 ms, optional
Embed query10–100 ms
Hybrid search10–40 ms
Rerank top 5050–150 ms
LLM TTFT (4–8K-token prompt)300–800 ms
Generation2–6 s (streamed)

Retrieval costs a few hundred milliseconds, and the LLM dominates. Retrieval quality, though, decides whether the answer is right.

The generation prompt

A solid RAG prompt:

  • States the task and tells the model to answer only from the provided sources, and to say so when they don't contain the answer.
  • Presents chunks with IDs and source metadata (title, date, URL) so the model can cite them.
  • Puts static instructions first and the retrieved chunks and question last. This maximises prompt-cache hits.
  • Asks for citations in a parseable format, so you can verify that each cited ID was actually provided and render links.

How to present RAG in an interview

  1. Draw both halves. Many candidates forget ingestion.
  2. State the scale: number of documents and chunks, update rate, QPS. Size the index with the vector sizing method.
  3. Name the retrieval design: hybrid, rerank, top-k and a context budget in tokens.
  4. Cover permissions, freshness, citations and evaluation. These separate a senior answer from a tutorial answer.

Key takeaways

  • RAG retrieves relevant content at query time and puts it in the prompt, so the model answers from your data rather than only its training data.
  • It has two halves: an offline ingestion pipeline (parse, chunk, embed, index) and an online query path (rewrite, retrieve, rerank, generate, cite).
  • Most RAG quality problems are retrieval problems, such as bad chunks, missing documents or no hybrid search, not LLM problems.
  • Design citations, permissions and freshness in from the start. They are hard to add later.

Go deeper

Finished reading? Mark it done to track your progress.