GenAI System Design
4. RAG Architecture

Retrieval and Context Assembly

The query-time path in depth, covering query rewriting, hybrid retrieval, reranking, choosing top-k, deduplication, ordering, and fitting context into a token budget.

Lesson 3 of 7 10 min

Step 1: query understanding

The user's message is often a poor search query:

  • Follow-ups: "and for Android?" means nothing on its own.
  • Vague or multi-part: "compare the Pro and Team plans and tell me how to upgrade".
  • Implicit filters: "last quarter's incident reports" implies a date filter and a document type.

A small, fast LLM call can turn the conversation into:

  • A standalone query: "How do I enable SSO on the Android app?"
  • Multiple sub-queries for multi-part questions, retrieved in parallel.
  • Structured filters: {doc_type: "incident_report", date >= "2026-07-01"}.

This adds about 150–400 ms and one cheap call, and it is often the single biggest retrieval improvement for chat interfaces. Skip it for single-shot search boxes.

HyDE (hypothetical document embeddings) is a variant: have the LLM write a plausible answer, then embed that to search. It helps when questions and documents are phrased very differently.

Step 2: retrieve wide

Run hybrid retrieval, BM25 plus vector search fused with RRF, as covered in the hybrid search lesson:

  • Retrieve 50–100 candidates. Recall matters most at this stage.
  • Apply permission filters here, never after generation.
  • Apply metadata filters from query understanding, and optionally boost recency for time-sensitive content.

Step 3: rerank narrow

Score candidates with a cross-encoder reranker and keep the top 5–10. Precision matters now: every irrelevant chunk costs tokens and can mislead the model.

Use the reranker's score as a confidence signal. If even the best chunk scores low, the corpus probably doesn't contain the answer.

Step 4: assemble the context

Set a token budget first, for example 4,000 tokens for retrieved context, then fill it:

  1. Deduplicate near-identical chunks. The same paragraph often appears in several versions of a document.
  2. Group by document and merge adjacent chunks, so the model sees coherent passages rather than fragments.
  3. Expand to parents if you index small chunks but want to show larger sections (small-to-big retrieval).
  4. Order deliberately. Models attend best to the beginning and end of the context, so put the strongest evidence first.
  5. Label each chunk with an ID, title, date and source for citations.
100 candidates
hybrid retrieval
Rerank → 10
cross-encoder
Dedup + merge → 6 passages
Fit 4K-token budget
Prompt with IDs
for citations

Choosing top-k

  • Too few: the answer may be in chunk 7 and you only sent 5.
  • Too many: higher cost and TTFT, and relevant facts get diluted.

Measure answer quality at k = 3, 5, 10 and 20 on your eval set. Many systems land between 5 and 10 reranked chunks. A dynamic k keeps every chunk above a score threshold, up to the budget, which adapts to easy and hard questions.

Handling "no good answer"

When retrieval confidence is low:

  • Tell the model to say it couldn't find the answer, rather than guessing. Put this in the system prompt and test it in evals.
  • Ask a clarifying question if the query is ambiguous.
  • Fall back to a broader search or a different source (web search, a human handoff).

Log these cases. They are a free list of content gaps to send to documentation owners.

Citations and verification

  • Ask the model to cite chunk IDs inline, such as [3].
  • Verify that each cited ID was actually in the prompt, and drop or flag any that weren't.
  • For high-stakes answers, run a cheap faithfulness check: a small model confirms each claim is supported by its cited chunk.
  • Render citations as links, so users can check the source. This builds trust and exposes retrieval errors quickly.

Key takeaways

  • Rewrite conversational questions into standalone search queries before retrieving.
  • Retrieve wide (hybrid, top 50–100), rerank narrow (top 5–10), then assemble within an explicit token budget.
  • Deduplicate, group chunks by document, and order by relevance, with the most important at the start or end.
  • Handle the no-answer case explicitly. Low retrieval scores should lead to "I don't know" or a clarifying question.

Go deeper