Retrieval and Context Assembly
The query-time path in depth, covering query rewriting, hybrid retrieval, reranking, choosing top-k, deduplication, ordering, and fitting context into a token budget.
Step 1: query understanding
The user's message is often a poor search query:
- Follow-ups: "and for Android?" means nothing on its own.
- Vague or multi-part: "compare the Pro and Team plans and tell me how to upgrade".
- Implicit filters: "last quarter's incident reports" implies a date filter and a document type.
A small, fast LLM call can turn the conversation into:
- A standalone query: "How do I enable SSO on the Android app?"
- Multiple sub-queries for multi-part questions, retrieved in parallel.
- Structured filters:
{doc_type: "incident_report", date >= "2026-07-01"}.
This adds about 150–400 ms and one cheap call, and it is often the single biggest retrieval improvement for chat interfaces. Skip it for single-shot search boxes.
HyDE (hypothetical document embeddings) is a variant: have the LLM write a plausible answer, then embed that to search. It helps when questions and documents are phrased very differently.
Step 2: retrieve wide
Run hybrid retrieval, BM25 plus vector search fused with RRF, as covered in the hybrid search lesson:
- Retrieve 50–100 candidates. Recall matters most at this stage.
- Apply permission filters here, never after generation.
- Apply metadata filters from query understanding, and optionally boost recency for time-sensitive content.
Step 3: rerank narrow
Score candidates with a cross-encoder reranker and keep the top 5–10. Precision matters now: every irrelevant chunk costs tokens and can mislead the model.
Use the reranker's score as a confidence signal. If even the best chunk scores low, the corpus probably doesn't contain the answer.
Step 4: assemble the context
Set a token budget first, for example 4,000 tokens for retrieved context, then fill it:
- Deduplicate near-identical chunks. The same paragraph often appears in several versions of a document.
- Group by document and merge adjacent chunks, so the model sees coherent passages rather than fragments.
- Expand to parents if you index small chunks but want to show larger sections (small-to-big retrieval).
- Order deliberately. Models attend best to the beginning and end of the context, so put the strongest evidence first.
- Label each chunk with an ID, title, date and source for citations.
Choosing top-k
- Too few: the answer may be in chunk 7 and you only sent 5.
- Too many: higher cost and TTFT, and relevant facts get diluted.
Measure answer quality at k = 3, 5, 10 and 20 on your eval set. Many systems land between 5 and 10 reranked chunks. A dynamic k keeps every chunk above a score threshold, up to the budget, which adapts to easy and hard questions.
Handling "no good answer"
When retrieval confidence is low:
- Tell the model to say it couldn't find the answer, rather than guessing. Put this in the system prompt and test it in evals.
- Ask a clarifying question if the query is ambiguous.
- Fall back to a broader search or a different source (web search, a human handoff).
Log these cases. They are a free list of content gaps to send to documentation owners.
Citations and verification
- Ask the model to cite chunk IDs inline, such as [3].
- Verify that each cited ID was actually in the prompt, and drop or flag any that weren't.
- For high-stakes answers, run a cheap faithfulness check: a small model confirms each claim is supported by its cited chunk.
- Render citations as links, so users can check the source. This builds trust and exposes retrieval errors quickly.
Key takeaways
- Rewrite conversational questions into standalone search queries before retrieving.
- Retrieve wide (hybrid, top 50–100), rerank narrow (top 5–10), then assemble within an explicit token budget.
- Deduplicate, group chunks by document, and order by relevance, with the most important at the start or end.
- Handle the no-answer case explicitly. Low retrieval scores should lead to "I don't know" or a clarifying question.