GenAI System Design
3. Embeddings and Vector Search

Hybrid Search and Reranking

Why vector search alone misses exact matches, how to fuse keyword and vector results with reciprocal rank fusion, and where a cross-encoder reranker fits in the retrieval pipeline.

Lesson 6 of 7 10 min

Where vector search falls short

Embeddings are trained to capture meaning, which is exactly why they struggle with things that aren't about meaning:

  • Identifiers and codes: "error E-4012", "SKU 88-2231-B", "RFC 9110"
  • Names and rare terms the embedding model barely saw in training
  • Exact phrases and negations: "not compatible with iOS"
  • Very short queries, where there is little context to embed

Keyword search, the BM25 ranking over an inverted index behind Elasticsearch, OpenSearch and Postgres full-text search, handles these well. It is weak on synonyms and paraphrases, which is exactly where embeddings are strong. The two fail in different places, so combining them is powerful.

The retrieval funnel

Query
rewrite / expand (optional)
BM25 top 100
≈ 5–30 ms
Vector top 100
≈ 5–20 ms incl. embedding
Fuse (RRF)
≈ 150 unique candidates
Rerank top 50
cross-encoder, ≈ 50–150 ms
Top 5–10 to the LLM
Each stage is more expensive per document and sees fewer documents.

Fusing results: reciprocal rank fusion

BM25 scores and cosine similarities live on different scales, so adding them directly doesn't work without careful normalisation. Reciprocal rank fusion uses only ranks:

RRF score(d) = Σ over result lists of 1 ÷ (k + rank of d in that list), with k ≈ 60

A document ranked 1st by BM25 and 5th by vectors scores 1/61 + 1/65. One that appears in only one list scores less. It needs no training and no score calibration, and it is surprisingly hard to beat, which is why most engines ship it as the default hybrid method. Weighted score fusion, with scores normalised to [0, 1], is the main alternative when you can tune weights on labelled data.

Reranking with a cross-encoder

The embedding model is a bi-encoder. It encodes the query and each document separately, so document vectors can be precomputed and indexed. That speed comes at a cost: the model never sees the query and document together.

A cross-encoder reranker takes the (query, document) pair as one input and outputs a relevance score. It can notice that the document mentions "iOS" but says "not compatible". The cost: every candidate needs its own model forward pass at query time.

LLMs can rerank too, by asking a model to score or order candidates. This is higher quality for complex relevance but much slower and costlier, so it suits low-QPS, high-value queries.

Other retrieval improvements

  • Query rewriting: use a small LLM to turn a conversational follow-up ("what about for Android?") into a standalone query, or to generate several query variants.
  • Metadata filters and boosts: recency, source authority, document type.
  • Contextual chunk enrichment: prepend a short summary of the document or section to each chunk before embedding, so chunks keep context such as "this section is about the 2025 pricing plan".
  • Learned sparse retrieval (SPLADE-style): sparse term vectors learned by a model, combining keyword precision with some semantic expansion.

Key takeaways

  • Embeddings capture meaning but blur exact terms such as product codes, names, error messages and rare words. BM25 catches those.
  • Hybrid search runs both and merges them. Reciprocal rank fusion (RRF) is the simple, robust default.
  • A cross-encoder reranker scores query and document together. It is far more accurate, but only affordable on the top 50–200 candidates.
  • The standard pipeline is a two-stage funnel: cheap, high-recall retrieval, then an expensive, high-precision reranker.

Go deeper

Finished reading? Mark it done to track your progress.