GenAI System Design
10. GenAI System Design Problems

Designing Semantic Product Search

A full walkthrough for e-commerce search over 50M products at 5K QPS with a 200 ms budget, covering hybrid retrieval, query understanding, where LLMs fit (and don't), learning-to-rank, freshness, and evaluation.

Lesson 5 of 10 14 min

1. Requirements

Functional

  • Search a catalogue of 50M products by keyword and by natural language ("waterproof hiking boots for wide feet under $150")
  • Filters and facets (brand, price, size, in stock), and sort options
  • Personalisation, such as preferred brands and sizes

Non-functional

  • 5K QPS at peak, p99 search latency under 200 ms
  • Inventory and price changes reflected within about a minute
  • Relevance improves conversion. This is a revenue-critical system

2. Operating point

Latency and cost are fixed constraints. Quality is the objective within them. At 5K QPS, even a $0.0005 LLM call per query is $2.5 a second, or about $6.5M a month, and adds hundreds of milliseconds. LLMs belong offline and on the long tail.

3. Estimates

QuantityValue
Products50M
Vectors (768-dim)float32 154 GB, int8 38 GB
HNSW graph (M = 32)50M × 64 × 4 B × 1.1 ≈ 14 GB
Per-replica index≈ 55–60 GB with int8, which fits one large node, so no sharding needed
QPS per node (filtered HNSW)≈ 1K (benchmark to confirm)
Replicas5K ÷ 1K = 5, plus headroom and availability, so ≈ 8
Query embedding5K QPS of short queries, batched on 1–2 small-GPU encoders, with head queries cached
Catalogue enrichment50M products × ≈ 1K tokens with a small LLM ≈ 50B tokens one-off (≈ $12K at $0.25 per million), then only changes

Query frequency is heavy-tailed: the top ~10% of distinct queries make up most of the traffic, so result caching for head queries removes a large share of the work.

4. Architecture

Online

Query
Normalise + cache
head queries hit here
Query understanding
category, attributes, filters
Parallel retrieval
BM25 + ANN, filtered
Fusion → 500
Ranker
LTR: relevance, popularity, personal

Offline

Catalogue changes
LLM enrichment
attributes, synonyms, clean titles
Embed products
Index update
vectors + fields
Click logs → training
ranker, encoders

5. Deep dives

Deep dive A: where LLMs fit

UseOnline or offlineWhy
Catalogue enrichment: extract attributes (material, fit, use case), normalise titles, generate synonymsOfflineImproves both BM25 and vector matching for every query, at one-off cost
Query understanding for head and torso queries: parse "boots under $150 wide fit" into filtersOffline, precomputed and cachedCovers most traffic at zero online cost
Long-tail query parsingOnline, small fast model, ≈ 30–60 msOnly for cache misses on complex queries
Generating embeddings for relevance judgments and training dataOfflineLabels at scale for the ranker
Conversational shopping assistantSeparate product surfaceDifferent latency expectations

Deep dive B: latency budget (p99 200 ms)

StepBudget
Network and gateway20 ms
Cache lookup2 ms
Query understanding (cached, or small model for the tail)5–60 ms
Query embedding (cached for head queries)5–15 ms
BM25 and ANN in parallel20–40 ms
Fusion and feature fetch15 ms
Ranker on 500 candidates (GBDT or small DNN)20–40 ms
Total≈ 100–190 ms

Deep dive C: filters and freshness

  • Filtering inside retrieval: "in stock in the user's region" is highly selective for some queries, so use in-index filtering with a planner that falls back to brute force over small filtered sets.
  • Inventory and price updates stream from the inventory service (CDC) into index metadata within seconds. Vectors only change when product content changes.
  • New products: embed on creation, and give them a freshness boost with exploration traffic so they can gather clicks.

6. Evaluation

  • Offline: a judged query set (head, torso and tail samples) with graded relevance labels, from human raters and LLM-assisted labelling calibrated to raters. Metrics are NDCG@10 and recall@100 for the candidate stage.
  • Online A/B: CTR, add-to-cart, conversion, revenue per search, zero-result rate, reformulation rate.
  • Guard metrics: latency p99, and diversity (avoid showing only one brand).

7. Bottlenecks and follow-ups

Likely questionAnswer sketch
"Why not ask an LLM to rank results?"Latency and cost at 5K QPS. Use it offline to create training labels for a fast ranker instead
"Catalogue grows to 1B items?"Shard the index, use IVF-PQ or disk-based indexes, and route by category
"Multilingual?"A multilingual embedding model, per-language BM25 analysers, and language-specific evals
"Zero results for long queries?"Relax filters progressively, rely on vector recall, and use LLM parsing for tail queries

Key takeaways

  • At 5K QPS with a 200 ms p99 budget, an LLM can't sit in every request. Use embeddings, hybrid retrieval and a fast ranker online.
  • Use LLMs offline to enrich the catalogue and to precompute query understanding for head and torso queries. Use small models online for the long tail.
  • Filters (in stock, region, price) must be applied during retrieval, and inventory changes must propagate in seconds to minutes.
  • Measure with NDCG on judged queries offline, and CTR, conversion and revenue per search in A/B tests.

Go deeper

Finished reading? Mark it done to track your progress.