Designing Semantic Product Search
A full walkthrough for e-commerce search over 50M products at 5K QPS with a 200 ms budget, covering hybrid retrieval, query understanding, where LLMs fit (and don't), learning-to-rank, freshness, and evaluation.
Lesson 5 of 10 14 min
1. Requirements
Functional
- Search a catalogue of 50M products by keyword and by natural language ("waterproof hiking boots for wide feet under $150")
- Filters and facets (brand, price, size, in stock), and sort options
- Personalisation, such as preferred brands and sizes
Non-functional
- 5K QPS at peak, p99 search latency under 200 ms
- Inventory and price changes reflected within about a minute
- Relevance improves conversion. This is a revenue-critical system
2. Operating point
Latency and cost are fixed constraints. Quality is the objective within them. At 5K QPS, even a $0.0005 LLM call per query is $2.5 a second, or about $6.5M a month, and adds hundreds of milliseconds. LLMs belong offline and on the long tail.
3. Estimates
| Quantity | Value |
|---|---|
| Products | 50M |
| Vectors (768-dim) | float32 154 GB, int8 38 GB |
| HNSW graph (M = 32) | 50M × 64 × 4 B × 1.1 ≈ 14 GB |
| Per-replica index | ≈ 55–60 GB with int8, which fits one large node, so no sharding needed |
| QPS per node (filtered HNSW) | ≈ 1K (benchmark to confirm) |
| Replicas | 5K ÷ 1K = 5, plus headroom and availability, so ≈ 8 |
| Query embedding | 5K QPS of short queries, batched on 1–2 small-GPU encoders, with head queries cached |
| Catalogue enrichment | 50M products × ≈ 1K tokens with a small LLM ≈ 50B tokens one-off (≈ $12K at $0.25 per million), then only changes |
Query frequency is heavy-tailed: the top ~10% of distinct queries make up most of the traffic, so result caching for head queries removes a large share of the work.
4. Architecture
Online
Query
Normalise + cache
head queries hit here
Query understanding
category, attributes, filters
Parallel retrieval
BM25 + ANN, filtered
Fusion → 500
Ranker
LTR: relevance, popularity, personal
Offline
Catalogue changes
LLM enrichment
attributes, synonyms, clean titles
Embed products
Index update
vectors + fields
Click logs → training
ranker, encoders
5. Deep dives
Deep dive A: where LLMs fit
| Use | Online or offline | Why |
|---|---|---|
| Catalogue enrichment: extract attributes (material, fit, use case), normalise titles, generate synonyms | Offline | Improves both BM25 and vector matching for every query, at one-off cost |
| Query understanding for head and torso queries: parse "boots under $150 wide fit" into filters | Offline, precomputed and cached | Covers most traffic at zero online cost |
| Long-tail query parsing | Online, small fast model, ≈ 30–60 ms | Only for cache misses on complex queries |
| Generating embeddings for relevance judgments and training data | Offline | Labels at scale for the ranker |
| Conversational shopping assistant | Separate product surface | Different latency expectations |
Deep dive B: latency budget (p99 200 ms)
| Step | Budget |
|---|---|
| Network and gateway | 20 ms |
| Cache lookup | 2 ms |
| Query understanding (cached, or small model for the tail) | 5–60 ms |
| Query embedding (cached for head queries) | 5–15 ms |
| BM25 and ANN in parallel | 20–40 ms |
| Fusion and feature fetch | 15 ms |
| Ranker on 500 candidates (GBDT or small DNN) | 20–40 ms |
| Total | ≈ 100–190 ms |
Deep dive C: filters and freshness
- Filtering inside retrieval: "in stock in the user's region" is highly selective for some queries, so use in-index filtering with a planner that falls back to brute force over small filtered sets.
- Inventory and price updates stream from the inventory service (CDC) into index metadata within seconds. Vectors only change when product content changes.
- New products: embed on creation, and give them a freshness boost with exploration traffic so they can gather clicks.
6. Evaluation
- Offline: a judged query set (head, torso and tail samples) with graded relevance labels, from human raters and LLM-assisted labelling calibrated to raters. Metrics are NDCG@10 and recall@100 for the candidate stage.
- Online A/B: CTR, add-to-cart, conversion, revenue per search, zero-result rate, reformulation rate.
- Guard metrics: latency p99, and diversity (avoid showing only one brand).
7. Bottlenecks and follow-ups
| Likely question | Answer sketch |
|---|---|
| "Why not ask an LLM to rank results?" | Latency and cost at 5K QPS. Use it offline to create training labels for a fast ranker instead |
| "Catalogue grows to 1B items?" | Shard the index, use IVF-PQ or disk-based indexes, and route by category |
| "Multilingual?" | A multilingual embedding model, per-language BM25 analysers, and language-specific evals |
| "Zero results for long queries?" | Relax filters progressively, rely on vector recall, and use LLM parsing for tail queries |
Key takeaways
- At 5K QPS with a 200 ms p99 budget, an LLM can't sit in every request. Use embeddings, hybrid retrieval and a fast ranker online.
- Use LLMs offline to enrich the catalogue and to precompute query understanding for head and torso queries. Use small models online for the long tail.
- Filters (in stock, region, price) must be applied during retrieval, and inventory changes must propagate in seconds to minutes.
- Measure with NDCG on judged queries offline, and CTR, conversion and revenue per search in A/B tests.
Go deeper
Finished reading? Mark it done to track your progress.