Designing Embedding-Based Recommendations
A full walkthrough for a recommendation feed serving 100M users over 10M items, covering two-tower embeddings, ANN candidate generation, multi-stage ranking, cold start with content embeddings, freshness, feedback loops, and where LLMs help.
Lesson 10 of 10 14 min
1. Requirements
Functional
- A personalised home feed and "more like this" for a content platform (articles, videos or products)
- New items should be recommendable within minutes of publishing
- Filters: region, language, content-safety settings, already-seen items
Non-functional
- 100M users, 10M active items
- 20 feed requests per user per day, p99 under 150 ms for the recommendation service
- Continuous improvement through A/B testing
2. Operating point
Latency and throughput are hard constraints at 70K QPS peak. Relevance and long-term engagement are the objective. LLMs sit offline (content understanding, embeddings, labels), not in the per-request path.
3. Estimates
| Quantity | Value |
|---|---|
| Feed requests per day | 100M × 20 = 2B, ≈ 23K QPS average, ≈ 70K peak |
| Item embeddings (256-dim, fp16) | 10M × 512 B ≈ 5 GB, which fits in RAM on every node, so replicate rather than shard |
| HNSW graph (M = 16) | 10M × 128 B × 1.1 ≈ 1.4 GB |
| ANN QPS per node | Several thousand at small dims, so ≈ 20–30 nodes for 70K QPS with headroom |
| Ranking | 70K requests × 1,000 candidates = 70M candidate scores per second, the real compute cost, on GPU or CPU ranking servers |
| New items per day | ≈ 500K, embedded within minutes |
4. Architecture
Online
Feed request
user_id, context
User embedding
user tower on recent activity
Candidate generation
ANN top 1,000 + other sources
Filter
seen, region, safety
Ranker
deep model, top 100
Re-rank
diversity, freshness, rules
Offline and near-line
Interaction logs
views, clicks, dwell
Train two-tower + ranker
daily
Item embeddings
batch + streaming for new items
ANN index
rebuild + incremental inserts
Feature store
user + item features
5. Deep dives
Deep dive A: two-tower candidate generation
- A user tower encodes user features and recent interaction history into a vector. An item tower encodes item features and content. They are trained so that the dot product predicts engagement (contrastive loss with in-batch negatives).
- Item vectors are precomputed and indexed with ANN. The user vector is computed per request (a few ms), or near-line after each session.
- Multiple candidate sources are merged: two-tower ANN, "similar to recently engaged items" (item-to-item ANN), follow or subscription graph, trending by region. This improves coverage and robustness.
Deep dive B: ranking funnel
| Stage | Candidates | Model | Budget |
|---|---|---|---|
| Candidate generation | 10M → 1,000 | ANN dot product | ≈ 10–20 ms |
| Light ranker (optional) | 1,000 → 300 | Small model | ≈ 10 ms |
| Heavy ranker | 300 → 50 | Deep model with cross features | ≈ 30–50 ms |
| Re-rank | 50 → 20 shown | Diversity (MMR), freshness, business rules | ≈ 5 ms |
The ranker predicts several objectives (click, dwell, like, hide) combined into a score whose weights encode product goals.
Deep dive C: cold start and freshness
- New items have no interaction history. Their item-tower input includes content embeddings (a text encoder on title and description, an image encoder on thumbnails, or LLM-generated summaries and tags embedded), so they land near similar items immediately. Insert them into the ANN index incrementally within minutes.
- Exploration: reserve a small share of impressions for new or uncertain items, using bandit-style strategies, so they gather signal.
- New users: onboarding choices, popular-by-region items, and session-based signals that update the user vector within the first few interactions.
6. Where LLMs help
- Content understanding at ingestion: topics, entities, quality and safety signals, and summaries, which become ranking features and content embeddings.
- Label generation: LLM-judged relevance for (user interest, item) pairs to augment sparse feedback, calibrated against real engagement.
- Explanations: "Because you read X", generated offline or cheaply for the shown items.
- Conversational discovery ("find me something like X but shorter") as a separate surface that queries the same embedding index.
7. Evaluation
- Offline: recall@K for candidate generation, NDCG and AUC for the ranker, on time-split held-out interactions (train on the past, test on the future, to avoid leakage).
- Online A/B: engagement (CTR, dwell), long-term retention, hides and reports, creator-side distribution fairness.
- Feedback loops: models trained on their own recommendations amplify popularity bias. Use exploration traffic, debiasing (inverse propensity weighting), and diversity constraints.
8. Bottlenecks and follow-ups
| Likely question | Answer sketch |
|---|---|
| "Catalogue grows to 1B items?" | Shard the ANN index or use IVF-PQ, route by region or category, and add more candidate sources |
| "Real-time interests (just watched X)?" | Near-line user embedding updates from the event stream, plus item-to-item candidates from recent views |
| "Filter bubbles?" | Diversity re-ranking, exploration, measuring topic spread per user |
| "Ranking cost too high?" | A light ranker stage, feature caching, distilling the heavy ranker, GPU batching |
Key takeaways
- Recommendations are a funnel. ANN over item embeddings retrieves about 1,000 candidates, a ranking model scores them, and a re-ranker applies diversity and business rules.
- Two-tower models embed users and items in one space, so item vectors can be precomputed and indexed.
- Cold start for new items uses content embeddings from text, image or LLM encoders, until interaction data accumulates.
- Evaluate offline with recall@K and NDCG, and online with engagement and long-term satisfaction. Beware feedback loops and popularity bias.
Go deeper
Finished reading? Mark it done to track your progress.