GenAI System Design
10. GenAI System Design Problems

Designing Embedding-Based Recommendations

A full walkthrough for a recommendation feed serving 100M users over 10M items, covering two-tower embeddings, ANN candidate generation, multi-stage ranking, cold start with content embeddings, freshness, feedback loops, and where LLMs help.

Lesson 10 of 10 14 min

1. Requirements

Functional

  • A personalised home feed and "more like this" for a content platform (articles, videos or products)
  • New items should be recommendable within minutes of publishing
  • Filters: region, language, content-safety settings, already-seen items

Non-functional

  • 100M users, 10M active items
  • 20 feed requests per user per day, p99 under 150 ms for the recommendation service
  • Continuous improvement through A/B testing

2. Operating point

Latency and throughput are hard constraints at 70K QPS peak. Relevance and long-term engagement are the objective. LLMs sit offline (content understanding, embeddings, labels), not in the per-request path.

3. Estimates

QuantityValue
Feed requests per day100M × 20 = 2B, ≈ 23K QPS average, ≈ 70K peak
Item embeddings (256-dim, fp16)10M × 512 B ≈ 5 GB, which fits in RAM on every node, so replicate rather than shard
HNSW graph (M = 16)10M × 128 B × 1.1 ≈ 1.4 GB
ANN QPS per nodeSeveral thousand at small dims, so ≈ 20–30 nodes for 70K QPS with headroom
Ranking70K requests × 1,000 candidates = 70M candidate scores per second, the real compute cost, on GPU or CPU ranking servers
New items per day≈ 500K, embedded within minutes

4. Architecture

Online

Feed request
user_id, context
User embedding
user tower on recent activity
Candidate generation
ANN top 1,000 + other sources
Filter
seen, region, safety
Ranker
deep model, top 100
Re-rank
diversity, freshness, rules

Offline and near-line

Interaction logs
views, clicks, dwell
Train two-tower + ranker
daily
Item embeddings
batch + streaming for new items
ANN index
rebuild + incremental inserts
Feature store
user + item features

5. Deep dives

Deep dive A: two-tower candidate generation

  • A user tower encodes user features and recent interaction history into a vector. An item tower encodes item features and content. They are trained so that the dot product predicts engagement (contrastive loss with in-batch negatives).
  • Item vectors are precomputed and indexed with ANN. The user vector is computed per request (a few ms), or near-line after each session.
  • Multiple candidate sources are merged: two-tower ANN, "similar to recently engaged items" (item-to-item ANN), follow or subscription graph, trending by region. This improves coverage and robustness.

Deep dive B: ranking funnel

StageCandidatesModelBudget
Candidate generation10M → 1,000ANN dot product≈ 10–20 ms
Light ranker (optional)1,000 → 300Small model≈ 10 ms
Heavy ranker300 → 50Deep model with cross features≈ 30–50 ms
Re-rank50 → 20 shownDiversity (MMR), freshness, business rules≈ 5 ms

The ranker predicts several objectives (click, dwell, like, hide) combined into a score whose weights encode product goals.

Deep dive C: cold start and freshness

  • New items have no interaction history. Their item-tower input includes content embeddings (a text encoder on title and description, an image encoder on thumbnails, or LLM-generated summaries and tags embedded), so they land near similar items immediately. Insert them into the ANN index incrementally within minutes.
  • Exploration: reserve a small share of impressions for new or uncertain items, using bandit-style strategies, so they gather signal.
  • New users: onboarding choices, popular-by-region items, and session-based signals that update the user vector within the first few interactions.

6. Where LLMs help

  • Content understanding at ingestion: topics, entities, quality and safety signals, and summaries, which become ranking features and content embeddings.
  • Label generation: LLM-judged relevance for (user interest, item) pairs to augment sparse feedback, calibrated against real engagement.
  • Explanations: "Because you read X", generated offline or cheaply for the shown items.
  • Conversational discovery ("find me something like X but shorter") as a separate surface that queries the same embedding index.

7. Evaluation

  • Offline: recall@K for candidate generation, NDCG and AUC for the ranker, on time-split held-out interactions (train on the past, test on the future, to avoid leakage).
  • Online A/B: engagement (CTR, dwell), long-term retention, hides and reports, creator-side distribution fairness.
  • Feedback loops: models trained on their own recommendations amplify popularity bias. Use exploration traffic, debiasing (inverse propensity weighting), and diversity constraints.

8. Bottlenecks and follow-ups

Likely questionAnswer sketch
"Catalogue grows to 1B items?"Shard the ANN index or use IVF-PQ, route by region or category, and add more candidate sources
"Real-time interests (just watched X)?"Near-line user embedding updates from the event stream, plus item-to-item candidates from recent views
"Filter bubbles?"Diversity re-ranking, exploration, measuring topic spread per user
"Ranking cost too high?"A light ranker stage, feature caching, distilling the heavy ranker, GPU batching

Key takeaways

  • Recommendations are a funnel. ANN over item embeddings retrieves about 1,000 candidates, a ranking model scores them, and a re-ranker applies diversity and business rules.
  • Two-tower models embed users and items in one space, so item vectors can be precomputed and indexed.
  • Cold start for new items uses content embeddings from text, image or LLM encoders, until interaction data accumulates.
  • Evaluate offline with recall@K and NDCG, and online with engagement and long-term satisfaction. Beware feedback loops and popularity bias.

Go deeper

Finished reading? Mark it done to track your progress.