GenAI System Design
3. Embeddings and Vector Search

Sizing and Sharding Vector Indexes

Size a vector index's memory, pick a sharding and replication scheme, and plan QPS capacity, freshness and re-embedding for a 100M-document system.

Lesson 7 of 7 12 min

Step 1: memory

Per replica ≈ N × (vector bytes + graph bytes + metadata bytes)

Vector index sizing

How much RAM does your vector store need? Change the dimensions, precision and index type.

Vectors in RAM 614 GBHNSW graph links 14 GBIDs + metadata 10 GB
Raw vectors
614 GB
100M × 1536 × 4 B
RAM per replica
638 GB
RAM, all replicas
1.3 TB
Shards (256 GB nodes)
4 × 2 = 8 nodes
70% RAM fill target

Worked example: 100M chunks, 1,536-dim float32, HNSW (M = 16), 100 bytes of metadata:

  • Vectors: 100M × 6,144 B = 614 GB
  • Graph: 100M × 2 × 16 × 4 B × 1.1 ≈ 14 GB
  • IDs and metadata: 100M × 100 B = 10 GB
  • ≈ 640 GB per replica

With int8 vectors (and full vectors on SSD for rescoring), vectors drop to 154 GB, so ≈ 180 GB per replica. That single choice cuts the fleet by more than 3×.

Step 2: shards

A shard holds a subset of the vectors. Keep each shard's index under about 60–70% of the node's RAM, leaving room for the OS, query buffers and rebuilds.

With 256 GB nodes and ≈ 640 GB per replica: 640 ÷ (256 × 0.7) ≈ 4 shards.

Sharding strategies

StrategyHow queries workProsCons
By document ID (hash)Scatter to all shards, gather each shard's top k, mergeEven data and loadEvery query touches every shard. Tail latency = the slowest shard
By tenantRoute to the tenant's shard onlyIsolation, easy per-tenant deletion, cheap queriesUneven shard sizes. Big tenants need splitting
By semantic cluster (IVF-style routing)Route to the few shards whose centroids are nearestFewer shards per queryHot clusters. Recall loss at boundaries. Complex

For multi-tenant SaaS, a common hybrid is to put many small tenants on shared shards (filtered by tenant ID) and give the largest tenants dedicated shards.

Step 3: replicas and QPS

Replicas are full copies of each shard. They serve two purposes: more query throughput, and surviving a node failure.

Estimate QPS per replica from benchmarks on your data and settings, for example ≈ 1,000 QPS per node for HNSW at 95% recall on 1,536-dim vectors (it varies widely with hardware, ef_search and filters). Then:

Replicas = max(2 for availability, peak QPS ÷ QPS per replica)

For a RAG product at 350 peak QPS: 2 replicas × 4 shards = 8 nodes, with room to spare. Vector search QPS is usually far cheaper to serve than the LLM calls that follow it.

Step 4: freshness and writes

  • Upserts: document edited, so re-chunk, re-embed, and upsert its chunks. Delete chunks that no longer exist. Key chunks by (document ID, chunk number) so updates replace rather than duplicate.
  • Deletes: tombstone, then compact. For GDPR and right-to-be-forgotten, make sure compaction actually runs and backups age out.
  • Write path: changes flow through a queue (CDC from the source database, or webhooks). Embedding workers batch requests and write to the index. Typical freshness targets are seconds to minutes.
  • Read-your-writes: a user who just uploaded a file expects to search it. Either wait for the index acknowledgement before confirming the upload, or query a small "recent" buffer alongside the main index.

Step 5: re-embedding and rebuilds

At some point you will change the embedding model, the chunking or the index parameters. Plan for it:

  1. Build the new index in parallel (a blue/green index).
  2. Backfill by re-embedding all documents in a batch pipeline. At 100M chunks and ≈ 5,000 chunks/s per GPU, that is ≈ 5.5 GPU-hours of embedding. The hours of index building and data movement usually take longer.
  3. Dual-write new changes to both indexes during the backfill.
  4. Compare recall and answer quality on your eval set, then switch reads with an alias. Keep the old index briefly for rollback.

Key takeaways

  • Index memory = vectors (or their compressed codes) + graph links + IDs and metadata, per replica.
  • Shard by document ID for even load (queries fan out to all shards), or by tenant for isolation (queries hit one shard).
  • Replicas add QPS capacity and availability. Shards add memory capacity.
  • Plan freshness (upserts and deletes), rebuilds and re-embedding migrations as first-class workloads.

Go deeper

Finished reading? Mark it done to track your progress.