Sizing and Sharding Vector Indexes
Size a vector index's memory, pick a sharding and replication scheme, and plan QPS capacity, freshness and re-embedding for a 100M-document system.
Step 1: memory
Per replica ≈ N × (vector bytes + graph bytes + metadata bytes)
Vector index sizing
How much RAM does your vector store need? Change the dimensions, precision and index type.
Worked example: 100M chunks, 1,536-dim float32, HNSW (M = 16), 100 bytes of metadata:
- Vectors: 100M × 6,144 B = 614 GB
- Graph: 100M × 2 × 16 × 4 B × 1.1 ≈ 14 GB
- IDs and metadata: 100M × 100 B = 10 GB
- ≈ 640 GB per replica
With int8 vectors (and full vectors on SSD for rescoring), vectors drop to 154 GB, so ≈ 180 GB per replica. That single choice cuts the fleet by more than 3×.
Step 2: shards
A shard holds a subset of the vectors. Keep each shard's index under about 60–70% of the node's RAM, leaving room for the OS, query buffers and rebuilds.
With 256 GB nodes and ≈ 640 GB per replica: 640 ÷ (256 × 0.7) ≈ 4 shards.
Sharding strategies
| Strategy | How queries work | Pros | Cons |
|---|---|---|---|
| By document ID (hash) | Scatter to all shards, gather each shard's top k, merge | Even data and load | Every query touches every shard. Tail latency = the slowest shard |
| By tenant | Route to the tenant's shard only | Isolation, easy per-tenant deletion, cheap queries | Uneven shard sizes. Big tenants need splitting |
| By semantic cluster (IVF-style routing) | Route to the few shards whose centroids are nearest | Fewer shards per query | Hot clusters. Recall loss at boundaries. Complex |
For multi-tenant SaaS, a common hybrid is to put many small tenants on shared shards (filtered by tenant ID) and give the largest tenants dedicated shards.
Step 3: replicas and QPS
Replicas are full copies of each shard. They serve two purposes: more query throughput, and surviving a node failure.
Estimate QPS per replica from benchmarks on your data and settings, for example ≈ 1,000 QPS per node for HNSW at 95% recall on 1,536-dim vectors (it varies widely with hardware, ef_search and filters). Then:
Replicas = max(2 for availability, peak QPS ÷ QPS per replica)
For a RAG product at 350 peak QPS: 2 replicas × 4 shards = 8 nodes, with room to spare. Vector search QPS is usually far cheaper to serve than the LLM calls that follow it.
Step 4: freshness and writes
- Upserts: document edited, so re-chunk, re-embed, and upsert its chunks. Delete chunks that no longer exist. Key chunks by (document ID, chunk number) so updates replace rather than duplicate.
- Deletes: tombstone, then compact. For GDPR and right-to-be-forgotten, make sure compaction actually runs and backups age out.
- Write path: changes flow through a queue (CDC from the source database, or webhooks). Embedding workers batch requests and write to the index. Typical freshness targets are seconds to minutes.
- Read-your-writes: a user who just uploaded a file expects to search it. Either wait for the index acknowledgement before confirming the upload, or query a small "recent" buffer alongside the main index.
Step 5: re-embedding and rebuilds
At some point you will change the embedding model, the chunking or the index parameters. Plan for it:
- Build the new index in parallel (a blue/green index).
- Backfill by re-embedding all documents in a batch pipeline. At 100M chunks and ≈ 5,000 chunks/s per GPU, that is ≈ 5.5 GPU-hours of embedding. The hours of index building and data movement usually take longer.
- Dual-write new changes to both indexes during the backfill.
- Compare recall and answer quality on your eval set, then switch reads with an alias. Keep the old index briefly for rollback.
Key takeaways
- Index memory = vectors (or their compressed codes) + graph links + IDs and metadata, per replica.
- Shard by document ID for even load (queries fan out to all shards), or by tenant for isolation (queries hit one shard).
- Replicas add QPS capacity and availability. Shards add memory capacity.
- Plan freshness (upserts and deletes), rebuilds and re-embedding migrations as first-class workloads.