How vector index memory is calculated
Memory for one loaded copy of an index has three parts:
- Vectors:
vectors × dimensions × bytes per dimension. With product quantization, the in-memory part is the PQ code instead, and full vectors stay on disk for rescoring. - Graph links (HNSW): about
vectors × 2 × M × 4 bytesfor layer 0, plus roughly 10% for upper layers. At M = 16 that is about 141 bytes per vector. - IDs and scalar fields: primary keys plus any fields loaded for filtering. This is easy to underestimate when you filter on text or JSON fields.
Nodes, replicas and headroom
Adding nodes spreads one copy of the index. Adding replicas adds whole copies, each needing its own group of nodes. This calculator sizes each replica group to hold a full copy within the usable memory, then adds one spare node so every replica still fits after a node fails. Try the interactive version in the sizing and sharding lesson.
Memory is the floor, not the plan
Memory tells you the smallest fleet that can hold the index. Query throughput, filter selectivity, ingestion rate and index-build capacity decide the rest, and they have to be benchmarked on your data. In a distributed database like Milvus, inserts also consume memory on the query nodes before data is indexed. The Milvus incident patterns show how that plays out.