Vector Database Sizing Calculator

Work out how much memory a vector index needs, how many nodes it takes with replicas, and whether it survives losing a node. Works for any vector database, with a Milvus preset.

Qdrant, Weaviate, pgvector, OpenSearch and others.

50M vectors (chunks, not documents)

Links per node; layer 0 keeps 2 × M

bytes

IDs plus loaded scalar fields

Full loaded copies, for QPS and failover

%

For growing data, queries, rebuilds and the OS

Raw vectors
205 GB
50M × 1024 × 4 B
Index memory per replica
217 GB
Vectors + graph + metadata
Search nodes
11
5 per replica × 2 + 1 spare
Memory, all replicas
434 GB
68% of each node filled
Vectors (or PQ codes) 205 GBHNSW graph links 7.0 GBIDs + scalar fields 5.0 GB
Node size options for this index
Memory per search nodePer replicaTotal incl. spareTotal memory
32 GB1021672 GB
64 GB511704 GB
128 GB37896 GB
256 GB251,280 GB
512 GB131,536 GB

Each replica needs its own group of nodes holding a complete copy, so nodes = ⌈index per replica ÷ usable memory⌉ × replicas, plus one spare so every replica still fits after losing a node. Here each node offers 45 GB usable of 64 GB. Benchmark QPS separately: memory tells you the minimum fleet, not the throughput.

How vector index memory is calculated

Memory for one loaded copy of an index has three parts:

  • Vectors: vectors × dimensions × bytes per dimension. With product quantization, the in-memory part is the PQ code instead, and full vectors stay on disk for rescoring.
  • Graph links (HNSW): about vectors × 2 × M × 4 bytes for layer 0, plus roughly 10% for upper layers. At M = 16 that is about 141 bytes per vector.
  • IDs and scalar fields: primary keys plus any fields loaded for filtering. This is easy to underestimate when you filter on text or JSON fields.

Nodes, replicas and headroom

Adding nodes spreads one copy of the index. Adding replicas adds whole copies, each needing its own group of nodes. This calculator sizes each replica group to hold a full copy within the usable memory, then adds one spare node so every replica still fits after a node fails. Try the interactive version in the sizing and sharding lesson.

Memory is the floor, not the plan

Memory tells you the smallest fleet that can hold the index. Query throughput, filter selectivity, ingestion rate and index-build capacity decide the rest, and they have to be benchmarked on your data. In a distributed database like Milvus, inserts also consume memory on the query nodes before data is indexed. The Milvus incident patterns show how that plays out.

Frequently asked questions

How much RAM do 10 million 1,536-dimension embeddings need?

The raw float32 vectors are 10M × 1,536 × 4 bytes ≈ 61 GB. An HNSW index with M = 16 adds about 1.4 GB of graph links, plus whatever IDs and scalar fields you load. Plan on roughly 64 GB per replica, or about 18 GB if you store int8 vectors.

Does adding replicas reduce memory per node?

No. A replica is a complete loaded copy of the index, so each replica multiplies total memory. Adding nodes to a replica is what spreads its memory. Replicas buy query throughput and failover.

Why hold back 25–30% of node memory?

Vector databases need room beyond the loaded index: growing segments that haven't been indexed yet, query buffers, temporary memory for index builds and compaction, and the operating system. Running nodes near 100% is how an insert burst turns into out-of-memory crashes.

How does product quantization (PQ) change the size?

PQ replaces each vector with a short code, often 1 byte per 16 dimensions, so a 1,024-dimension float32 vector drops from 4,096 bytes to 64. Full vectors move to disk and are read only to rescore the top candidates, trading some recall and latency for a much smaller memory footprint.

What is the Milvus collection unit limit?

Milvus limits the total of shards × partitions across every collection, 65,536 by default (rootCoord.maxGeneralCapacity). 4,000 collections with 2 shards and 4 partitions each already use 32,000 units, so a collection-per-tenant design can hit this limit while QueryNodes still have plenty of memory.

Related reading

More AI builder tools

  • LLM Token Counter — Count tokens for GPT, Claude, and Gemini, check context-window fit, and see what the prompt costs.
  • LLM API Cost Calculator — Compare per-request and monthly API costs across OpenAI, Anthropic, and Google models, with prompt caching.
  • LLM VRAM Calculator — Estimate GPU memory for weights, KV cache, and fine-tuning, and see which GPUs a model fits on.
  • AI Agent Cost Calculator — Model the real cost of agent loops: growing context, tool results, prompt caching, retries, and sub-agents.