Vendor case study · Milvus 2.5

Scaling Milvus: Nodes, Replicas and Tenancy

Model QueryNodes versus collection replicas, count shard × partition capacity units, and choose between partition keys, collections and separate clusters for multi-tenancy.

01 / Change the shape of the cluster

More nodes. More replicas. Different effects.

A simplified memory model for one loaded collection. QueryNodes spread one copy; collection replicas add whole copies, each needing its own group of nodes.

Memory per node
Replica placementFits, but can't lose a node
Replica 12 nodes · 60 GB full copy
QN-0130.0 GB
QN-0230.0 GB
Replica 22 nodes · 60 GB full copy
QN-0330.0 GB
QN-0430.0 GB
Total memory demand
120 GB
2 copies
Smallest group holds
96 GB
Usable per node
48 GB
25% held back

Each replica group holds a complete, searchable copy. More replicas add read throughput and resilience but multiply memory. Losing one node would leave a replica that can't fit. Add headroom before scaling in.

How this model works

Nodes are split as evenly as possible between replicas. Each replica must hold the whole index within its own group, and 25% of each node's memory is held back for growing data, queries and rebuilds. The smallest group limits capacity. It ignores uneven segment sizes, other collections and placement constraints, and predicts neither QPS nor failover time.

Sizing a whole deployment? The vector database sizing calculator works out index memory from vector count, dimensions and index type first.

02 / Scale the right role

Each worker role scales for a different reason

Choose the bottleneck before the knob

Scale out

Add QueryNodes to spread loaded segments and compute. Add collection replicas separately, when extra complete copies serve a throughput or availability goal.

Check replica placement, CPU, memory and p99.

Scale up

Give each node more CPU or memory for a working-set or per-task constraint, and make sure Kubernetes can place the larger pods.

Trade-off: larger failure domains and slower restarts.

Scale in

Remove one node at a time. Confirm each replica still fits, wait for load and balance to converge, and run search probes.

Low CPU alone does not prove spare capacity.

03 / Find the scaling boundary

Yes, Milvus scales horizontally, up to a boundary

Worker capacity, metadata capacity and tenant isolation need different scaling decisions.

Within a cluster

Add Proxy, QueryNode, DataNode or IndexNode capacity for whichever role is saturated.

At the metadata boundary

etcd members replicate the keyspace. Adding members neither divides its size nor multiplies its quota.

Across clusters

Application routing can place tenants or collections on independently budgeted clusters.

Start with the symptom

What should you scale?

Scale the query path · QueryNode

Add QueryNodes for capacity; add collection replicas for more copies.

Check first

Check loaded segment memory, query CPU, p95/p99 latency, replica placement and the Proxy before scaling. Benchmark representative filters, top-k, indexes, concurrency and growing data.

Next decision

Add QueryNodes to spread work within a replica. Increase collection replicas when the extra copies fit and measured throughput benefits. Scale Proxies when request handling or result merging is saturated.

A replica must collectively fit its assigned data. More nodes do not give a single query a linear speedup.

Replica model ↗Worker scaling ↗

04 / Collection capacity is a product

Count shards × partitions, not just collections

Milvus guards the sum of shards × partitions across every collection. Enter your own layout; for mixed collections, add up each collection’s shards × partitions.

8,000 / 65,536 units

1,000 collections × 2 shards × 4 partitions

Within the guard

12.2% of the guard. At most 7,192 more collections of this shape fit. The guard is an admission limit, not a safe operating target: metadata and control-plane latency usually bind earlier.

Count real physical partitions, including the default partition. This is a guardrail calculation, not an etcd size or throughput estimate. Calculation reference ↗

05 / Multi-tenancy

Choose the boundary that matches what you isolate

Partition keys, separate collections and separate clusters isolate different resources. Pick by tenant isolation, query scope and the resource you need to separate.

Application enforces tenant identity + filter
One collectionTenant A · B · CPartition-key routing

One control plane · one etcd budget

Share physical partitions across compatible tenants.

Use when

Tenants share a schema and index policy, and application-enforced tenant filtering meets your isolation requirements. A partition key narrows search without one physical partition per tenant.

Trade-off

Tenants may hash into the same physical partition. Every request must carry tenant authorization and a filter; the partition key is not a security boundary. Very large tenants may still need their own placement.

Collection overhead is consolidated. The control plane and etcd keyspace are still shared.

Partition-key behavior ↗Tenancy trade-offs ↗