Vendor case study · Milvus 2.5
Milvus Tuning Reference
The Milvus 2.5 settings that matter for capacity, search, ingestion and resilience, with what each changes, its trade-off and the documented defaults.
01 / Tune with intent
The knobs that matter
Names are specific to Milvus 2.5. Check the effective values for your exact patch and Helm chart before changing anything.
| Knob / scope | What it changes | Trade-off |
|---|---|---|
queryNode.replicasHelm deployment | Changes the total QueryNode pod count. | This is not collection replication. Wait for segment placement and load readiness. Reference ↗ |
replica_numberCollection load API | Sets the number of loaded collection copies. | Each copy needs its own memory. Check replica and resource-group placement before raising it. Reference ↗ |
queryCoord.autoBalanceMilvus configuration | Enables automatic segment balancing. | Balancing moves data and consumes resources; inspect placement before disabling it. Reference ↗ |
rootCoord.maxGeneralCapacityMilvus configuration | Guards the sum of shards × partitions across collections. | Default 65,536 units. Raising it does not add worker or etcd capacity. Reference ↗ |
consistency_levelCollection / request | Chooses Strong, Bounded, Session or Eventually visibility. | Stronger visibility may wait for stream progress. Bounded is the documented default. Reference ↗ |
ef / nprobeIndex search parameters | Controls search effort for HNSW and IVF respectively. | More effort trades latency for recall. Validate on representative ground truth. Reference ↗ |
queryNode.scheduler.maxReadConcurrentRatioMilvus configuration | Sets read-task concurrency relative to CPU count. | More concurrency can add contention. Compare throughput, queueing and p99 together. Reference ↗ |
queryNode.mmap.vectorField / vectorIndexMilvus configuration | Lets supported vector data or indexes use memory mapping. | Cuts resident memory at the cost of local disk and page-fault sensitivity. Benchmark warm and cold searches. Reference ↗ |
num_shardsCollection creation API | Sets write sharding for the collection. | Shards add ingestion parallelism and channel overhead. Plan before creation; adding pods does not split a hot shard. Reference ↗ |
dataCoord.segment.maxSize / sealProportionMilvus configuration | Influences when segments are sealed. | Small segments add scheduling and index overhead; large ones make each task and load heavier. Reference ↗ |
dataCoord.compaction.enableAutoCompactionMilvus configuration | Controls automatic segment compaction. | Turning it off defers cost while fragmentation and storage grow. Reference ↗ |
dataNode.replicas / indexNode.replicasHelm deployment | Changes ingestion and indexing worker counts independently. | Follow channel distribution for DataNodes and task backlog for IndexNodes; both share storage bandwidth. Reference ↗ |
quotaAndLimits.limitWriting.memProtection.enabledMilvus configuration | Applies write backpressure based on worker memory. | Find the resource that triggered it. Don't remove protection as a substitute for capacity. Reference ↗ |
--auto-compaction-retentionetcd server flag | Controls how much revision history etcd keeps. | Retention 0 disables automatic history compaction. Separate from Milvus segment compaction. Reference ↗ |
--quota-backend-bytesetcd server flag | Sets the backend space quota. | A larger volume does not raise this quota, and extra members do not shard it. Reference ↗ |
rootCoordinator.activeStandby.enabledHelm chart | Enables coordinator HA together with a replica count. | Standbys speed up failover; they do not add serving throughput. Reference ↗ |
Deployment replica counts, collection replicas and internal concurrency settings are separate controls. Dynamic reload support varies by key. Configuration guide ↗
02 / Know the guardrails
Documented limits, and how to read them
Defaults pinned to Milvus 2.5.0 and etcd 3.5. A guard is an admission limit; your measured operating limit is usually lower.
| Boundary | Documented baseline | How to interpret it |
|---|---|---|
| Aggregate object guard | rootCoord.maxGeneralCapacity65,536 units | Sum of shards × partitions across all collections. Raising it only relaxes admission; it does not add memory, channels or metadata space. Capacity formula ↗ |
| Collection count guards | quotaAndLimits.limits.maxCollectionNum65,536 | A separate collection ceiling. The unit guard and your measured operating budget usually bind much earlier. 2.5.0 configuration ↗ |
| Per-collection shape | proxy.maxShardNum: 16rootCoord.maxPartitionNum: 1,024 | Partitions multiply the unit count, so a manual partition per tenant is not an unlimited tenancy model. 2.5.0 configuration ↗ |
| Logical databases | rootCoord.maxDatabaseNum: 64 | A namespace and access boundary inside one cluster. More databases do not add etcd capacity. 2.5.0 configuration ↗ |
| etcd backend quota | etcd default: 2 GiBSuggested max: 8 GiB | Read the effective quota for your deployment. Larger quotas need validation, and adding members does not multiply them. etcd system limits ↗ |
03 / Keep it running
Change one failure domain at a time
- 01Establish a baseline
Record p95/p99 latency, errors, lag, loaded replicas and pending tasks.
- 02Check recovery prerequisites
Healthy etcd quorum, durable storage, broker retention and a tested restore path.
- 03Roll a small change
Use your pinned Helm chart or Operator spec, and respect disruption budgets.
- 04Wait for convergence
Confirm segment loading, replica health and real insert and search probes before the next step.