Vendor case study · Milvus 2.5

Milvus Tuning Reference

The Milvus 2.5 settings that matter for capacity, search, ingestion and resilience, with what each changes, its trade-off and the documented defaults.

01 / Tune with intent

The knobs that matter

Names are specific to Milvus 2.5. Check the effective values for your exact patch and Helm chart before changing anything.

Knob / scopeWhat it changesTrade-off
queryNode.replicasHelm deploymentChanges the total QueryNode pod count.This is not collection replication. Wait for segment placement and load readiness. Reference ↗
replica_numberCollection load APISets the number of loaded collection copies.Each copy needs its own memory. Check replica and resource-group placement before raising it. Reference ↗
queryCoord.autoBalanceMilvus configurationEnables automatic segment balancing.Balancing moves data and consumes resources; inspect placement before disabling it. Reference ↗
rootCoord.maxGeneralCapacityMilvus configurationGuards the sum of shards × partitions across collections.Default 65,536 units. Raising it does not add worker or etcd capacity. Reference ↗
consistency_levelCollection / requestChooses Strong, Bounded, Session or Eventually visibility.Stronger visibility may wait for stream progress. Bounded is the documented default. Reference ↗
ef / nprobeIndex search parametersControls search effort for HNSW and IVF respectively.More effort trades latency for recall. Validate on representative ground truth. Reference ↗
queryNode.scheduler.maxReadConcurrentRatioMilvus configurationSets read-task concurrency relative to CPU count.More concurrency can add contention. Compare throughput, queueing and p99 together. Reference ↗
queryNode.mmap.vectorField / vectorIndexMilvus configurationLets supported vector data or indexes use memory mapping.Cuts resident memory at the cost of local disk and page-fault sensitivity. Benchmark warm and cold searches. Reference ↗
num_shardsCollection creation APISets write sharding for the collection.Shards add ingestion parallelism and channel overhead. Plan before creation; adding pods does not split a hot shard. Reference ↗
dataCoord.segment.maxSize / sealProportionMilvus configurationInfluences when segments are sealed.Small segments add scheduling and index overhead; large ones make each task and load heavier. Reference ↗
dataCoord.compaction.enableAutoCompactionMilvus configurationControls automatic segment compaction.Turning it off defers cost while fragmentation and storage grow. Reference ↗
dataNode.replicas / indexNode.replicasHelm deploymentChanges ingestion and indexing worker counts independently.Follow channel distribution for DataNodes and task backlog for IndexNodes; both share storage bandwidth. Reference ↗
quotaAndLimits.limitWriting.memProtection.enabledMilvus configurationApplies write backpressure based on worker memory.Find the resource that triggered it. Don't remove protection as a substitute for capacity. Reference ↗
--auto-compaction-retentionetcd server flagControls how much revision history etcd keeps.Retention 0 disables automatic history compaction. Separate from Milvus segment compaction. Reference ↗
--quota-backend-bytesetcd server flagSets the backend space quota.A larger volume does not raise this quota, and extra members do not shard it. Reference ↗
rootCoordinator.activeStandby.enabledHelm chartEnables coordinator HA together with a replica count.Standbys speed up failover; they do not add serving throughput. Reference ↗

Deployment replica counts, collection replicas and internal concurrency settings are separate controls. Dynamic reload support varies by key. Configuration guide ↗

02 / Know the guardrails

Documented limits, and how to read them

Defaults pinned to Milvus 2.5.0 and etcd 3.5. A guard is an admission limit; your measured operating limit is usually lower.

BoundaryDocumented baselineHow to interpret it
Aggregate object guardrootCoord.maxGeneralCapacity65,536 unitsSum of shards × partitions across all collections. Raising it only relaxes admission; it does not add memory, channels or metadata space. Capacity formula ↗
Collection count guardsquotaAndLimits.limits.maxCollectionNum65,536A separate collection ceiling. The unit guard and your measured operating budget usually bind much earlier. 2.5.0 configuration ↗
Per-collection shapeproxy.maxShardNum: 16rootCoord.maxPartitionNum: 1,024Partitions multiply the unit count, so a manual partition per tenant is not an unlimited tenancy model. 2.5.0 configuration ↗
Logical databasesrootCoord.maxDatabaseNum: 64A namespace and access boundary inside one cluster. More databases do not add etcd capacity. 2.5.0 configuration ↗
etcd backend quotaetcd default: 2 GiBSuggested max: 8 GiBRead the effective quota for your deployment. Larger quotas need validation, and adding members does not multiply them. etcd system limits ↗

03 / Keep it running

Change one failure domain at a time

  1. 01Establish a baseline

    Record p95/p99 latency, errors, lag, loaded replicas and pending tasks.

  2. 02Check recovery prerequisites

    Healthy etcd quorum, durable storage, broker retention and a tested restore path.

  3. 03Roll a small change

    Use your pinned Helm chart or Operator spec, and respect disruption budgets.

  4. 04Wait for convergence

    Confirm segment loading, replica health and real insert and search probes before the next step.