Milvus and etcd: Compaction vs Defrag, and Why Metadata Keeps Growing
Milvus segment compaction, etcd revision compaction and etcd defrag are three different operations that get mixed up during a NOSPACE incident. Here is what each reclaims, and what to do when none of them helps.

Here is an incident pattern that comes up often with Milvus in production. A cluster has been running happily for months. Tenants are added, each with their own collection. Then etcd raises a NOSPACE alarm and metadata writes start failing, even though QueryNodes have plenty of memory and CPU is quiet.
The instinctive fixes are to "run compaction", add etcd members or grow the volume. Often none of them help, because three different operations get lumped together as "compaction". This post untangles them, using the Milvus 2.5 architecture. To try the scenarios yourself, the etcd space explorer in the Milvus lab animates each one.
Where Milvus keeps its metadata
Milvus stores vectors in object storage, but everything about the data lives in etcd: collections and schemas, partitions, segment records, channel checkpoints, index tasks, and service registrations. (See how the Milvus architecture fits together for the full picture.)
Three facts about etcd drive this incident:
- Every member holds the whole keyspace. etcd replicates for fault tolerance; it does not shard. Five members hold five copies of the same data.
- It keeps history. Every update creates a new revision, and old revisions stay until revision compaction retires them.
- Its backend file doesn't shrink on its own. Freed pages are reused internally, but the file only gets smaller after defragmentation.
etcd also has a backend quota, 2 GiB by default, with 8 GiB as the suggested maximum for normal environments. When a member's backend exceeds it, etcd raises NOSPACE and rejects writes.
Three operations with similar names
Operation | What it does | What it doesn't do |
|---|---|---|
Milvus segment compaction | Rewrites segments in object storage: merges small ones, drops deleted rows | Touch etcd's revision history or backend file |
etcd revision compaction | Removes superseded key versions up to a retention boundary | Shrink the backend file, or remove current keys |
etcd defragmentation | Rewrites the backend so unused pages go back to disk | Remove history, or remove current keys |
Milvus segment compaction is about vectors and deletes. The other two are etcd housekeeping. Only the etcd pair affects the quota, and they measure different things:
- In-use bytes fall after revision compaction.
- Allocated bytes (the file size, which the quota checks) fall after defrag.
Three ways the same alarm happens
Look at allocated size, in-use size, the quota and volume free space on each member, not summed across members. The pattern tells you which fix applies.
1. Retained history. In-use is high, but much of it is old revisions because compaction isn't running or its retention window is long. Revision compaction drops in-use sharply, then defrag shrinks the file. This is the easy case.
2. Fragmentation. Compaction already ran, so in-use is low, but the file is still large. Repeating compaction does nothing. Defrag, one member at a time, returns the space.
3. The footprint is real. In-use stays high after a successful compaction. Defrag reclaims only the small gap. There is no maintenance fix, because the cluster genuinely has that much metadata. This is the case that surprises teams, and it is common with collection-per-tenant designs.
allocated in-use
History retained 1.9 GiB 1.7 GiB → compact, then defrag
Fragmented 1.9 GiB 0.6 GiB → defrag
Real footprint 1.9 GiB 1.7 GiB → even after compaction: capacity problemIllustrative numbers against the 2 GiB default quota, not measurements.
When the metadata is real
If in-use stays high after maintenance, the fix is fewer objects or more budget.
Count shards × partitions, not collections. Milvus guards the total of shards × partitions across all collections, 65,536 units by default. 4,000 tenant collections with 2 shards and 4 partitions each use 32,000 units, and every one of those partitions brings segments, checkpoints and index records into etcd. The collection capacity counter lets you plug in your own layout.
Consolidate compatible tenants. Tenants with the same schema and index settings can share one collection routed by a partition key, with a tenant filter enforced on every request. That collapses thousands of collections into a handful. The partition key is not an authorization boundary, so the application must still enforce tenant identity.
Split at the metadata boundary. When tenants need isolation, or growth outpaces consolidation, place new tenants on a separate Milvus cluster with its own etcd. Two clusters sharing one etcd ensemble still share its quota. The application then owns a durable tenant-to-cluster map, so treat any later move as a migration.
Treat the quota as a capacity change. Raising --quota-backend-bytes can buy time, but larger backends affect recovery and defrag time, so validate it like any other capacity change rather than treating it as a cleanup.
What not to do
- Don't delete Milvus keys from etcd by hand. The coordinators assume their metadata is consistent.
- Don't point Milvus at a new
rootPathas a "cleanup". That is a new empty cluster sharing the same etcd quota, not a migration. - Don't grow the volume and expect the alarm to clear. A bigger disk doesn't raise the backend quota.
- Don't defrag every member at once. A member blocks reads and writes while it defrags, so go one at a time and keep quorum.
The lessons
- Milvus compaction, etcd compaction and etcd defrag are three different tools. Name the one you mean.
- Measure allocated and in-use bytes per member before and after each step. The gap tells you which tool applies.
- If in-use stays high after compaction, stop doing maintenance and plan capacity: fewer collections, partition keys, or separately budgeted clusters.
- Metadata growth is a design decision made when you choose a tenancy model, long before the alarm.
The full scenario, with the interactive etcd explorer and four other incident patterns, is in the Milvus incident patterns lab.


