Milvus / segments
Segment compaction
Rewrites segment data, merges small segments and drops deleted rows. It does not touch etcd’s revision history or backend file.
Vendor case study · Milvus 2.5
How a Milvus segment moves from growing to sealed, flushed, indexed and compacted, and why deletes and storage reclamation happen at different times.
01 / The unit of work
A collection is sharded into channels, and its data is stored, indexed and scheduled in segments. Step through the five stages.
Mutations accumulate in a growing segment. QueryNodes consume the stream and can serve this data once the request's visibility requirement is met.
Milvus 2.5 can build interim indexes on growing data.
Reference ↗02 / Why it matters
An insert passes four separate milestones: acknowledged (in the log broker), visible (consumed by the QueryNode that owns the channel, subject to the request’s consistency level), durable as a segment file (flushed by a DataNode), and indexed and loaded. Each is driven by a different component, so each can lag independently.
Deletes work the same way. A delete is a mutation that hides rows immediately, but the bytes stay in sealed segments until compaction rewrites them, the replacement segment is indexed and loaded, and garbage collection removes the obsolete files after a retention window. “Deleted but storage didn’t shrink” is usually this pipeline still running, not a bug.
Segment shape matters too. Many small segments, often caused by forcing a flush after every batch, mean more index builds, more scheduling and more metadata keys in etcd. That links a data-path habit directly to the metadata-growth incident on the incident patterns page.
03 / Three kinds of cleanup
A full etcd database, old revisions and a large segment count are different diagnoses. Know which cleanup does what.
Milvus / segments
Rewrites segment data, merges small segments and drops deleted rows. It does not touch etcd’s revision history or backend file.
etcd / history
Removes superseded key versions up to a retention boundary. Current keys stay, and the backend file keeps the freed pages.
etcd / allocated space
Returns unused backend pages to the filesystem on each member. It neither removes live keys nor replaces history compaction.
DataCoord plans compaction and DataNodes execute it. Replacement segments may need index builds and loading before the old objects can be collected, and this background work competes with ingestion for compute and object-store bandwidth.
| Milvus 2.5.0 setting | Why an operator cares |
|---|---|
dataCoord.enableCompaction, dataCoord.compaction.enableAutoCompaction | Turn compaction and its automatic scheduling on. Check whether tasks are disabled, queued or failing before adding workers. |
dataCoord.compaction.maxParallelTaskNum | Bounds parallel compaction work. Raise it only after checking DataNode memory, CPU, storage I/O and ingestion latency. |
dataCoord.segment.maxSize and seal policies | Shape segments. Frequent explicit flushes and tiny segments multiply indexing, scheduling and metadata churn. |
dataCoord.enableGarbageCollection, dataCoord.gc.dropTolerance | Retention and delayed cleanup mean a finished delete or compaction does not free object storage immediately. |