Vendor case study · Milvus 2.5

The Milvus Segment Lifecycle

How a Milvus segment moves from growing to sealed, flushed, indexed and compacted, and why deletes and storage reclamation happen at different times.

01 / The unit of work

A segment’s journey

A collection is sharded into channels, and its data is stored, indexed and scheduled in segments. Step through the five stages.

Recent data is already part of the search path.

Mutations accumulate in a growing segment. QueryNodes consume the stream and can serve this data once the request's visibility requirement is met.

Milvus 2.5 can build interim indexes on growing data.

Reference ↗

02 / Why it matters

Four milestones that are easy to confuse

An insert passes four separate milestones: acknowledged (in the log broker), visible (consumed by the QueryNode that owns the channel, subject to the request’s consistency level), durable as a segment file (flushed by a DataNode), and indexed and loaded. Each is driven by a different component, so each can lag independently.

Deletes work the same way. A delete is a mutation that hides rows immediately, but the bytes stay in sealed segments until compaction rewrites them, the replacement segment is indexed and loaded, and garbage collection removes the obsolete files after a retention window. “Deleted but storage didn’t shrink” is usually this pipeline still running, not a bug.

Segment shape matters too. Many small segments, often caused by forcing a flush after every batch, mean more index builds, more scheduling and more metadata keys in etcd. That links a data-path habit directly to the metadata-growth incident on the incident patterns page.

03 / Three kinds of cleanup

Compaction and defrag are different operations

A full etcd database, old revisions and a large segment count are different diagnoses. Know which cleanup does what.

Milvus / segments

Segment compaction

Rewrites segment data, merges small segments and drops deleted rows. It does not touch etcd’s revision history or backend file.

Milvus compaction ↗

etcd / history

Revision compaction

Removes superseded key versions up to a retention boundary. Current keys stay, and the backend file keeps the freed pages.

etcd history ↗

etcd / allocated space

Defragmentation

Returns unused backend pages to the filesystem on each member. It neither removes live keys nor replaces history compaction.

etcd defrag ↗

Compaction, retention and garbage-collection settings

DataCoord plans compaction and DataNodes execute it. Replacement segments may need index builds and loading before the old objects can be collected, and this background work competes with ingestion for compute and object-store bandwidth.

Milvus 2.5.0 settingWhy an operator cares
dataCoord.enableCompaction, dataCoord.compaction.enableAutoCompactionTurn compaction and its automatic scheduling on. Check whether tasks are disabled, queued or failing before adding workers.
dataCoord.compaction.maxParallelTaskNumBounds parallel compaction work. Raise it only after checking DataNode memory, CPU, storage I/O and ingestion latency.
dataCoord.segment.maxSize and seal policiesShape segments. Frequent explicit flushes and tiny segments multiply indexing, scheduling and metadata churn.
dataCoord.enableGarbageCollection, dataCoord.gc.dropToleranceRetention and delayed cleanup mean a finished delete or compaction does not free object storage immediately.

Pinned 2.5.0 configuration ↗