Vendor case study · Milvus 2.5

Milvus Incident Patterns and Lessons

Five common Milvus production incidents, from a hot Proxy and insert bursts to broker throttling and etcd growth, each traced from symptom to mechanism to lesson.

These are anonymised, generic patterns built from the Milvus 2.5 architecture, not a diagnosis of any particular cluster. Each follows the same shape: the symptom you see, the mechanism behind it, and the lesson. Thresholds depend on your workload, so none of this implies a universal safe rate.

01 / Follow the pressure

Five incident patterns

Pattern 1

One Proxy runs hot while the fleet looks idle

Symptom: Some clients time out; fleet-average Proxy CPU looks fine.

A connection-level load balancer distributes connections, not requests. Long-lived gRPC connections can pin heavy work to a few Proxies.

Proxy Client / SDK

  1. 1

    Connections concentrate work: A few busy connections favour a few Proxies.

    gRPC carries many requests over one long-lived HTTP/2 connection. A Layer 4 load balancer picks a backend per connection, so every request on that connection keeps reaching the same Proxy.

    Evidence and next action
    Evidence to collect
    Compare connections, requests, bytes and in-flight work per Proxy. Check session affinity, endpoint readiness and the SDK's balancing policy. Equal connection counts do not mean equal request cost.
    Next action
    Work out whether the imbalance is connection placement, one hot tenant, or unusually expensive requests before changing replica counts.

    Watch out: A Kubernetes Service or network load balancer does not rebalance individual gRPC calls on existing connections.

  2. 2

    One Proxy becomes the queue: Average CPU looks healthy while one Proxy struggles.

    The busy Proxy validates and serializes inserts, dispatches searches and merges results. Large batches, many query vectors, big top-k values and many output fields make some requests far more expensive than others.

    Evidence and next action
    Evidence to collect
    Compare per-Proxy CPU, throttled CPU time, memory, bytes, latency, errors and operation mix. Separate time spent locally from time waiting on QueryNodes.
    Next action
    Bound expensive batches and client concurrency. Scale Proxies only if the access layer is genuinely saturated and traffic can actually reach new instances.

    Watch out: Spreading requests evenly can still leave one expensive tenant dominating a Proxy.

  3. 3

    Retries amplify the pressure: Timeouts create more work than the original load.

    If clients retry slow calls immediately, more requests and connections land on the already busy path. A client timeout also doesn't tell you whether a write failed before or after it was accepted.

    Evidence and next action
    Evidence to collect
    Track attempted versus successful calls, retries per original operation, outstanding requests and new connections over the same window.
    Next action
    Use bounded concurrency, deadlines, exponential backoff with jitter and a retry budget. Reconcile ambiguous writes using application semantics.

    Watch out: Restarting every client or Proxy to reshuffle connections causes a reconnect storm and drops in-flight work.

  4. 4

    Rebalance deliberately: Fix distribution first, then check total capacity.

    A request-aware Layer 7 gRPC balancer distributes individual calls. Client-side balancing also works when the SDK supports it and can discover every real Proxy endpoint.

    Evidence and next action
    Evidence to collect
    After a controlled change, compare the busiest Proxy with the fleet, then check success rate and tail latency. Verify TLS, auth, deadlines and draining through the new path.
    Next action
    Test the supported balancing approach, keep connection pools bounded and drain gradually. Setting round-robin against a single virtual IP proves nothing about per-Proxy balance.

    Watch out: Better Proxy distribution admits more writes and can expose the next bottleneck in the broker, DataNodes or QueryNodes.

Lessons

  • Compare the busiest instance with the fleet, never just the average.
  • Layer 4 balancing places connections; gRPC multiplexes many requests over each one.
  • Fixing distribution admits more work, so expect the bottleneck to move downstream next.

gRPC load balancing ↗Proxy configuration ↗

Pattern 2

An insert burst slows down search

Symptom: Search p99 climbs during a bulk load, although search QPS is flat.

In Milvus 2.5 the write stream also feeds the query layer, so recent data is searchable. Inserts therefore cost QueryNode memory and CPU.

Log broker QueryNode Proxy

  1. 1

    One stream, two consumers: Every insert creates work on both the persistence and the query path.

    The Proxy publishes changes to the broker. DataNodes consume them for persistence, and QueryNodes consume the same channels to maintain growing segments. QueryNodes are doing ingestion work even when search QPS is flat.

    Evidence and next action
    Evidence to collect
    Correlate insert bytes and rows with DataNode and QueryNode channel lag. Identify which QueryNodes own the growing data for the loaded collections.
    Next action
    Budget QueryNode memory and CPU for fresh writes as well as historical data. Put ingestion and search metrics on the same timeline.

    Watch out: Adding DataNodes speeds up persistence without adding any capacity to an overloaded QueryNode.

  2. 2

    Fresh data competes with searches: Memory, CPU or one hot channel limits the query layer.

    Applying mutations and maintaining growing data use the same resources as searching. Uneven channel ownership overloads a few QueryNodes. A slow consumer also delays the visibility timestamp a search needs, so latency is not always vector computation.

    Evidence and next action
    Evidence to collect
    Separate search execution and queue time from consistency waiting. Check per-node memory, CPU, growing-segment size, channel ownership, consumer lag and handoff progress.
    Next action
    Slow the insert burst while you diagnose. Add query capacity only when work can actually be placed on it; look for hot channels first.

    Watch out: More nodes don't split a hot channel. Weakening consistency returns older results but doesn't remove the ingestion bottleneck.

  3. 3

    Protection reaches the writer: Query-layer pressure throttles or rejects inserts.

    Milvus reduces admitted writes when memory protection or stream-progress protection triggers. Time-tick delay measures how far internal progress timestamps lag; it is not a Kafka offset count.

    Evidence and next action
    Evidence to collect
    Read the actual rejection reason. Inspect effective quota settings, QueryNode and DataNode watermarks and time-tick delay alongside broker health.
    Next action
    Keep protection on while reducing insert traffic. Set explicit insert, upsert and delete limits for the operations you use, and reserve query capacity for your freshness target.

    Watch out: Raising watermarks or disabling protection turns controlled rejection into out-of-memory crashes. Longer timeouts don't help consumers catch up.

  4. 4

    Recover both paths: Recovery is done when lag falls and search latency recovers.

    Reducing new traffic leaves room to process the backlog. Persistence, indexing and the handoff to sealed data must also progress.

    Evidence and next action
    Evidence to collect
    Confirm sustained catch-up, memory headroom, fresh-data visibility and search latency under representative traffic.
    Next action
    Ramp writes gradually once lag and queues fall. If execution is still saturated, add correctly placed QueryNodes or replicas within memory and broker budgets.

    Watch out: Extra query replicas multiply memory and broker egress. They are not a free fix for an insert-driven overload.

Pattern 3

The log broker throttles and everything downstream lags

Symptom: Inserts slow down, consumer lag grows, and searches wait for freshness.

A managed Kafka or Pulsar cluster has its own network, quota and storage limits. Find the limiting broker resource before touching Milvus settings.

Log broker QueryNode

  1. 1

    Identify the kind of limit: Network shaping, quotas and slow consumers need different remedies.

    Cloud-managed brokers can queue or drop packets once an instance exceeds its network allowance. Kafka quotas deliberately delay clients that exceed configured rates. High consumer lag alone proves neither.

    Evidence and next action
    Evidence to collect
    Look at the affected broker in the matching time window. Separate bandwidth, packet-rate, connection-tracking, quota and CPU or disk pressure using your provider's broker metrics.
    Next action
    Confirm the broker type, Kafka version and which exact metric or alarm fired before changing anything.

    Watch out: A broker network alarm is not evidence that a Milvus rate limit is responsible.

  2. 2

    Broker delay spreads: Slow produces hurt inserts; slow fetches hurt consumers.

    Proxies publish changes while DataNodes and QueryNodes fetch them, and Kafka spends resources on its own replication. Hot partition leaders can concentrate load even when cluster averages look normal.

    Evidence and next action
    Evidence to collect
    Compare per-broker ingress and egress, partition-leader placement, produce and fetch latency, request queues and replication health.
    Next action
    Reduce offered traffic at the application or Milvus admission layer while you find the constrained resource. Check for connection churn if connection tracking is involved.

    Watch out: Changing the Proxy load balancer does nothing for Kafka partition leaders.

  3. 3

    Consumers lose freshness: Broker delay becomes query latency and write backpressure.

    When fetching falls behind, DataNodes persist later and QueryNodes learn about new data later. Searches that need a newer visibility timestamp wait, and Milvus protection may then reduce writes even though the broker was the original cause.

    Evidence and next action
    Evidence to collect
    Line up broker delay, consumer lag, time-tick delay, search waiting and insert rejections on one timeline. Missing consumer-lag metrics do not mean zero lag.
    Next action
    Restore a sustainable arrival rate with headroom to catch up. Confirm the oldest position a consumer needs is still inside broker retention.

    Watch out: Blind retries add traffic, and shortening retention can delete changes a recovering consumer still needs.

  4. 4

    Relieve the measured bottleneck: Scale and rebalance for the resource that was actually limited.

    More broker capacity addresses a demonstrated network, CPU or storage limit. After expanding, partitions must be reassigned so existing traffic uses the new brokers.

    Evidence and next action
    Evidence to collect
    Check the constrained metric improves per broker, placement converges, replication is healthy and consumers catch up.
    Next action
    Follow your provider's supported expansion and partition-reassignment procedures. Validate batching, compression or quota changes against the latency budget.

    Watch out: Reassignment itself consumes network and disk. Don't weaken acknowledgements or replication to buy throughput.

Lessons

  • Network shaping, Kafka quotas and a slow consumer look alike but need different fixes.
  • Broker delay turns into search waiting and then write backpressure inside Milvus.
  • New brokers only help once hot topic partitions actually move onto them.

Kafka quotas ↗Milvus Kafka settings ↗

Pattern 4

etcd fills up as collections grow

Symptom: etcd raises a NOSPACE alarm or approaches its quota while workers have spare capacity.

All Milvus metadata lives in one etcd keyspace, copied to every member. More collections, partitions and segments mean more keys, and old revisions linger until compacted.

etcd RootCoord

  1. 1

    Separate allocated from in-use: The backend file can be much bigger than the live data.

    etcd keeps superseded revisions until history is compacted, and freed pages stay inside the backend file until defrag. Allocated size, in-use size and the quota are three different numbers.

    Evidence and next action
    Evidence to collect
    Compare allocated backend bytes, in-use bytes, the effective quota and volume free space on every member, not summed across members.
    Next action
    Check whether history compaction is running and when it last succeeded before interpreting in-use bytes.

    Watch out: Adding members or growing the volume changes neither the keyspace size nor the backend quota.

  2. 2

    Retire old history: Revision compaction frees space inside the file.

    Compaction removes superseded key versions up to a retention boundary. Current keys stay. The freed pages become reusable but the file does not shrink.

    Evidence and next action
    Evidence to collect
    Confirm the retention mode and window, and allow for watchers that are still catching up.
    Next action
    Compact history using the procedure for your etcd version, then compare in-use bytes again.

    Watch out: Milvus segment compaction does not compact etcd history. They are unrelated operations.

  3. 3

    Return pages to disk: Defrag shrinks the allocated file.

    Defrag rewrites the backend so unused pages go back to the filesystem. It does not delete live Milvus metadata.

    Evidence and next action
    Evidence to collect
    Measure allocated size per member before and after, plus how quickly it grows back.
    Next action
    Defrag one member at a time in a controlled window, keeping quorum, because a member blocks reads and writes while it runs.

    Watch out: Never delete Milvus keys directly from etcd, and don't treat a new rootPath as a cleanup.

  4. 4

    Plan for the real footprint: If in-use stays high, the metadata is simply large.

    Collections, shards, partitions, segments and index tasks all create keys. When maintenance leaves the footprint high, the cluster has reached a metadata boundary.

    Evidence and next action
    Evidence to collect
    Track in-use bytes after maintenance against object counts and DDL churn over weeks.
    Next action
    Reduce object overhead (fewer tiny collections, partition keys), or route new tenants to a separately budgeted cluster.

    Watch out: Raising the quota is a capacity change to validate, not a fix, and routing new tenants away does not shrink the old cluster.

etcd space explorer

See what each maintenance step can reclaim

One member’s backendDefault quota: 2 GiB
In useReusable pagesBelow quota
Allocated
1.90 GiB
In use
1.70 GiB
Reusable
0.20 GiB

Illustrative sizes, not measurements. These controls do not run anything.

Example 1 / 3 · Retained revision history

In-use bytes may include old history.

1.7 GiB is in use, but superseded revisions have not been retired yet. The allocated backend is close to the 2 GiB default quota.

What it means: Check history retention and the last successful compaction.

Lessons

  • Milvus segment compaction, etcd revision compaction and etcd defrag are three different operations.
  • Adding etcd members copies the keyspace; it never divides it.
  • If in-use bytes stay high after compaction, the footprint is real: plan capacity, don't repeat maintenance.

Milvus etcd settings ↗etcd maintenance ↗etcd limits ↗

Pattern 5

One collection per tenant stops scaling

Symptom: Creating collections gets slower and hits guards, while QueryNodes are mostly idle.

Per-tenant collections multiply shards × partitions, segments and metadata. The control plane saturates long before search compute does.

RootCoord QueryCoord Client / SDK

  1. 1

    Count the real units: Each collection costs shards × partitions units.

    Milvus guards the total of shards × partitions across all collections (65,536 by default). 4,000 collections with 2 shards and 4 partitions each already use 32,000 units.

    Evidence and next action
    Evidence to collect
    Sum shards × partitions for every collection, including default partitions, and compare with rootCoord.maxGeneralCapacity.
    Next action
    Use the capacity counter on the scaling page with your real shape.

    Watch out: Raising the guard relaxes admission; it adds no memory, channels or metadata capacity.

  2. 2

    Consolidate compatible tenants: A partition key shares physical partitions.

    Tenants with the same schema and index policy can share one collection, routed by a partition key, with a mandatory tenant filter on every request.

    Evidence and next action
    Evidence to collect
    Group tenants by schema, index settings, size and isolation requirements.
    Next action
    Move small compatible tenants onto a shared collection; keep separate collections only where policy genuinely differs.

    Watch out: A partition key is not an authorization boundary. The application must enforce tenant identity.

  3. 3

    Isolate by the right boundary: Resource groups isolate compute, not metadata.

    Resource groups give tenants dedicated QueryNodes, but databases, resource groups and partitions all share one control plane and one etcd.

    Evidence and next action
    Evidence to collect
    Name the resource you are isolating: query compute, metadata capacity, upgrades or blast radius.
    Next action
    Use resource groups for compute isolation and separate clusters for metadata or failure-domain isolation.

    Watch out: Two clusters sharing one etcd ensemble still share its quota.

  4. 4

    Route by placement, not hash: New clusters need explicit, stable routing.

    Once tenants live on several clusters, the application owns routing, migration and any cross-cluster query merge.

    Evidence and next action
    Evidence to collect
    Track per-cluster metadata growth, DDL latency and unit usage against a tested admission budget.
    Next action
    Keep a durable tenant → cluster map. Treat any move as a migration: backfill, capture changes, verify, switch, with a rollback plan.

    Watch out: Changing a modulo hash silently re-routes existing tenants to clusters that don't have their data.

02 / A recurring lesson

Why “add more workers” can make it worse

Every change relieves one resource and loads another. Test these consequences against your workload.

ChangePressure it can create elsewhere
Add Proxies or fix client balancingMore accepted writes can overwhelm the broker or either consumer path. Watch the bottleneck move.
Add collection replicasEach copy consumes memory and consumes the stream, raising broker egress.
Add DataNodesPersistence may improve while QueryNodes still lag. More flush and compaction work adds object-store load.
Add brokersHot partitions that aren't moved leave new capacity idle, and reassignment itself uses network and disk.
Increase queues, timeouts or retriesWaiting work piles up. Completed throughput stays flat while tail latency and memory rise.
Force frequent flushesMany small sealed segments add index builds, scheduling and metadata, and don't fix a slow broker.

Stabilize → catch up → scale → ramp

  1. Limit offered work. Pause nonessential bulk ingestion and cap concurrency, keeping memory and progress protections on.
  2. Find the busiest instance and the limiting stage, not just fleet averages.
  3. Change one constrained resource at a time, allowing for cold loads, broker rebalancing and background compaction.
  4. Require recovery evidence: falling lag and queues, returning memory headroom, and real insert and search probes that meet your targets.
  5. Ramp gradually, and scale down only after everything has converged.

03 / From symptom to subsystem

Quick triage

Start with the failing path, then narrow the scope of the change.

01Search latency rises after an ingestion burstQuery / stream · QueryNode

Inspect first

Compare client and server latency with QueryNode CPU, memory, growing-data volume and consumer lag. Is the request waiting on visibility or executing?

Respond, then verify

Reduce incoming pressure, restore stream progress, then tune search effort or add query capacity. Only weaken consistency if the application accepts stale reads.

Supporting reference ↗

02Inserts are rejected or ingestion stops progressingIngestion · DataNode

Inspect first

Read the actual rejection reason. Check quotas, time-tick delay, QueryNode and DataNode memory, and broker publish failures.

Respond, then verify

Fix the limiting resource and confirm lag falls before increasing traffic. Disabling backpressure turns controlled rejection into worker failure.

Supporting reference ↗

03A collection is stuck loadingPlacement · QueryCoord

Inspect first

Inspect load progress, replica count, resource-group membership and per-node capacity. Look for object-store download failures or missing index files.

Respond, then verify

Fix capacity or placement constraints and let loading converge. Restarting all QueryNodes at once amplifies downloads and memory pressure.

Supporting reference ↗

04Index builds or compactions keep accumulatingBackground work · IndexNode

Inspect first

Separate queued work from tasks that keep failing. Check worker resources, object-store errors and the rate of new segments.

Respond, then verify

Fix task errors first, then add workers or reschedule the workload. Frequent small flushes create avoidable indexing overhead.

Supporting reference ↗

05A coordinator or storage dependency failsRecovery · etcd

Inspect first

Identify the failed dependency before restarting workers: etcd quorum and leases, coordinator leadership, broker health and object-store access.

Respond, then verify

Restore the dependency with its own recovery procedure. A standby coordinator must take leadership and reload state.

Supporting reference ↗

06Deletes succeed but storage usage stays highData lifecycle · DataCoord

Inspect first

Check delete visibility separately from compaction and garbage collection. Inspect retention settings.

Respond, then verify

Let background work finish. Never delete objects by hand; use logical backup and test restores into a separate target.

Supporting reference ↗