Pattern 1
One Proxy runs hot while the fleet looks idle
Symptom: Some clients time out; fleet-average Proxy CPU looks fine.
A connection-level load balancer distributes connections, not requests. Long-lived gRPC connections can pin heavy work to a few Proxies.
Proxy Client / SDK
- 1
Connections concentrate work: A few busy connections favour a few Proxies.
gRPC carries many requests over one long-lived HTTP/2 connection. A Layer 4 load balancer picks a backend per connection, so every request on that connection keeps reaching the same Proxy.
Evidence and next action
- Evidence to collect
- Compare connections, requests, bytes and in-flight work per Proxy. Check session affinity, endpoint readiness and the SDK's balancing policy. Equal connection counts do not mean equal request cost.
- Next action
- Work out whether the imbalance is connection placement, one hot tenant, or unusually expensive requests before changing replica counts.
Watch out: A Kubernetes Service or network load balancer does not rebalance individual gRPC calls on existing connections.
- 2
One Proxy becomes the queue: Average CPU looks healthy while one Proxy struggles.
The busy Proxy validates and serializes inserts, dispatches searches and merges results. Large batches, many query vectors, big top-k values and many output fields make some requests far more expensive than others.
Evidence and next action
- Evidence to collect
- Compare per-Proxy CPU, throttled CPU time, memory, bytes, latency, errors and operation mix. Separate time spent locally from time waiting on QueryNodes.
- Next action
- Bound expensive batches and client concurrency. Scale Proxies only if the access layer is genuinely saturated and traffic can actually reach new instances.
Watch out: Spreading requests evenly can still leave one expensive tenant dominating a Proxy.
- 3
Retries amplify the pressure: Timeouts create more work than the original load.
If clients retry slow calls immediately, more requests and connections land on the already busy path. A client timeout also doesn't tell you whether a write failed before or after it was accepted.
Evidence and next action
- Evidence to collect
- Track attempted versus successful calls, retries per original operation, outstanding requests and new connections over the same window.
- Next action
- Use bounded concurrency, deadlines, exponential backoff with jitter and a retry budget. Reconcile ambiguous writes using application semantics.
Watch out: Restarting every client or Proxy to reshuffle connections causes a reconnect storm and drops in-flight work.
- 4
Rebalance deliberately: Fix distribution first, then check total capacity.
A request-aware Layer 7 gRPC balancer distributes individual calls. Client-side balancing also works when the SDK supports it and can discover every real Proxy endpoint.
Evidence and next action
- Evidence to collect
- After a controlled change, compare the busiest Proxy with the fleet, then check success rate and tail latency. Verify TLS, auth, deadlines and draining through the new path.
- Next action
- Test the supported balancing approach, keep connection pools bounded and drain gradually. Setting round-robin against a single virtual IP proves nothing about per-Proxy balance.
Watch out: Better Proxy distribution admits more writes and can expose the next bottleneck in the broker, DataNodes or QueryNodes.