Below the API

Platform and Fleet

Advanced Advanced 1h Difficulty 4/5

Prerequisites 02, 03; helpful: Inference Engineering XI.04 and XII.07

The idea in one minute#

One engine’s metrics tell you about one replica. A platform has many replicas, several models, a router deciding who gets which request, an autoscaler deciding how many replicas exist, and tenants competing for all of it. Fleet observability is about the differences: between replicas, between what the scheduler was asked for and what it delivered, between tenants.

Three things deserve most of the attention: routing quality (did requests land where their cache was?), scaling signals (queue and KV pressure, not GPU utilization), and capacity that is paid for but not serving (pending pods, cold starts, idle GPUs).

An analogy#

Air traffic control. Each pilot watches their own instruments. The controller watches spacing, queues for each runway, which runway is closed, and whether another one should open. No single cockpit shows any of that.

A picture#

flowchart TB
  CL["Clients, tenants"] --> GW["Gateway<br/>auth, quotas, per-tenant accounting"]
  GW --> RT["Router / endpoint picker<br/>scores replicas by queue, KV usage, prefix match"]
  RT --> R1["Replica A<br/>engine metrics"]
  RT --> R2["Replica B"]
  RT --> R3["Replica C"]
  R1 --- G1[("GPU metrics")]
  R2 --- G2[("GPU metrics")]
  R3 --- G3[("GPU metrics")]
  AS["Autoscaler"] -->|"reads waiting, KV usage"| R1
  AS -->|"adds or removes replicas"| K8S["Kubernetes<br/>scheduler, DRA, node pools"]
  K8S --> R3
  class CL neutral
  class GW,RT,AS queue
  class R1,R2,R3 compute
  class G1,G2,G3 memory
  class K8S io

How it really works#

What each component should tell you#

ComponentMetricsQuestion answered
GatewayRequests, tokens and errors by tenant, model, route; 429s; time spent in the gatewayWho is using what; who is being throttled
RouterDecision time; chosen replica; score inputs; prefix-match rate; retries and fallbacksIs routing helping or scattering caches
EnginesLesson 03, per replicaWhich replica is the outlier
AutoscalerDesired vs current replicas; scale events; the signal value it acted onIs it reacting, and to what
KubernetesPod phase, pending time, restarts, OOM kills, node conditions (kube-state-metrics)Is capacity materializing
GPUsLesson 02, joined to podsIs the hardware under each replica healthy
Model lifecycleLoad time, download time, readiness time per model versionHow long is a cold start, really

Outliers: compare replicas with each other#

Fleet averages hide the broken one. For every important engine metric, plot the spread:

# Slowest replica's TTFT p95 against the fleet's
max by (model_name) (
  histogram_quantile(0.95, sum by (le, pod, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m]))))
/
histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))

# Load imbalance: busiest replica's share of waiting requests
max by (model_name) (vllm:num_requests_waiting) / clamp_min(avg by (model_name) (vllm:num_requests_waiting), 1)

A persistent outlier is usually hardware (lesson 02: throttling, a degraded link, a different GPU type), placement (a noisy neighbour, a wrong NUMA node), or routing (one replica attracts the long prompts).

Routing quality#

LLM routing is not round-robin: sending a conversation back to the replica that already holds its prefix in the KV cache skips most of prefill. In Kubernetes this is the job of the Gateway API Inference Extension — whose InferencePool API has been stable (v1) since its 1.0 release — and its endpoint picker, which as of 2026 is maintained in the llm-d project (a CNCF sandbox project since March 2026). The picker scrapes each replica’s queue depth and KV-cache usage and scores endpoints per request.

What to watch:

  • Fleet prefix-cache hit rate, and its split by replica. Good routing raises it; a rollout or a scale-up temporarily lowers it.
  • Router decision latency — it is on the path of every request.
  • Staleness of the router’s view: it acts on scraped metrics a second or more old.
  • Saturation-driven spills: how often the preferred (cache-holding) replica was too busy and the request went elsewhere.

Scaling signals#

SignalUse it?Why
GPU utilizationNoPinned near 100% under light load (lesson 02)
GPU memory usedNoPre-allocated
Requests per secondWeakRequests are not a unit of work
Requests waiting (queue depth)YesDirect measure of unmet demand
KV-cache usageYesThe real memory pressure
Running requests ÷ configured concurrencyYesHeadroom
TTFT p95 vs SLOAs a guardA lagging symptom; scale before it moves
Tokens/s vs measured capacityYes, for planningNeeds a capacity figure from a load test

In Kubernetes these reach the autoscaler through KEDA or the Prometheus adapter. Two timings decide whether autoscaling works at all, and both must be measured:

time to scale = detection delay + pod scheduling + node provisioning (if no spare GPU)
                + image pull + model download + weight load + warm-up

On a cold node this is commonly several minutes. Record each stage as a metric or a span. If the total exceeds how long your queue can absorb a burst, you need warm spare capacity, not a faster autoscaler.

Capacity that is not serving#

WasteHow it shows up
Pods pending for a GPUkube_pod_status_phase{phase="Pending"} with a GPU request; time-in-pending histogram
GPUs allocated to a pod that is idleGPU power near idle with a pod label (lesson 02)
GPUs not allocated at allAllocatable minus requested, per node pool
Replicas loading a modelReadiness lag after start
FragmentationFree GPUs spread so that no node can fit the next multi-GPU pod
Over-provisioned concurrencyKV usage and running count persistently far below limits

Allocation ratio (GPUs requested ÷ allocatable) and effective use (tokens produced ÷ tokens the allocated GPUs could produce) are the two fleet efficiency numbers. The first is a scheduler metric; the second needs a capacity figure per model and GPU type.

With Dynamic Resource Allocation — core APIs stable since Kubernetes 1.34, device taints stable in 1.37, partitionable devices still maturing — GPUs are requested by attribute through ResourceClaim objects rather than as an opaque count. Your allocation dashboards must follow: count claims and devices, not just nvidia.com/gpu requests.

Tenants#

Per-tenant accounting belongs at the gateway, where identity is known: tokens in, tokens out, cached tokens, requests, errors, throttles, TTFT. Keep cardinality in check (II.05): metrics for the top tenants and tiers, wide events for everyone. Watch for noisy neighbours — one tenant’s long prompts raising everyone’s ITL — by comparing a tenant’s token share with the fleet’s latency.

Rollouts#

A model or engine upgrade is the most common cause of incidents. Put model_revision, engine_version and config_hash on every metric and span, and compare canary with baseline on: TTFT, TPOT, error ratio, finish reasons, output length distribution, prefix hit rate — and quality (lesson 07). A canary can pass every latency check and still produce worse answers.

Multi-node and disaggregated serving#

When one model spans several GPUs or nodes, or prefill and decode run on separate pools, add: KV transfer time and bytes, transfer failures, the balance between the prefill and decode pools (one of them is always the bottleneck), and link health (lesson 02). vLLM exports transfer metrics for its NIXL connector (vllm:nixl_xfer_time_seconds, vllm:nixl_bytes_transferred, failure counters); NVIDIA Dynamo and llm-d expose their own.

Code#

Why queue depth is a better scaling signal than a utilization gauge: simulate both autoscalers on the same traffic.

// autoscale.go — scale on GPU utilization vs on unmet demand, with a slow cold start.
package main

import (
	"fmt"
	"math"
)

func main() {
	const (
		perReplica  = 50.0 // requests a replica can serve per tick within the SLO
		coldStart   = 6    // ticks until a new replica serves traffic
		maxReplicas = 12
		ticks       = 60
	)
	demand := func(t int) float64 {
		if t >= 15 && t < 40 {
			return 260 // a burst
		}
		return 40
	}

	policies := []struct {
		name string
		want func(util, served, waiting float64, replicas int) int
	}{
		{"GPU utilization (up >80%, down <30%)", func(util, _, _ float64, r int) int {
			switch {
			case util > 0.8:
				return r + 1
			case util < 0.3:
				return r - 1
			}
			return r
		}},
		{"served + waiting requests", func(_, served, waiting float64, _ int) int {
			return int(math.Ceil((served + waiting) / perReplica))
		}},
	}

	fmt.Println("policy                                 replica-ticks paid  max replicas  ticks with a queue  peak queue")
	for _, p := range policies {
		ready := 1
		var starting []int // tick at which each new replica becomes ready
		queue, peakQueue := 0.0, 0.0
		lateTicks, paid, most := 0, 0, 1
		for t := 0; t < ticks; t++ {
			for len(starting) > 0 && starting[0] <= t {
				ready++
				starting = starting[1:]
			}
			queue += demand(t)
			served := math.Min(float64(ready)*perReplica, queue)
			queue -= served

			// A GPU serving even one request reports ~100% "utilization".
			util := 0.0
			if served > 0 {
				util = 1.0
			}
			target := p.want(util, served, queue, ready+len(starting))
			target = max(1, min(target, maxReplicas))
			for ready+len(starting) < target {
				starting = append(starting, t+coldStart)
			}
			for ready+len(starting) > target {
				if len(starting) > 0 {
					starting = starting[:len(starting)-1]
				} else {
					ready--
				}
			}
			if queue > 0 {
				lateTicks++
			}
			peakQueue = math.Max(peakQueue, queue)
			paid += ready + len(starting)
			most = max(most, ready+len(starting))
		}
		fmt.Printf("%-38s %18d  %12d  %18d  %10.0f\n", p.name, paid, most, lateTicks, peakQueue)
	}
	fmt.Println("\nUtilization reads 100% whenever anything is served, so that policy climbs to the")
	fmt.Println("maximum and stays there: it is not autoscaling, it is paying for peak all day.")
	fmt.Println("The demand-based policy follows the load; the queue it shows during the burst is")
	fmt.Println("the cold start, and only warm capacity or a faster start removes it.")
}

Remember this#

  • Fleet observability is about differences: replica vs replica, desired vs actual, tenant vs tenant.
  • Scale on waiting requests and KV-cache usage. Measure every stage of a cold start.
  • Routing quality shows up as prefix-cache hit rate.
  • Track capacity that is paid for and not serving: pending, idle, loading, fragmented.
  • Version labels on everything; compare canary and baseline on latency and quality.

Try it#

  1. Run autoscale.go. Set coldStart to 1, then to 15. What changes for each policy, and what does that say about warm capacity?
  2. Write PromQL for: pods pending with a GPU request for more than five minutes; GPUs allocated per node pool; prefix-cache hit rate per replica.
  3. List every stage of a cold start in a system you know and how you would time each.

Check yourself#

  1. Why are GPU utilization and GPU memory poor autoscaling signals?
  2. What does a falling fleet prefix-cache hit rate after a scale-up tell you?
  3. Name four kinds of capacity that is paid for but not serving.

↑↓ navigate ↵ open