Below the API

Designing Observability for an Inference Platform

Expert Expert 1h Difficulty 4/5

Prerequisites the whole course

The idea in one minute#

This lesson assembles the course into one design. The method is the same for any system: write down the questions people will ask, decide which signal answers each, define a common schema so the signals join, lay out the pipeline, set SLOs and alerts, and budget the cost. Tools come last.

The worked example is a multi-tenant LLM serving platform: a gateway, a router, a fleet of engine replicas on GPUs in Kubernetes, serving several models to interactive and batch tenants.

A picture#

flowchart TB
  subgraph DATA["Data plane"]
    GW["Gateway"] --> RT["Router"] --> EN["Engine replicas"] --> GP[("GPUs")]
  end
  GW -->|"wide event per request,<br/>spans, tenant token metrics"| AG["Node agents<br/>OTel Collector"]
  RT -->|"decision spans and metrics"| AG
  EN -->|"/metrics scrape, phase spans"| AG
  GP -->|"dcgm-exporter, Xid logs"| AG
  K8S["Kubernetes state"] --> AG
  AG --> GWC["Collector gateways<br/>enrich, redact, tail-sample, per-tenant limits"]
  GWC --> TS[("Metrics TSDB<br/>alerts, SLOs, long-term")]
  GWC --> EV[("Column store<br/>wide events, traces")]
  GWC --> CT[("Restricted store<br/>opt-in prompt content")]
  TS --> UI["Dashboards, SLO alerts"]
  EV --> UI
  EV --> EVAL["Evaluators"] --> EV
  class GW,RT,EN compute
  class GP memory
  class AG,GWC io
  class TS,EV,CT memory
  class UI,EVAL queue
  class K8S neutral

How it really works#

Step 1 — the questions#

WhoAsks
On-callAre users affected? Which model, tenant, replica? Since which change?
Capacity plannerHow close is each model to its limit? What will next month need?
Platform engineerIs routing working? Why is this replica slow? How long is a cold start?
Finance / productWhat does each tenant and model cost? What is the margin per token?
Model ownerDid the new revision or quantization change quality?
TenantWhat did I use, how fast was it, was I throttled?
SecurityWho accessed which prompts? Was anything leaked across tenants?

Every signal in the design must trace back to a row here. Anything that does not is cost.

Step 2 — one schema#

Agree on the attribute names before the first line of instrumentation, and attach them everywhere: metrics labels (where bounded), span and event attributes, log fields, GPU metric labels.

identity     tenant.id, tenant.tier, workload.class (interactive | batch)
model        gen_ai.request.model, model.revision, model.precision, lora.adapter
serving      engine.name, engine.version, config.hash, replica (k8s.pod.name)
placement    cloud.region, k8s.cluster.name, k8s.node.name, node.pool,
             accelerator.type, gpu.uuid
request      trace_id, request.id, gen_ai.response.id, stream, finish_reason, error.type
size         input_tokens, cached_input_tokens, output_tokens
timing       gateway_ms, queue_ms, prefill_ms, ttft_ms, decode_ms, tpot_ms, total_ms

Use OTel semantic-convention names where they exist; keep your own names in one shared package so a rename upstream is a one-line change.

Step 3 — signals per component#

ComponentMetricsEvents / tracesNotes
GatewayRED by route, model, tier; tokens by tenant (top-N)One wide event per request with the whole schema; root spanThe accounting source of truth
RouterDecision latency; spills; prefix-match rateA span with the chosen replica and scores
EngineThe five families of V.03Spans for queue / prefill / decode, linked to the caller’s traceScrape every 15 s; these also drive routing and autoscaling
GPUV.02 set, with the off-by-default fields enabledXid events as structured logsLabelled with pod and gpu.uuid
Kuberneteskube-state-metrics, node, DRA claimsScheduling and scaling events
Model lifecycleDownload, load, warm-up durationsOne span per cold start, with stages
EvaluatorsScore distributions by revisionAn evaluation event keyed by response ID

Step 4 — the pipeline#

  • Agents on every node: scrape engine and GPU endpoints, receive OTLP, tail kernel and container logs, add Kubernetes metadata.
  • Collector gateways: redact, enforce per-tenant limits, tail-sample traces (keep all errors, SLO violations and a baseline), derive RED metrics from spans before sampling.
  • Stores: a Prometheus-compatible TSDB for metrics and SLO rules; a column store for wide events and traces; a separate, access-controlled store with short retention for opt-in prompt content.
  • Isolation: the telemetry stack runs outside the serving clusters’ failure domain, with a dead man’s switch.

Step 5 — SLOs and alerts#

SLO (per model, interactive class)Target
Availability: requests completing without server error99.9%
TTFT under 800 ms99%
TPOT under 50 ms99%
Batch class: jobs finished by deadline99%

Pages — a short list, all symptoms: fast and slow burn on each SLO; the telemetry pipeline’s dead man’s switch; a hardware-fault rate above the fleet’s norm. Everything else — KV pressure, cache hit rate, pending pods, idle GPUs, cost — is a ticket or a dashboard.

Step 6 — dashboards#

  1. Fleet overview: SLO status and budget per model; goodput; tokens/s; spend today.
  2. Model: the engine row (V.03) with the replica spread; routing and cache; canary vs baseline.
  3. Replica / node: engine metrics beside the GPU metrics of that pod.
  4. Capacity and cost: occupancy, allocation, $ per M tokens, tokens per joule, idle GPU-hours.
  5. Tenant: usage, latency, throttles — also what the tenant sees.
  6. Quality: finish reasons, parse failures, evaluation scores by revision.

Step 7 — cost of the design#

Estimate before building. For a fleet of 400 GPUs at 20,000 requests per second across models:

metrics   400 GPUs × ~60 GPU series + 250 replicas × ~400 engine series
          + Kubernetes and gateway                      ≈ 0.3 M active series
events    20,000/s × 1.2 kB × 86,400                     ≈ 2 TB/day raw, ~170 GB/day compressed
traces    tail-sampled to ~3% plus all errors            ≈ 200 GB/day raw
content   opt-in, sampled, 7-day retention               sized per tenant agreement

The program below prices this at roughly 2% of what the fleet’s GPUs cost per day — about what two points of occupancy are worth, and observability of this kind typically recovers more than that (V.06). If the estimate had come out at 10%, the design would need more sampling before it needed more hardware.

Step 8 — privacy and access#

Prompts and completions are customer data. Default off; opt-in per tenant; redact in the Collector; separate store, short retention, audited access. Tenant-facing usage views must be filtered server-side by tenant, never by a dashboard variable.

Step 9 — rollout order#

A design is not adopted in one step. A realistic order:

  1. Engine /metrics, dcgm-exporter and Kubernetes metrics into Prometheus; the model dashboard.
  2. Gateway wide events and tenant token accounting.
  3. SLOs and burn-rate alerts; delete threshold alerts.
  4. Trace propagation gateway → router → engine; exemplars.
  5. Cost and efficiency views; showback.
  6. Quality signals and canary gates.
  7. Hardware-health automation; continuous profiling of gateway and engine hosts.

Each step is useful on its own, which is what keeps the project alive.

Review checklist#

  • Every signal maps to a question in step 1.
  • One schema; model, replica, GPU and tenant join across all layers.
  • SLOs on TTFT and TPOT, not end-to-end latency; pages only on symptoms.
  • Autoscaling and routing read queue depth and KV usage, not GPU utilization.
  • Cardinality ceilings computed; tenants bounded in metrics, unbounded in events.
  • Metrics derived before sampling; sampling rates recorded.
  • Prompt content off by default, with a written policy.
  • The pipeline is monitored and outside the failure domain it watches.
  • A version label on everything; canaries compared on quality as well as latency.
  • Telemetry cost estimated and attributed.

Code#

A sizing calculator for the design, so the estimate in step 7 is reproducible.

// sizing.go — estimate the telemetry volume of an inference platform design.
package main

import "fmt"

func main() {
	const (
		gpus            = 400.0
		replicas        = 250.0
		reqPerSec       = 20000.0
		gpuSeries       = 60.0  // per GPU, after enabling health and profiling fields
		engineSeries    = 400.0 // per replica, mostly histogram buckets
		platformSeries  = 150000.0
		eventBytes      = 1200.0
		compression     = 12.0
		spansPerRequest = 9.0
		spanBytes       = 450.0
		traceKeep       = 0.03
		gpuHourPrice    = 3.20
	)
	series := gpus*gpuSeries + replicas*engineSeries + platformSeries
	eventsRaw := reqPerSec * eventBytes * 86400 / 1e9
	tracesRaw := reqPerSec * spansPerRequest * spanBytes * 86400 * traceKeep / 1e9
	fleetPerDay := gpus * gpuHourPrice * 24

	fmt.Printf("active metric series        %10.0f\n", series)
	fmt.Printf("wide events, raw            %10.0f GB/day\n", eventsRaw)
	fmt.Printf("wide events, compressed     %10.0f GB/day\n", eventsRaw/compression)
	fmt.Printf("traces after tail sampling  %10.0f GB/day raw\n", tracesRaw)
	fmt.Printf("GPU fleet cost              %10.0f $/day\n\n", fleetPerDay)

	// Illustrative unit prices for self-hosted storage and compute, all-in.
	telemetry := series/1e6*250 + (eventsRaw/compression)*30*0.03 + tracesRaw/compression*14*0.03 + 400
	fmt.Printf("telemetry, order of magnitude %8.0f $/day (%.2f%% of the fleet)\n", telemetry, 100*telemetry/fleetPerDay)
	fmt.Println("One point of occupancy on this fleet is worth about", int(fleetPerDay/100), "$/day.")
}

Remember this#

  • Design from questions to signals to schema to pipeline to SLOs to cost; pick tools last.
  • One shared schema is what makes six layers one system.
  • Page on TTFT, TPOT, availability and the pipeline’s own health. Everything else informs.
  • Build it in steps that each pay for themselves.

Try it#

  1. Run sizing.go for a platform a tenth the size. Which term dominates now?
  2. Write step 1’s table for a system you operate. Which questions have no signal today?
  3. Do the review checklist against an existing design and list the three most valuable fixes.

Check yourself#

  1. Why do tools come last in the design order?
  2. Which component should own per-tenant accounting, and why?
  3. What belongs on a pager for an inference platform, and what does not?

You have finished the course. Build the projects if you have not, then apply the design method above to a real system. The companion paths — Inference Engineering and GPU Engineering — cover the systems this course taught you to watch.

↑↓ navigate ↵ open