The idea in one minute#
This lesson assembles the course into one design. The method is the same for any system: write down the questions people will ask, decide which signal answers each, define a common schema so the signals join, lay out the pipeline, set SLOs and alerts, and budget the cost. Tools come last.
The worked example is a multi-tenant LLM serving platform: a gateway, a router, a fleet of engine replicas on GPUs in Kubernetes, serving several models to interactive and batch tenants.
A picture#
flowchart TB
subgraph DATA["Data plane"]
GW["Gateway"] --> RT["Router"] --> EN["Engine replicas"] --> GP[("GPUs")]
end
GW -->|"wide event per request,<br/>spans, tenant token metrics"| AG["Node agents<br/>OTel Collector"]
RT -->|"decision spans and metrics"| AG
EN -->|"/metrics scrape, phase spans"| AG
GP -->|"dcgm-exporter, Xid logs"| AG
K8S["Kubernetes state"] --> AG
AG --> GWC["Collector gateways<br/>enrich, redact, tail-sample, per-tenant limits"]
GWC --> TS[("Metrics TSDB<br/>alerts, SLOs, long-term")]
GWC --> EV[("Column store<br/>wide events, traces")]
GWC --> CT[("Restricted store<br/>opt-in prompt content")]
TS --> UI["Dashboards, SLO alerts"]
EV --> UI
EV --> EVAL["Evaluators"] --> EV
class GW,RT,EN compute
class GP memory
class AG,GWC io
class TS,EV,CT memory
class UI,EVAL queue
class K8S neutralHow it really works#
Step 1 — the questions#
| Who | Asks |
|---|---|
| On-call | Are users affected? Which model, tenant, replica? Since which change? |
| Capacity planner | How close is each model to its limit? What will next month need? |
| Platform engineer | Is routing working? Why is this replica slow? How long is a cold start? |
| Finance / product | What does each tenant and model cost? What is the margin per token? |
| Model owner | Did the new revision or quantization change quality? |
| Tenant | What did I use, how fast was it, was I throttled? |
| Security | Who accessed which prompts? Was anything leaked across tenants? |
Every signal in the design must trace back to a row here. Anything that does not is cost.
Step 2 — one schema#
Agree on the attribute names before the first line of instrumentation, and attach them everywhere: metrics labels (where bounded), span and event attributes, log fields, GPU metric labels.
identity tenant.id, tenant.tier, workload.class (interactive | batch)
model gen_ai.request.model, model.revision, model.precision, lora.adapter
serving engine.name, engine.version, config.hash, replica (k8s.pod.name)
placement cloud.region, k8s.cluster.name, k8s.node.name, node.pool,
accelerator.type, gpu.uuid
request trace_id, request.id, gen_ai.response.id, stream, finish_reason, error.type
size input_tokens, cached_input_tokens, output_tokens
timing gateway_ms, queue_ms, prefill_ms, ttft_ms, decode_ms, tpot_ms, total_msUse OTel semantic-convention names where they exist; keep your own names in one shared package so a rename upstream is a one-line change.
Step 3 — signals per component#
| Component | Metrics | Events / traces | Notes |
|---|---|---|---|
| Gateway | RED by route, model, tier; tokens by tenant (top-N) | One wide event per request with the whole schema; root span | The accounting source of truth |
| Router | Decision latency; spills; prefix-match rate | A span with the chosen replica and scores | |
| Engine | The five families of V.03 | Spans for queue / prefill / decode, linked to the caller’s trace | Scrape every 15 s; these also drive routing and autoscaling |
| GPU | V.02 set, with the off-by-default fields enabled | Xid events as structured logs | Labelled with pod and gpu.uuid |
| Kubernetes | kube-state-metrics, node, DRA claims | Scheduling and scaling events | |
| Model lifecycle | Download, load, warm-up durations | One span per cold start, with stages | |
| Evaluators | Score distributions by revision | An evaluation event keyed by response ID |
Step 4 — the pipeline#
- Agents on every node: scrape engine and GPU endpoints, receive OTLP, tail kernel and container logs, add Kubernetes metadata.
- Collector gateways: redact, enforce per-tenant limits, tail-sample traces (keep all errors, SLO violations and a baseline), derive RED metrics from spans before sampling.
- Stores: a Prometheus-compatible TSDB for metrics and SLO rules; a column store for wide events and traces; a separate, access-controlled store with short retention for opt-in prompt content.
- Isolation: the telemetry stack runs outside the serving clusters’ failure domain, with a dead man’s switch.
Step 5 — SLOs and alerts#
| SLO (per model, interactive class) | Target |
|---|---|
| Availability: requests completing without server error | 99.9% |
| TTFT under 800 ms | 99% |
| TPOT under 50 ms | 99% |
| Batch class: jobs finished by deadline | 99% |
Pages — a short list, all symptoms: fast and slow burn on each SLO; the telemetry pipeline’s dead man’s switch; a hardware-fault rate above the fleet’s norm. Everything else — KV pressure, cache hit rate, pending pods, idle GPUs, cost — is a ticket or a dashboard.
Step 6 — dashboards#
- Fleet overview: SLO status and budget per model; goodput; tokens/s; spend today.
- Model: the engine row (V.03) with the replica spread; routing and cache; canary vs baseline.
- Replica / node: engine metrics beside the GPU metrics of that pod.
- Capacity and cost: occupancy, allocation, $ per M tokens, tokens per joule, idle GPU-hours.
- Tenant: usage, latency, throttles — also what the tenant sees.
- Quality: finish reasons, parse failures, evaluation scores by revision.
Step 7 — cost of the design#
Estimate before building. For a fleet of 400 GPUs at 20,000 requests per second across models:
metrics 400 GPUs × ~60 GPU series + 250 replicas × ~400 engine series
+ Kubernetes and gateway ≈ 0.3 M active series
events 20,000/s × 1.2 kB × 86,400 ≈ 2 TB/day raw, ~170 GB/day compressed
traces tail-sampled to ~3% plus all errors ≈ 200 GB/day raw
content opt-in, sampled, 7-day retention sized per tenant agreementThe program below prices this at roughly 2% of what the fleet’s GPUs cost per day — about what two points of occupancy are worth, and observability of this kind typically recovers more than that (V.06). If the estimate had come out at 10%, the design would need more sampling before it needed more hardware.
Step 8 — privacy and access#
Prompts and completions are customer data. Default off; opt-in per tenant; redact in the Collector; separate store, short retention, audited access. Tenant-facing usage views must be filtered server-side by tenant, never by a dashboard variable.
Step 9 — rollout order#
A design is not adopted in one step. A realistic order:
- Engine
/metrics, dcgm-exporter and Kubernetes metrics into Prometheus; the model dashboard. - Gateway wide events and tenant token accounting.
- SLOs and burn-rate alerts; delete threshold alerts.
- Trace propagation gateway → router → engine; exemplars.
- Cost and efficiency views; showback.
- Quality signals and canary gates.
- Hardware-health automation; continuous profiling of gateway and engine hosts.
Each step is useful on its own, which is what keeps the project alive.
Review checklist#
- Every signal maps to a question in step 1.
- One schema; model, replica, GPU and tenant join across all layers.
- SLOs on TTFT and TPOT, not end-to-end latency; pages only on symptoms.
- Autoscaling and routing read queue depth and KV usage, not GPU utilization.
- Cardinality ceilings computed; tenants bounded in metrics, unbounded in events.
- Metrics derived before sampling; sampling rates recorded.
- Prompt content off by default, with a written policy.
- The pipeline is monitored and outside the failure domain it watches.
- A version label on everything; canaries compared on quality as well as latency.
- Telemetry cost estimated and attributed.
Code#
A sizing calculator for the design, so the estimate in step 7 is reproducible.
// sizing.go — estimate the telemetry volume of an inference platform design.
package main
import "fmt"
func main() {
const (
gpus = 400.0
replicas = 250.0
reqPerSec = 20000.0
gpuSeries = 60.0 // per GPU, after enabling health and profiling fields
engineSeries = 400.0 // per replica, mostly histogram buckets
platformSeries = 150000.0
eventBytes = 1200.0
compression = 12.0
spansPerRequest = 9.0
spanBytes = 450.0
traceKeep = 0.03
gpuHourPrice = 3.20
)
series := gpus*gpuSeries + replicas*engineSeries + platformSeries
eventsRaw := reqPerSec * eventBytes * 86400 / 1e9
tracesRaw := reqPerSec * spansPerRequest * spanBytes * 86400 * traceKeep / 1e9
fleetPerDay := gpus * gpuHourPrice * 24
fmt.Printf("active metric series %10.0f\n", series)
fmt.Printf("wide events, raw %10.0f GB/day\n", eventsRaw)
fmt.Printf("wide events, compressed %10.0f GB/day\n", eventsRaw/compression)
fmt.Printf("traces after tail sampling %10.0f GB/day raw\n", tracesRaw)
fmt.Printf("GPU fleet cost %10.0f $/day\n\n", fleetPerDay)
// Illustrative unit prices for self-hosted storage and compute, all-in.
telemetry := series/1e6*250 + (eventsRaw/compression)*30*0.03 + tracesRaw/compression*14*0.03 + 400
fmt.Printf("telemetry, order of magnitude %8.0f $/day (%.2f%% of the fleet)\n", telemetry, 100*telemetry/fleetPerDay)
fmt.Println("One point of occupancy on this fleet is worth about", int(fleetPerDay/100), "$/day.")
}
Remember this#
- Design from questions to signals to schema to pipeline to SLOs to cost; pick tools last.
- One shared schema is what makes six layers one system.
- Page on TTFT, TPOT, availability and the pipeline’s own health. Everything else informs.
- Build it in steps that each pay for themselves.
Try it#
- Run
sizing.gofor a platform a tenth the size. Which term dominates now? - Write step 1’s table for a system you operate. Which questions have no signal today?
- Do the review checklist against an existing design and list the three most valuable fixes.
Check yourself#
- Why do tools come last in the design order?
- Which component should own per-tenant accounting, and why?
- What belongs on a pager for an inference platform, and what does not?
You have finished the course. Build the projects if you have not, then apply the design method above to a real system. The companion paths — Inference Engineering and GPU Engineering — cover the systems this course taught you to watch.