Below the API

What Is Different

Advanced Advanced 45 min Difficulty 3/5

Prerequisites I–IV; helpful: Inference Engineering I.05 and V.03

The idea in one minute#

A web request costs microseconds of CPU and is either right or an error. An LLM request costs seconds on a device that rents for dollars an hour, its cost depends on how many tokens go in and come out, its latency has two parts (wait for the first token, then the speed of the rest), and its answer can be confidently wrong with a 200 status.

So AI infrastructure needs observability at six layers, three of which ordinary services do not have: the accelerator, the tokens, and the quality of the output.

An analogy#

Monitoring a web service is like monitoring a toll booth: cars per minute, how long each waits, how many barriers are broken. Monitoring LLM serving is like running a restaurant kitchen: every order is a different size, the oven is the expensive constraint, cooking many dishes together is the whole trick, and a dish can leave the kitchen on time and still be bad.

A picture#

flowchart TB
  Q["6. Quality<br/>is the answer good? evals, feedback, drift"]
  A["5. Application and agent<br/>traces of model calls, tools, retrieval"]
  P["4. Platform<br/>routing, autoscaling, tenants, rollout"]
  E["3. Engine<br/>queue, batch, TTFT, tokens/s, KV cache"]
  G["2. Accelerator<br/>memory, power, temperature, errors"]
  F["1. Facility and cost<br/>power, cooling, dollars per token"]
  Q --- A --- P --- E --- G --- F
  class Q queue
  class A,P compute
  class E memory
  class G io
  class F neutral

How it really works#

What breaks from the web-service playbook#

Web-service assumptionIn LLM serving
Requests are roughly the same sizeSizes span four orders of magnitude: 20 tokens to 200,000
Latency is one numberTTFT and per-token speed are separate experiences with separate causes
Requests per second is the unit of loadTokens per second is; a request is not a unit of anything
CPU utilization shows how busy a server is“GPU utilization” reads 100% at a fraction of capacity (lesson 02)
Memory usage shows pressureAn engine pre-allocates GPU memory; it is ~90% “used” when idle
Each request is served independentlyRequests are batched together; one long prompt slows its neighbours
Scale on CPUScale on queue depth and KV-cache pressure (lesson 05)
Cold start is secondsLoading tens of GB of weights takes minutes
A 200 means successThe answer may be truncated, off-topic, unsafe or fabricated
Cost per request is negligibleCost per request is the business model
Hardware failure is rareAt fleet scale, GPU faults, throttling and link errors are routine

The six layers and their questions#

LayerThe questionTypical sourceLesson
1. Facility and costWhat does this cost in dollars and watts?Power telemetry, billing, energy counters06
2. AcceleratorIs the hardware healthy and actually working?NVML / DCGM, vendor equivalents02
3. EngineIs the server keeping up, and what limits it?The engine’s /metrics03
4. PlatformIs the fleet routing, scaling and isolating well?Gateway, scheduler, Kubernetes05
5. Application / agentWhat did this request do, step by step?OTel traces with GenAI conventions04
6. QualityWas the output good?Evaluations, user feedback07

A complete picture needs a join key across them: the model, the replica (pod), the GPU and the tenant must be attributes on every layer’s telemetry. Without that you can see that tokens are slow and that a GPU is hot, and never learn they are the same GPU.

The unit of work is the token#

Three token counts matter, and they cost different amounts:

  • Input (prompt) tokens — processed in parallel during prefill; compute-bound.
  • Cached input tokens — a prefix already in the KV cache; nearly free.
  • Output tokens — generated one at a time during decode; memory-bound, and the bulk of latency.

Every rate, cost and capacity figure should be expressed per token, split by these three. A “requests per second” graph for an LLM service is close to meaningless.

Two workload shapes to keep apart#

Interactive (chat, agents, voice)Batch (offline scoring, summarization)
Cares aboutTTFT, per-token speed, tail latencyTokens per dollar, completion time
SLOPer-request latencyJob deadline
Good stateHeadroom, short queuesFull queues, saturated GPUs
Same metric, opposite meaningA long queue is an incidentA long queue is efficiency

Mixing them on one dashboard produces graphs nobody can interpret. Label by workload class.

Agents change the shape again#

An agent turns one user task into tens of model calls with tool calls between them. Three consequences for observability: the task, not the call, is the unit users judge; contexts grow with every step, so prompt tokens and cache hit rate dominate cost; and the tail of a single call becomes the median of a task (I.04). Lesson 04 traces them.

Code#

Why requests per second misleads: two traffic mixes with the same request rate and very different load.

// units.go — same requests/s, very different load. Tokens are the unit of work.
package main

import "fmt"

type Class struct {
	Name         string
	ReqPerSec    float64
	PromptTokens float64
	CachedShare  float64 // fraction of prompt tokens served from the prefix cache
	OutputTokens float64
}

func main() {
	// Rough per-GPU capacities for one model; replace with measurements (lesson 03).
	const prefillTokPerSec = 20000.0 // compute-bound, batched
	const decodeTokPerSec = 2500.0   // total across the batch

	mixes := map[string][]Class{
		"monday: short chat": {
			{"chat", 40, 600, 0.5, 250},
		},
		"tuesday: agents arrive": {
			{"chat", 20, 600, 0.5, 250},
			{"agent step", 20, 18000, 0.9, 400},
		},
	}
	for _, name := range []string{"monday: short chat", "tuesday: agents arrive"} {
		var rps, in, cached, out float64
		for _, c := range mixes[name] {
			rps += c.ReqPerSec
			in += c.ReqPerSec * c.PromptTokens
			cached += c.ReqPerSec * c.PromptTokens * c.CachedShare
			out += c.ReqPerSec * c.OutputTokens
		}
		uncached := in - cached
		gpus := uncached/prefillTokPerSec + out/decodeTokPerSec
		fmt.Printf("%s\n", name)
		fmt.Printf("  requests/s              %8.0f\n", rps)
		fmt.Printf("  prompt tokens/s         %8.0f  (%.0f%% cached)\n", in, 100*cached/in)
		fmt.Printf("  uncached prompt tok/s   %8.0f\n", uncached)
		fmt.Printf("  output tokens/s         %8.0f\n", out)
		fmt.Printf("  GPUs needed (rough)     %8.1f\n\n", gpus)
	}
	fmt.Println("The request rate did not change. The load did. Graph tokens, split three ways.")
}

Remember this#

  • Six layers: facility/cost, accelerator, engine, platform, application/agent, quality.
  • The unit of work is the token, split into input, cached input and output.
  • Latency has two parts: time to first token and per-token speed.
  • A 200 response says nothing about whether the answer was good.
  • Carry model, replica, GPU and tenant on every layer’s telemetry so they can be joined.

Try it#

  1. Run units.go. Lower the agents’ cached share from 0.9 to 0.5. What happens to GPUs needed? What does that tell you about which metric to alert on?
  2. For an AI product you use, list one signal you would want at each of the six layers.
  3. Write an SLO for an agent task rather than a model call. What makes a task “good”?

Check yourself#

  1. Name three web-service assumptions that fail for LLM serving.
  2. Why is “requests per second” a poor load metric here?
  3. Why must interactive and batch traffic be labelled separately?

↑↓ navigate ↵ open