The idea in one minute#
A web request costs microseconds of CPU and is either right or an error. An LLM request costs seconds on a device that rents for dollars an hour, its cost depends on how many tokens go in and come out, its latency has two parts (wait for the first token, then the speed of the rest), and its answer can be confidently wrong with a 200 status.
So AI infrastructure needs observability at six layers, three of which ordinary services do not have: the accelerator, the tokens, and the quality of the output.
An analogy#
Monitoring a web service is like monitoring a toll booth: cars per minute, how long each waits, how many barriers are broken. Monitoring LLM serving is like running a restaurant kitchen: every order is a different size, the oven is the expensive constraint, cooking many dishes together is the whole trick, and a dish can leave the kitchen on time and still be bad.
A picture#
flowchart TB Q["6. Quality<br/>is the answer good? evals, feedback, drift"] A["5. Application and agent<br/>traces of model calls, tools, retrieval"] P["4. Platform<br/>routing, autoscaling, tenants, rollout"] E["3. Engine<br/>queue, batch, TTFT, tokens/s, KV cache"] G["2. Accelerator<br/>memory, power, temperature, errors"] F["1. Facility and cost<br/>power, cooling, dollars per token"] Q --- A --- P --- E --- G --- F class Q queue class A,P compute class E memory class G io class F neutral
How it really works#
What breaks from the web-service playbook#
| Web-service assumption | In LLM serving |
|---|---|
| Requests are roughly the same size | Sizes span four orders of magnitude: 20 tokens to 200,000 |
| Latency is one number | TTFT and per-token speed are separate experiences with separate causes |
| Requests per second is the unit of load | Tokens per second is; a request is not a unit of anything |
| CPU utilization shows how busy a server is | “GPU utilization” reads 100% at a fraction of capacity (lesson 02) |
| Memory usage shows pressure | An engine pre-allocates GPU memory; it is ~90% “used” when idle |
| Each request is served independently | Requests are batched together; one long prompt slows its neighbours |
| Scale on CPU | Scale on queue depth and KV-cache pressure (lesson 05) |
| Cold start is seconds | Loading tens of GB of weights takes minutes |
| A 200 means success | The answer may be truncated, off-topic, unsafe or fabricated |
| Cost per request is negligible | Cost per request is the business model |
| Hardware failure is rare | At fleet scale, GPU faults, throttling and link errors are routine |
The six layers and their questions#
| Layer | The question | Typical source | Lesson |
|---|---|---|---|
| 1. Facility and cost | What does this cost in dollars and watts? | Power telemetry, billing, energy counters | 06 |
| 2. Accelerator | Is the hardware healthy and actually working? | NVML / DCGM, vendor equivalents | 02 |
| 3. Engine | Is the server keeping up, and what limits it? | The engine’s /metrics | 03 |
| 4. Platform | Is the fleet routing, scaling and isolating well? | Gateway, scheduler, Kubernetes | 05 |
| 5. Application / agent | What did this request do, step by step? | OTel traces with GenAI conventions | 04 |
| 6. Quality | Was the output good? | Evaluations, user feedback | 07 |
A complete picture needs a join key across them: the model, the replica (pod), the GPU and the tenant must be attributes on every layer’s telemetry. Without that you can see that tokens are slow and that a GPU is hot, and never learn they are the same GPU.
The unit of work is the token#
Three token counts matter, and they cost different amounts:
- Input (prompt) tokens — processed in parallel during prefill; compute-bound.
- Cached input tokens — a prefix already in the KV cache; nearly free.
- Output tokens — generated one at a time during decode; memory-bound, and the bulk of latency.
Every rate, cost and capacity figure should be expressed per token, split by these three. A “requests per second” graph for an LLM service is close to meaningless.
Two workload shapes to keep apart#
| Interactive (chat, agents, voice) | Batch (offline scoring, summarization) | |
|---|---|---|
| Cares about | TTFT, per-token speed, tail latency | Tokens per dollar, completion time |
| SLO | Per-request latency | Job deadline |
| Good state | Headroom, short queues | Full queues, saturated GPUs |
| Same metric, opposite meaning | A long queue is an incident | A long queue is efficiency |
Mixing them on one dashboard produces graphs nobody can interpret. Label by workload class.
Agents change the shape again#
An agent turns one user task into tens of model calls with tool calls between them. Three consequences for observability: the task, not the call, is the unit users judge; contexts grow with every step, so prompt tokens and cache hit rate dominate cost; and the tail of a single call becomes the median of a task (I.04). Lesson 04 traces them.
Code#
Why requests per second misleads: two traffic mixes with the same request rate and very different load.
// units.go — same requests/s, very different load. Tokens are the unit of work.
package main
import "fmt"
type Class struct {
Name string
ReqPerSec float64
PromptTokens float64
CachedShare float64 // fraction of prompt tokens served from the prefix cache
OutputTokens float64
}
func main() {
// Rough per-GPU capacities for one model; replace with measurements (lesson 03).
const prefillTokPerSec = 20000.0 // compute-bound, batched
const decodeTokPerSec = 2500.0 // total across the batch
mixes := map[string][]Class{
"monday: short chat": {
{"chat", 40, 600, 0.5, 250},
},
"tuesday: agents arrive": {
{"chat", 20, 600, 0.5, 250},
{"agent step", 20, 18000, 0.9, 400},
},
}
for _, name := range []string{"monday: short chat", "tuesday: agents arrive"} {
var rps, in, cached, out float64
for _, c := range mixes[name] {
rps += c.ReqPerSec
in += c.ReqPerSec * c.PromptTokens
cached += c.ReqPerSec * c.PromptTokens * c.CachedShare
out += c.ReqPerSec * c.OutputTokens
}
uncached := in - cached
gpus := uncached/prefillTokPerSec + out/decodeTokPerSec
fmt.Printf("%s\n", name)
fmt.Printf(" requests/s %8.0f\n", rps)
fmt.Printf(" prompt tokens/s %8.0f (%.0f%% cached)\n", in, 100*cached/in)
fmt.Printf(" uncached prompt tok/s %8.0f\n", uncached)
fmt.Printf(" output tokens/s %8.0f\n", out)
fmt.Printf(" GPUs needed (rough) %8.1f\n\n", gpus)
}
fmt.Println("The request rate did not change. The load did. Graph tokens, split three ways.")
}
Remember this#
- Six layers: facility/cost, accelerator, engine, platform, application/agent, quality.
- The unit of work is the token, split into input, cached input and output.
- Latency has two parts: time to first token and per-token speed.
- A 200 response says nothing about whether the answer was good.
- Carry model, replica, GPU and tenant on every layer’s telemetry so they can be joined.
Try it#
- Run
units.go. Lower the agents’ cached share from 0.9 to 0.5. What happens to GPUs needed? What does that tell you about which metric to alert on? - For an AI product you use, list one signal you would want at each of the six layers.
- Write an SLO for an agent task rather than a model call. What makes a task “good”?
Check yourself#
- Name three web-service assumptions that fail for LLM serving.
- Why is “requests per second” a poor load metric here?
- Why must interactive and batch traffic be labelled separately?