The idea in one minute#
An LLM server is a queue in front of a batch loop. A request waits, is prefilled (the prompt is read, producing the first token), then decodes one token per step alongside everyone else in the batch. Each phase has its own latency metric and its own cause of slowness, so the engine must report them separately.
Five families of metric tell you nearly everything: latency by phase (queue, TTFT, inter-token), throughput in tokens, saturation (waiting requests, KV-cache usage, preemptions), cache effectiveness (prefix cache hits), and errors and finish reasons.
An analogy#
A lift in a tall building. You wait for it (queue). It takes a moment to load everyone (prefill). Then it stops at every floor for every passenger (decode steps shared by the batch). “How long did my trip take” is the sum, but to fix a slow building you need to know which of the three is the problem — and how full the lift was.
A picture#
flowchart TB
ARR["Request arrives"] --> W["WAITING<br/>queue time"]
W --> PF["PREFILL<br/>read the prompt, cached prefix skipped"]
PF --> FT["First token sent<br/>TTFT = queue + prefill"]
FT --> DEC["DECODE<br/>one token per step, shared with the batch<br/>inter-token latency"]
DEC --> END["Finished<br/>stop, length, error"]
DEC -->|"KV cache full"| PRE["Preempted<br/>back to waiting"]
PRE --> W
KV[("KV cache blocks")] --- DEC
class ARR,END,FT neutral
class W queue
class PF,DEC compute
class KV memory
class PRE warnHow it really works#
The latency metrics#
| Metric | Definition | Dominated by | Users feel it as |
|---|---|---|---|
| Queue time | Arrival → scheduled | Load vs capacity | Part of the initial wait |
| Prefill time | Scheduled → first token computed | Uncached prompt length; compute-bound | Part of the initial wait |
| TTFT (time to first token) | Arrival → first token sent | Queue + prefill (+ network) | “Is it responding?” |
| ITL (inter-token latency) | Gap between consecutive tokens | Batch size, model size, memory bandwidth | Reading speed, stutter |
| TPOT (time per output token) | (end − first token) ÷ (output tokens − 1) | Same as ITL, averaged per request | Overall streaming speed |
| End-to-end | Arrival → last token | TTFT + TPOT × output tokens | Total wait for non-streaming use |
end-to-end ≈ TTFT + TPOT × (output_tokens − 1). Because output length varies so much,
end-to-end latency is mostly a measure of how much was asked for. Set SLOs on TTFT and TPOT,
and track end-to-end only per workload class.
ITL and TPOT are not the same thing: TPOT averages a request’s gaps, hiding a stall; ITL’s distribution shows it. A stream that pauses for two seconds mid-answer has a fine TPOT and a terrible ITL p99.
Throughput#
- Output tokens/s and prompt tokens/s, per replica and per model — the load figures.
- Tokens per engine step — how full the batch actually is.
- Requests running — the current batch size.
Saturation#
| Signal | Meaning | Why it matters |
|---|---|---|
| Requests waiting | Queue depth | The earliest, clearest overload signal; the best autoscaling input |
| KV-cache usage | Fraction of KV blocks allocated | The true memory pressure (GPU “memory used” is not) |
| Preemptions | A running request evicted to free KV blocks | The engine is thrashing: lots of wasted work, sharp latency tail |
| Waiting by reason | Why requests cannot start (capacity, KV space, …) | Tells you whether to add replicas or change limits |
Cache effectiveness#
Prefix cache hit rate = cached prompt tokens ÷ prompt tokens. For chat and agents with long repeated prefixes this is the single largest cost lever: a hit turns thousands of prefill tokens into a lookup. A drop in hit rate after a deploy (someone put a timestamp at the top of the system prompt) can double GPU cost with no other symptom than higher TTFT.
vLLM’s metrics, by name#
vLLM serves Prometheus metrics at /metrics, prefixed vllm:. Names below are from the
metrics reference for v0.30 (September 2026). Names have changed between releases — for
example KV-cache usage was once exported as vllm:gpu_cache_usage_perc — so check your
version’s page.
| Family | Metric | Type |
|---|---|---|
| Latency | vllm:time_to_first_token_seconds | Histogram |
vllm:inter_token_latency_seconds | Histogram | |
vllm:request_time_per_output_token_seconds | Histogram | |
vllm:e2e_request_latency_seconds | Histogram | |
vllm:request_queue_time_seconds, vllm:request_prefill_time_seconds, vllm:request_decode_time_seconds, vllm:request_inference_time_seconds | Histograms | |
| Throughput | vllm:prompt_tokens, vllm:generation_tokens | Counters |
vllm:iteration_tokens_total | Histogram (tokens per engine step) | |
vllm:request_prompt_tokens, vllm:request_generation_tokens | Histograms (size per request) | |
| Saturation | vllm:num_requests_running, vllm:num_requests_waiting | Gauges |
vllm:num_requests_waiting_by_reason | Gauge | |
vllm:kv_cache_usage_perc | Gauge, 0–1 | |
vllm:num_preemptions, vllm:request_num_preemptions | Counter, histogram | |
| Cache | vllm:prefix_cache_queries, vllm:prefix_cache_hits | Counters (tokens) |
vllm:prompt_tokens_cached, vllm:prompt_tokens_by_source | Counters | |
vllm:kv_block_lifetime_seconds, vllm:kv_block_idle_before_evict_seconds | Histograms | |
| Outcome | vllm:request_success (labelled by finish reason) | Counter |
vllm:corrupted_requests | Counter (NaNs in the logits) | |
| Efficiency | vllm:estimated_flops_per_gpu_total, vllm:estimated_read_bytes_per_gpu_total | Counters |
| Speculative decoding | vllm:spec_decode_num_accepted_tokens_per_pos | Counter |
Other engines expose the same ideas under other names: SGLang (sglang: prefix), NVIDIA’s
Dynamo and Triton, TensorRT-LLM, llama.cpp’s server, and hosted APIs through their usage
fields. Hugging Face’s TGI — archived in March 2026 — used tgi_ names you will still meet in
older dashboards. Learn the families, then map the names.
The queries to keep#
# TTFT p95 per model
histogram_quantile(0.95, sum by (le, model_name) (rate(vllm:time_to_first_token_seconds_bucket[5m])))
# Output tokens per second, per replica
sum by (pod) (rate(vllm:generation_tokens_total[1m]))
# Prefix cache hit rate
sum(rate(vllm:prefix_cache_hits_total[5m])) / sum(rate(vllm:prefix_cache_queries_total[5m]))
# Queue pressure per replica, and preemptions per minute
max by (pod) (vllm:num_requests_waiting)
sum by (pod) (rate(vllm:num_preemptions_total[5m])) * 60
# Where does request time go? Mean share spent queued
sum(rate(vllm:request_queue_time_seconds_sum[5m])) / sum(rate(vllm:e2e_request_latency_seconds_sum[5m]))
Prometheus appends _total to counters and _bucket / _sum / _count to histograms on
exposition; confirm the exact series names by reading your own /metrics.
Reading the combinations#
| Symptom | With | Likely cause | Action |
|---|---|---|---|
| TTFT up | Waiting up, KV usage moderate | Not enough compute for arrivals | Scale out |
| TTFT up | Waiting low, prefill time up | Longer or less-cached prompts | Check prefix hit rate, prompt changes |
| ITL up | Running count up | Bigger batch: the designed trade-off | Lower the concurrency limit, or scale |
| ITL p99 spikes | A long-prompt request arrived | Prefill is stalling decode steps | Chunked prefill; separate prefill and decode |
| Preemptions above zero | KV usage near 1.0 | KV cache exhausted | Fewer concurrent sequences, shorter contexts, more memory |
| Throughput down | Clock or power down on the GPU | Thermal or power throttling | Lesson 02 |
| Tokens/s fine, users unhappy | Finish reason length rising | Truncated outputs | A quality problem (lesson 07) |
Goodput#
Throughput counts every token. Goodput counts only requests that met the SLO: TTFT under its threshold and TPOT under its threshold and completed. As load rises, throughput keeps climbing while goodput peaks and collapses — the peak is your real capacity. Compute it per request (a gateway or client-side event is the natural place) and make it the number you size, autoscale and benchmark against.
Measure from outside as well#
Server-side TTFT stops at the engine’s socket. Users also pay for the gateway, the network and client buffering. Record TTFT and ITL in the client or with a synthetic prober, and compare: the difference is everything in front of the engine.
Code#
Compute the streaming metrics and goodput from token timestamps, as a load tester or gateway would.
// streammetrics.go — TTFT, TPOT, ITL p99 and goodput from per-token timestamps.
package main
import (
"fmt"
"math/rand"
"sort"
)
type Request struct {
Arrival float64 // seconds
Tokens []float64 // time each output token was received
}
func simulate(rng *rand.Rand, n int, load float64) []Request {
reqs := make([]Request, n)
for i := range reqs {
queue := rng.ExpFloat64() * 0.05 * load * load // queueing grows sharply with load
prefill := 0.08 + rng.Float64()*0.25
itl := 0.018 * (1 + 0.6*load) // bigger batches → slower steps
t := queue + prefill
toks := make([]float64, 40+rng.Intn(400))
for j := range toks {
toks[j] = t
gap := itl * (0.9 + 0.2*rng.Float64())
if rng.Float64() < 0.002*load { // another request's long prefill stalls the batch
gap += 0.8
}
t += gap
}
reqs[i] = Request{0, toks}
}
return reqs
}
func p(xs []float64, q float64) float64 {
s := append([]float64{}, xs...)
sort.Float64s(s)
return s[int(q*float64(len(s)-1))]
}
func main() {
const ttftSLO, tpotSLO = 0.5, 0.040
rng := rand.New(rand.NewSource(4))
fmt.Println("load TTFT p95 TPOT p95 ITL p99 tokens/s/req good requests")
for _, load := range []float64{0.5, 1, 2, 3, 4} {
reqs := simulate(rng, 2000, load)
var ttft, tpot, itl []float64
good := 0
for _, r := range reqs {
first, last := r.Tokens[0], r.Tokens[len(r.Tokens)-1]
tt := first - r.Arrival
tp := (last - first) / float64(len(r.Tokens)-1)
ttft, tpot = append(ttft, tt), append(tpot, tp)
for i := 1; i < len(r.Tokens); i++ {
itl = append(itl, r.Tokens[i]-r.Tokens[i-1])
}
if tt <= ttftSLO && tp <= tpotSLO {
good++
}
}
fmt.Printf("%4.1f %6.0f ms %5.1f ms %6.0f ms %10.1f %10.1f%%\n", load,
p(ttft, .95)*1000, p(tpot, .95)*1000, p(itl, .99)*1000, 1/p(tpot, .5), 100*float64(good)/float64(len(reqs)))
}
fmt.Printf("\nSLO: TTFT <= %.0f ms and TPOT <= %.0f ms. Capacity is the load where 'good' is still ~99%%.\n",
ttftSLO*1000, tpotSLO*1000)
}
Remember this#
- Three phases — wait, prefill, decode — each with its own metric and its own cause.
- SLOs on TTFT and TPOT; watch the ITL distribution for stalls.
- Saturation is requests waiting, KV-cache usage and preemptions — not GPU memory or utilization.
- Prefix cache hit rate is a cost metric. Goodput, not throughput, is capacity.
- Metric names differ by engine and version; the families do not.
Try it#
- Run
streammetrics.go. At which load does goodput fall below 99%? Which SLO breaks first? - Start any engine with a small model and read
/metrics. Map every family in this lesson to a series name. No GPU: do the same against the vLLM metrics reference page. - Build a six-panel dashboard from the queries above, in the order: traffic, errors, TTFT, ITL, saturation, cache.
Check yourself#
- Which two phases make up TTFT, and what drives each?
- How can TPOT look healthy while users see the stream freeze?
- What do preemptions tell you, and what is the usual fix?