1. What is it?#
Latency is how long one thing takes. Throughput is how many things you finish per unit time. They are different, they are often in conflict, and for LLMs you need at least four numbers, not one.
LLM-specific metrics:
TTFT Time To First Token request arrives → first token emitted
ITL Inter-Token Latency gap between consecutive tokens (a.k.a. TPOT)
E2E End-to-End latency request arrives → last token emitted
TPS Tokens Per Second output tokens/sec — per request, or system-wide
QPS Queries Per Second requests/sec (a weak metric — see §9)
Goodput requests/sec that MET their SLOAnd the relationship you should be able to write from memory:
E2E ≈ TTFT + (output_tokens − 1) × ITLDiagram — The life of one request, and which metric measures which part#
sequenceDiagram
autonumber
participant C as Client
participant Q as Queue
participant E as Engine
C->>Q: request arrives
Note over Q: queue wait
Q->>E: admitted
Note over E: prefill - whole prompt in one pass
E-->>C: first token
Note over C,E: TTFT = queue wait + prefill + first decode step
loop every decode step
E-->>C: next token
end
Note over C,E: ITL = gap between tokens, TPOT is its average
E-->>C: last token + usage
Note over C,E: E2E latency = TTFT + remaining tokens x ITL2. Why does it exist?#
Because “latency” alone is meaningless for a streaming, generative workload.
Consider two systems answering the same 500-token request:
System A: TTFT = 2000 ms, ITL = 10 ms → E2E = 2000 + 499×10 = 6,990 ms
System B: TTFT = 200 ms, ITL = 30 ms → E2E = 200 + 499×30 = 15,170 msSystem B has 2.2x worse end-to-end latency and feels dramatically better to a human, because text appears almost instantly and then streams faster than you can read (30 ms/token ≈ 33 tokens/sec ≈ 25 words/sec, roughly double comfortable reading speed).
A single “latency” SLO cannot express that. You need TTFT and ITL as separate objectives. This is the first genuinely LLM-specific piece of engineering judgment in the curriculum.
3. Simple analogy#
A restaurant.
- TTFT = how long until the bread arrives. Sets the mood for the entire meal.
- ITL = how quickly courses follow one another. Too slow feels neglectful; faster than you can eat is wasted.
- E2E = total time at the table.
- Throughput = covers served per night. What the owner cares about.
- Goodput = covers served without complaints. What the owner should care about.
A restaurant can raise throughput by seating everyone at once and serving in giant batches. The food arrives cold and late. Every LLM serving system faces exactly this decision, every millisecond, in its scheduler.
4. Tiny example#
Concrete traces. Request: 1,000-token prompt, 200-token response.
BATCH SIZE 1
t=0 request arrives
t=0 prefill starts (1000 tokens, one big matmul-heavy pass)
t=50ms prefill done, token 1 emitted ← TTFT = 50 ms
t=60ms token 2 ← ITL = 10 ms
t=70ms token 3
...
t=2040ms token 200 ← E2E = 50 + 199×10 = 2040 ms
Per-request TPS = 200 / 2.04 = 98 tok/s
System TPS = 98 tok/s (only one request!)
BATCH SIZE 32 (all arrive together)
t=0 32 requests arrive
t=0 prefill for all 32 (32,000 tokens — compute bound, takes longer)
t=600ms prefills done, 32 first tokens ← TTFT = 600 ms (12x worse!)
t=615ms 32 second tokens ← ITL = 15 ms (1.5x worse)
...
t=3585ms all done ← E2E = 3585 ms
Per-request TPS = 200 / 3.585 = 56 tok/s (worse per user)
System TPS = 32 × 56 = 1786 tok/s ← 18x better!Read that comparison until it is obvious. Batching made every individual user’s experience worse and made the business 18x more efficient. Every serving decision you make lives on this tradeoff curve. The rest of the curriculum is about bending the curve so you give up less latency for the same throughput — that is precisely what continuous batching, chunked prefill, and disaggregation accomplish.
5. Technical explanation#
Where TTFT comes from#
TTFT = queue_wait + tokenize + prefill_compute + first_sample + network
└─ often └─ ~1ms └─ the real work └─ ~0.1ms └─ 5-50ms
dominant! ∝ prompt_lengthUnder load, queueing usually dominates. A well-tuned system with a bad queue has bad TTFT. This is why admission control and scheduling (Section VIII.03) matter more than kernel micro-optimizations for TTFT.
Prefill compute scales roughly linearly with prompt length (plus a quadratic attention term that matters at long context):
prefill_FLOPs ≈ 2 × P × S_prompt + attention term (∝ L·h·S²·d_head)Where ITL comes from#
ITL = time for one decode step
= max( weight_read_time , compute_time , kv_read_time ) + overheadsAt small batch, weight_read_time dominates:
ITL_floor = model_bytes / HBM_bandwidthFor 70B FP16 on H100: 140 GB / 3.35 TB/s = 42 ms. You cannot beat that at batch 1, no matter
what. To beat it you must either shrink model_bytes (quantization) or amortize it over more
sequences (batching) or read fewer weights per token (MoE, speculative decoding).
This one formula explains most of Section VII. Sit with it.
Percentiles, not averages#
Never report a mean latency. Report p50, p95, p99, and p99.9.
Why: with 20 backend calls per user action, a p99 backend latency
becomes roughly a p80 user experience. (Dean & Barroso, "The Tail at Scale")For LLMs the tail is structural, not noise: a request that generates 4,000 tokens genuinely takes 40x longer than one generating 100. So:
Normalize before percentiling. Report:
- TTFT percentiles (comparable across requests — good)
- ITL percentiles (comparable — good)
- E2E percentiles bucketed by output length (raw E2E percentiles are dominated by length distribution and tell you nothing about system health)
This is a mistake almost every team makes once.
Little’s Law — the one queueing result you must know#
L = λ × W
L = average number of requests in the system (concurrency)
λ = arrival rate (requests/sec)
W = average time in system (seconds)Example: 10 req/s arriving, each taking 4 s → 40 requests in flight on average. If your engine can only hold 20 concurrent sequences in KV cache, you have a problem — half your requests are queueing, and TTFT will climb without bound until arrivals drop.
Rearranged, it is a capacity planning tool:
max_λ = max_concurrency / avg_service_timeYou will use this in Section XI constantly.
Utilization and the latency wall#
Queueing theory’s most important practical fact: as utilization ρ approaches 1, waiting time goes to infinity.
W_queue ≈ W_service × ρ / (1 − ρ)
ρ = 0.5 → wait = 1.0 × service
ρ = 0.8 → wait = 4.0 × service
ρ = 0.9 → wait = 9.0 × service
ρ = 0.95 → wait = 19 × service latency
^
| |
| /
| /
| _____/
| ______________/
|___________/
+------------------------------------------> utilization
0 0.5 0.8 0.9 0.95 1.0Never plan to run GPUs at 95% utilization if you have a latency SLO. Somewhere between 60% and 80% is where cost and latency balance for most interactive workloads. Engineers who come from cost-optimization instincts push for 95% and then cannot understand why p99 exploded.
6. Under the hood#
Instrument the real boundaries, not convenient ones:
// One request's timeline in an OpenAI-compatible server.
type Timeline struct {
Arrive time.Time // at the HTTP handler, before anything
Admitted time.Time // when the scheduler accepts it into a batch
PrefillEnd time.Time // first token produced
Tokens []time.Time // each streamed token
Complete time.Time
}
func (t Timeline) TTFT() time.Duration { return t.PrefillEnd.Sub(t.Arrive) }
func (t Timeline) QueueWait() time.Duration { return t.Admitted.Sub(t.Arrive) }
func (t Timeline) PrefillTime() time.Duration { return t.PrefillEnd.Sub(t.Admitted) }
func (t Timeline) ITL(i int) time.Duration { return t.Tokens[i].Sub(t.Tokens[i-1]) }Emit queue_wait separately from prefill_time. If you only emit TTFT, you cannot tell “we
are overloaded” from “prompts got longer” — and the fixes are completely different (add
capacity vs. optimize prefill / chunk it).
Measure at the client too. Server-side TTFT excludes TLS handshake, network RTT, and any proxy buffering. A proxy that buffers your SSE stream will destroy perceived TTFT while your server dashboards look perfect. This is a real and common production failure.
7. Performance implications#
Sensitivity of each metric to the knobs you control:
| Knob | TTFT | ITL | Throughput | Cost/token |
|---|---|---|---|---|
| ↑ batch size | ↑ worse | ↑ worse | ↑↑ better | ↓↓ better |
| ↑ tensor parallel degree | ↓ better | ↓ better | ↔ or worse | ↑ worse |
| ↓ precision (quantize) | ↓ better | ↓↓ better | ↑ better | ↓ better |
| prefix caching hit | ↓↓ better | ↔ | ↑ better | ↓ better |
| speculative decoding | ↔ | ↓ better | ↔ or ↓ | depends |
| chunked prefill | ↑ slightly | ↓ better | ↑ better | ↓ better |
| ↑ context length used | ↑ worse | ↑ worse | ↓ worse | ↑ worse |
Two entries deserve comment.
Tensor parallelism improves latency but hurts efficiency: splitting a model across 4 GPUs makes each step faster (more aggregate bandwidth) but adds AllReduce communication and leaves each GPU less well utilized. You buy latency with money. (Section IX.)
Chunked prefill trades a little TTFT for a lot of ITL stability: by splitting long prefills into chunks interleaved with decodes, you stop long prompts from stalling everyone else. (Section XIII.05.)
8. Production implications#
Define separate SLOs. A usable spec looks like:
p95 TTFT < 500 ms for prompts ≤ 2,000 tokens
p95 ITL < 50 ms
p99 ITL < 100 ms
Availability > 99.9%
Goodput > 95% of requests meet both TTFT and ITL SLOsNote the prompt-length qualifier on TTFT. Without it, one user pasting a book breaks your SLO and the number becomes meaningless.
Goodput is the metric that actually aligns engineering with product. Raw throughput can be gamed by making everyone slow. Goodput cannot.
Different products want different points on the curve:
| Workload | Priority | Configuration |
|---|---|---|
| Interactive chat | TTFT, ITL | smaller batches, chunked prefill, prefix cache |
| Coding autocomplete | TTFT above all | tiny model, aggressive caching, low batch |
| Batch summarization | throughput only | max batch, no latency SLO, spot instances |
| Agentic / tool loops | E2E per step | short outputs, so TTFT dominates — cache prefixes |
| RAG | TTFT with long prompts | prefill-heavy — chunked prefill, prefix caching |
9. Common mistakes#
Reporting a single “latency” number. Meaningless. Which latency?
Using QPS as the capacity metric. A “query” varies 1000x in cost. Two systems both doing “100 QPS” can differ 50x in real work. Use tokens/sec, and report the input/output length distribution alongside it. Any benchmark quoting QPS without token counts is uninterpretable.
Averaging ITL over the whole response. ITL usually degrades as context grows (more KV to read) and jumps when other requests join the batch. Report the distribution, and ideally ITL as a function of position.
Measuring with a fixed synthetic workload. All-2048-token prompts and all-128-token outputs is not your traffic. Sample real length distributions. Benchmarks built on uniform shapes have misled entire teams into deploying the wrong configuration.
Optimizing E2E when the user sees streaming. If the UI streams, the user’s experience is TTFT + reading speed. Cutting E2E by finishing faster than the user can read buys nothing.
Ignoring queue time. The most common cause of bad TTFT is not slow prefill; it is waiting. Instrument it separately or you will optimize the wrong thing for a month.
Planning for 95% utilization. See the latency wall above.
10. Hands-on exercise#
A. Compute the tradeoffs. A model does prefill at 10,000 tok/s and decode at 40 ms/step (batch 1), 50 ms/step (batch 8), 80 ms/step (batch 32). Requests are 500-token prompts, 150-token outputs.
- TTFT, ITL, E2E, per-request TPS, and system TPS for batch 1, 8, 32.
- If your SLO is p95 TTFT < 400 ms and p95 ITL < 60 ms, which batch sizes are legal?
- Which legal configuration has the lowest cost per token?
B. Little’s Law. Your engine holds at most 48 concurrent sequences. Average request takes 6 s. What is the maximum sustainable arrival rate? At 90% of that rate, using the queueing formula, what is the average wait? What is your practical rate limit?
C. Instrument something real. Run any local model server (vLLM, llama.cpp, Ollama) and
write a client that measures TTFT and per-token ITL for 50 requests with varied prompt lengths.
Plot: (i) TTFT vs prompt length, (ii) ITL vs token position, (iii) the ITL histogram. Write
three sentences explaining each plot. Save the numbers in your numbers.md.
D. Break it. Increase concurrency until TTFT rises 10x. Record the utilization at which the knee occurs. That number is your system’s real capacity.
11. Interview questions#
- Define TTFT, ITL, E2E, and give the formula relating them.
- A user complains the model “feels slow.” What three measurements do you take first?
- Why can improving throughput hurt latency? Give the mechanism, not just the fact.
- What is goodput and why is it a better target than throughput?
- Your p99 E2E latency doubled but p99 TTFT and p99 ITL are unchanged. What happened?
- State Little’s Law and apply it to sizing an LLM serving cluster.
- Why is it dangerous to run an inference cluster at 95% utilization?
12. Further reading#
- [FUNDAMENTAL] Dean & Barroso, “The Tail at Scale,” CACM 2013
- [FUNDAMENTAL] Any queueing theory primer covering M/M/1 and Little’s Law
- [ESTABLISHED] vLLM benchmark scripts (
benchmarks/benchmark_serving.py) — a good model of what to measure - [REFERENCE] MLPerf Inference rules — for how the industry standardizes these metrics
- Next: 06 — Batching: the central tradeoff