Eight worked investigations. Read each symptom, stop, and write down your hypotheses before reading the diagnosis. That exercise is the point of this file.
Case 1 — “TTFT is 4 seconds”#
SYMPTOM p95 TTFT 4,100 ms (target 800 ms). p95 ITL 24 ms (fine).
Throughput 3,900 tok/s. No errors.
STOP. What are your hypotheses?
MEASUREMENTS
queue_wait p95: 3,650 ms ← 89% of TTFT
prefill p95: 380 ms
tokenize p95: 3 ms
running batch avg: 62 (max_num_seqs = 64)
KV usage: 94%
arrival rate: 18 req/s
Little's Law check: L = λW = 18 × 6.2 s = 112 in system, capacity 64
→ 48 always waiting
DIAGNOSIS
Genuinely at capacity. Not a tuning problem — an arithmetic one.
The system can serve ~10 req/s at the current request shape; 18 are arriving.
WRONG FIXES
✗ optimize kernels (they're fine)
✗ increase max_num_seqs (KV is at 94%; you'd cause preemption)
✗ increase TP degree (doesn't help throughput; Section IX.12)
RIGHT FIXES, in order
1. Check prefix caching. Chat workload, 1,100-token system prompt,
hit rate was 0% (round-robin routing).
→ enable prefix-aware routing: prefill drops 68%, capacity +40%
2. Enable FP8: capacity +80%
3. Add replicas for the remainder.
RESULT after (1) and (2): 18 req/s served at p95 TTFT 640 ms,
with no new hardware.Lesson: “at capacity” is a real diagnosis, but check whether you’re at capacity efficiently before buying hardware.
Case 2 — “ITL spikes to 300 ms randomly”#
SYMPTOM p50 ITL 22 ms, p99 ITL 310 ms. Spikes are irregular.
TTFT and throughput fine.
STOP. Hypotheses?
MEASUREMENTS
Plotted ITL over time: spikes are ~450 ms and occur 3-8 times/minute.
Correlated with: arrival of requests with prompt > 12,000 tokens.
chunked prefill: DISABLED
max_num_batched_tokens: 32,768
DIAGNOSIS
Head-of-line blocking. A 12k-token prefill runs as one unit, occupying
the GPU for ~450 ms, during which every decoding sequence stalls.
FIX
--enable-chunked-prefill --max-num-batched-tokens 2048
→ the prefill is split into 6 chunks, each riding along with a decode step
→ each step is ~28 ms instead of 22 ms, but there are no 450 ms stalls
RESULT p50 ITL 26 ms (slightly worse), p99 ITL 51 ms (6x better).
p95 TTFT for long prompts: 480 → 620 ms (slightly worse).
Overall throughput: +6%.Lesson: a p99 that’s 14x the p50 is almost always a blocking event, not general slowness. Correlate the spikes with something.
Case 3 — “Throughput is a third of what we calculated”#
SYMPTOM Predicted 3,000 tok/s from the roofline; measured 980 tok/s.
Batch 32, 8B model, one H100.
STOP. Hypotheses?
MEASUREMENTS
nsys: GPU busy 38% of wall clock
API summary: 4,200 cudaLaunchKernel per step, 1 cudaStreamSynchronize
py-spy: 41% of time in a custom logits processor (Python)
which calls .cpu() on the logits every step
CUDA graphs: disabled (--enforce-eager was in the config)
DIAGNOSIS
Two compounding problems:
1. --enforce-eager left in the config from a debugging session
2. a per-step D2H copy + sync in a custom stopping-criterion callback
FIX
1. Remove --enforce-eager → 980 → 1,510 tok/s
2. Move the stopping check to the GPU → 1,510 → 2,740 tok/s
RESULT 2,740 tok/s (91% of the roofline prediction).
Total effort: 3 hours.Lesson: a 3x gap from the roofline is almost never kernel efficiency. It’s CPU, launches, or synchronization. Check GPU-busy percentage first.
Case 4 — “OOM three times a day”#
SYMPTOM Server OOMs 3-4 times daily, always near peak. Restarts cleanly.
STOP. Hypotheses?
MEASUREMENTS
memory_summary() at failure: reserved 78.1 GB, allocated 71.4 GB
→ 6.7 GB fragmentation, but that's not the
whole story
Correlation: every OOM within 2 s of a request with prompt > 28,000 tokens
gpu_memory_utilization: 0.94
chunked prefill: disabled
Reproduced: 30k prompt + 40 decoding sequences → OOM
DIAGNOSIS
Type 2 (Section X.06): activation spike during unchunked long prefill.
The KV pool was sized using a profiling pass with max_num_batched_tokens
tokens, but without chunking the actual prefill processed all 30k at once,
producing an activation tensor 5x larger than profiled.
FIX
1. --enable-chunked-prefill (bounds activation to max_num_batched_tokens)
2. --gpu-memory-utilization 0.90 (headroom)
3. Cap max input at 32,000 at the gateway with a clear 400
RESULT Zero OOMs in 30 days. Bonus: p99 ITL improved 40% (same root cause
as Case 2).Lesson: OOMs correlate with something. Find the correlation before tuning memory settings.
Case 5 — “Adding GPUs made it slower”#
SYMPTOM Went from TP=4 to TP=8 to increase throughput. Throughput per GPU
dropped 35%; total throughput up only 30%.
STOP. Hypotheses?
MEASUREMENTS
nvidia-smi topo -m:
GPU0-3: NV18 to each other
GPU4-7: NV18 to each other
GPU0-4: SYS ← across the CPU interconnect
nsys with NCCL trace: AllReduce time 4.2 ms of a 9.1 ms step (46%)
At TP=4 (within one NVLink island): AllReduce 0.6 ms of 12.8 ms (5%)
DIAGNOSIS
The node has two NVLink islands, not one NVSwitch fabric.
TP=8 forces AllReduce across the CPU interconnect.
FIX
Revert to TP=4, run TWO instances (DP=2 × TP=4), one per NVLink island.
Pin each with numactl to its NUMA node.
RESULT Total throughput +95% vs the original TP=4 single instance,
vs +30% for TP=8. Same 8 GPUs.Lesson: always check nvidia-smi topo -m before choosing a TP degree. This exact mistake is
common on non-DGX hardware.
Case 6 — “The GPUs are at 100%, we need more”#
SYMPTOM Capacity request for 64 additional GPUs. Justification:
"all 64 existing GPUs at 100% utilization."
STOP. Hypotheses?
MEASUREMENTS
DCGM_FI_PROF_DRAM_ACTIVE: 0.24
DCGM_FI_PROF_SM_ACTIVE: 0.97
avg_running_batch_size: 7 (max_num_seqs = 96)
kv_cache_usage_ratio: 0.09
queue depth: 0
replica count: 16, each receiving ~4 req/s
DIAGNOSIS
Not capacity-limited at all. Traffic is spread too thinly across too many
replicas; each reads the full model per step to serve 7 sequences
(Section IX.12, way 5).
FIX
1. Consolidate 16 → 5 replicas
2. Prefix-aware routing across them
3. Keep 1 spare for headroom
RESULT avg batch 7 → 26. DRAM_ACTIVE 0.24 → 0.71. Throughput per GPU 3.4x.
44 of 64 GPUs freed. No purchase.Lesson: nvidia-smi utilization is not a capacity signal (Section X.04). Require
DRAM_ACTIVE and avg_running_batch_size in every capacity request.
Case 7 — “Quality dropped after a deploy, but all tests passed”#
SYMPTOM User complaints about "worse answers" starting Tuesday.
Error rate, latency, throughput all normal. All CI tests green.
STOP. Hypotheses?
MEASUREMENTS
Deploy log: Tuesday's release changed the serving precision from BF16 to
INT4 (W4A16, GPTQ, group 128) for cost reasons.
Offline eval at deploy time: MMLU -0.4%, perplexity +0.08. Accepted.
Output length distribution: mean 218 → 341 tokens (+56%) ← the signal
Regeneration rate: 4.1% → 11.3%
Structured output validity: 99.2% → 91.7%
DIAGNOSIS
INT4 quantization degraded instruction-following and long-form coherence
in a way that MMLU and perplexity did not capture (Section VII.06).
The model rambles and violates JSON schemas more often.
FIX
1. Immediate: roll back to BF16 (old checkpoint still cached — good)
2. Re-evaluate with the Level 3 tests (Section IV.12): long generations,
instruction following, structured output validity
3. Deploy FP8 instead: -0.1% MMLU, output length unchanged,
structured validity 99.1%, and 1.9x throughput.
RESULT Cost target met with FP8. Added output-length distribution and
structured-validity to the canary guards.Lesson: output length distribution is a cheap, sensitive early quality signal. Perplexity and MMLU are not sufficient validation for a precision change.
Case 8 — “Client sees 8 seconds, server says 200 ms”#
SYMPTOM Users report the response appears all at once after ~8 seconds.
Server-side TTFT metric: p95 210 ms. Server-side E2E: p95 7.9 s.
STOP. Hypotheses?
MEASUREMENTS
curl -N -w '%{time_starttransfer}' through the production ingress: 7.82 s
curl -N directly to the pod IP: 0.19 s ← the answer
nginx ingress config: proxy_buffering not set → defaults to ON
DIAGNOSIS
The ingress buffers the entire SSE response before forwarding.
Server metrics are correct; the user experience is not.
FIX
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 3600s;
add_header X-Accel-Buffering no;
(and verify the CDN and WAF layers too)
RESULT Client TTFT p95: 7.82 s → 0.26 s.
Added a synthetic client-side TTFT check through the production
ingress to monitoring.Lesson: always measure from a real client through the real path. Server-side metrics cannot see this class of problem (Section II.08).
The patterns across all eight#
1. DECOMPOSE FIRST. Six of eight were resolved by phase decomposition
or a single correlation, before any deep profiling.
2. CHECK THE CONFIGURATION. Cases 3, 4, 8 were configuration errors.
Cases 2 and 5 were configuration choices made without measurement.
3. MEASURE FROM THE OUTSIDE. Case 8 is invisible from the inside.
4. THE OBVIOUS METRIC LIES. Case 6: GPU utilization. Case 7: MMLU.
5. CORRELATE ANOMALIES WITH SOMETHING. Cases 2 and 4 were both
"correlate the spikes with the request shape."
6. THE FIX IS OFTEN A FLAG. Cases 2, 3, 4, 8. Total engineering time
for those four: about a day. Combined impact: large.
7. AMDAHL. In every case, the right fix addressed the dominant term.
In Case 3, someone could have spent a month on kernels for 5%.Hands-on exercise#
A. Predict before reading. Re-read each case’s symptom and measurements, covering the diagnosis. Write your hypothesis. How often were you right? Which cases fooled you and why?
B. Reproduce one. Pick Case 2, 3, or 4 and reproduce it deliberately on a test system. Confirm the symptom, apply the fix, measure the improvement.
C. Write your own. Take a real performance problem you’ve encountered (in any system) and write it up in this format: symptom, measurements, diagnosis, wrong fixes, right fix, result, lesson. This is the format for your team’s incident postmortems.
D. Build the checklist. From the eight cases, derive a “first 15 minutes” checklist for an inference performance incident. What do you measure, in what order?
E. The capacity request template. Write the template for a GPU capacity request at your organization, requiring the evidence that Cases 1 and 6 show is necessary.
Interview questions#
- p95 TTFT is 4 s, p95 ITL is fine. Walk me through your investigation.
- p99 ITL is 14x p50, irregularly. What’s your first hypothesis?
- Measured throughput is a third of the roofline prediction. What do you check first?
- A team requests 64 more GPUs citing 100% utilization. What evidence do you require?
- Quality dropped after a deploy but all tests passed. How do you find it, and how do you prevent it next time?
- Client-side latency is 40x server-side latency. Diagnose.
- Adding GPUs reduced per-GPU throughput 35%. What happened?
Further reading#
- [FUNDAMENTAL] Google SRE Book, “Effective Troubleshooting” and the postmortem chapters
- [FUNDAMENTAL] Brendan Gregg, Systems Performance — the case studies chapter
- Next: Section XI — Production Inference Engineering