★ If you internalize one thing from Section X, make it this procedure.
1. The procedure#
STEP 0 — REPRODUCE
Can you make it slow on demand? If not, you're debugging a ghost.
Capture: the workload shape, the concurrency, the configuration.
STEP 1 — DEFINE THE METRIC
Not "slow." Which number, at which percentile, versus what target?
p95 TTFT is 3.2 s, target 800 ms.
p99 ITL is 180 ms, target 50 ms.
Throughput is 1,200 tok/s, expected 3,000.
Cost is $1.80/M tokens, budgeted $0.60.
Different metrics have completely different causes.
STEP 2 — DECOMPOSE BY PHASE
Application metrics: queue_wait, tokenize, prefill, decode, detokenize,
network. Which phase holds the time?
→ This single step resolves ~50% of investigations.
STEP 3 — GPU, CPU, OR WAITING?
Nsight Systems timeline, or:
sum(kernel_time) / wall_clock
Near 1.0 → GPU-bound; go to step 4.
Well below → CPU-bound or waiting; go to file 09.
STEP 4 — WHICH REGIME?
Nsight Compute on the top kernels:
sm__throughput% and dram__throughput%
high/low → compute bound
low/high → memory bound
low/low → latency bound
→ determines WHICH optimizations can possibly help (file 03)
STEP 5 — WHICH KERNEL, AND WHY?
Per-kernel: occupancy, sectors/request, stall reasons.
STEP 6 — ESTIMATE THE FIX
Amdahl: what's the maximum possible gain? Is it worth the effort?
STEP 7 — CHANGE ONE THING, RE-MEASURE, RECORD.Diagram — Why is it slow? A triage tree#
flowchart TB
S["Service is slow"] --> W{"Where does the time go?<br/>queue wait vs engine time"}
W -->|"queue wait"| CAP["Capacity or admission problem<br/>batch size, replicas, routing"]
W -->|"engine time"| H{"Is the GPU busy<br/>for the whole step?"}
H -->|"no: gaps between kernels"| HOST["Host-bound<br/>Python, tokenizer, launch overhead, syncs"]
H -->|"yes"| R{"Roofline: is bandwidth<br/>or FLOPs saturated?"}
R -->|"bandwidth"| MEM["Memory-bound<br/>batch more, quantize, shrink KV"]
R -->|"FLOPs"| CMP["Compute-bound<br/>precision, kernels, more GPUs"]
R -->|"neither"| KRN["Inefficient kernel<br/>occupancy, coalescing, fusion"]
class W,H,R,CAP queue
class HOST io
class MEM memory
class CMP compute
class KRN warn
class S neutral2. Why the order matters#
Because each step eliminates a class of hypotheses cheaply, and the expensive steps come last.
Step 2 costs 5 minutes and eliminates 5 of 6 phases.
Step 5 costs 2 hours and tells you about one kernel.
Doing step 5 first means you might spend 2 hours characterizing a kernel
in a phase that accounts for 3% of your latency.The most common failure mode in performance work is starting too deep. Engineers enjoy kernel analysis; it feels like real work. The phase decomposition feels trivial. But the phase decomposition is where the answer usually is.
3. Simple analogy#
A doctor’s diagnostic sequence. Vital signs, then history, then targeted tests, then imaging, then biopsy. Nobody starts with a biopsy — it’s invasive, slow, and you wouldn’t know where to take the sample.
Performance debugging has the same structure and the same failure mode when done backwards.
4. The phase decomposition (step 2 in detail)#
This is the highest-value step, so instrument it properly:
# In the API layer / engine, emit per-request:
{
"request_id": ...,
"arrival_ts": ...,
"tokenize_ms": 1.2,
"queue_wait_ms": 2340.0, # ← usually the answer
"prefill_ms": 87.3,
"first_token_ms": 2428.5, # = tokenize + queue + prefill + sample
"decode_ms": 4210.0,
"itl_p50_ms": 21.0,
"itl_p99_ms": 145.0, # ← the tail within one request
"detokenize_ms": 12.4,
"total_ms": 6650.9,
"prompt_tokens": 1843,
"output_tokens": 200,
"batch_size_at_admission": 47,
"kv_blocks_used": 128,
"prefix_cache_hit_tokens": 1536,
"model_version": "...",
"tenant": "..."
}
Emit this for every request (sampled if volume is high). With it, most investigations become a database query:
-- Where is TTFT going?
SELECT
percentile_cont(0.95) WITHIN GROUP (ORDER BY queue_wait_ms) AS p95_queue,
percentile_cont(0.95) WITHIN GROUP (ORDER BY prefill_ms) AS p95_prefill,
percentile_cont(0.95) WITHIN GROUP (ORDER BY tokenize_ms) AS p95_tokenize
FROM requests WHERE ts > now() - interval '1 hour';
If p95_queue is 2,300 ms of a 2,500 ms TTFT, you have a capacity problem and no amount of
kernel work will help.
5. The decision tree from the phase result#
QUEUE_WAIT dominates
→ capacity or admission problem
→ check: average running batch size vs the maximum
low → traffic-limited or scheduler-limited (Section VIII.03/04)
high → genuinely at capacity (Section XI.02)
→ also check: are long prefills blocking? (chunked prefill, Section XIII.05)
PREFILL dominates
→ check the prompt length distribution first (did it change?)
→ check prefix cache hit rate (Section V.11)
→ then: is prefill compute-bound and near the roofline? (file 02)
→ then: kernel analysis (file 08)
DECODE (ITL) dominates
→ compute the memory-bandwidth floor: bytes_per_step / bandwidth
→ measured / floor:
> 0.8 → you're near the physical limit; reduce bytes (quantize)
0.4-0.8 → kernel or launch overhead; profile
< 0.4 → something is structurally wrong
TOKENIZE / DETOKENIZE dominates
→ slow Python tokenizer, or per-token detokenization bug (Section V.02)
NETWORK dominates (client TTFT >> server TTFT)
→ buffering proxy (Section II.08)
NOTHING dominates (all phases small, total large)
→ you're measuring wrong, or the time is in gaps (step 3)6. Step 3 in detail — GPU, CPU, or waiting#
# The quick version
import torch
from torch.profiler import profile, ProfilerActivity
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
run_one_decode_step()
table = prof.key_averages()
cpu_total = sum(e.self_cpu_time_total for e in table)
gpu_total = sum(e.self_device_time_total for e in table)
print(f"CPU {cpu_total/1000:.1f} ms, GPU {gpu_total/1000:.1f} ms")
gpu_total ≈ wall_clock → GPU-bound. Go to step 4.
cpu_total ≈ wall_clock → CPU-bound. Go to file 09.
both << wall_clock → WAITING. On what?
- a lock (py-spy dump shows it)
- an implicit sync (.item())
- the network
- a collective (NCCL)The “both much less than wall clock” case is the interesting one and the one people miss. It means the process is blocked, and off-CPU profiling (file 09) is what finds it.
7. The estimation habit (step 6)#
Before doing the work, compute the ceiling:
FOUND: the RMSNorm kernels are 9% of decode step time and are unfused.
FIX: fuse them (3x faster on that operation).
GAIN: 1/((1-0.09) + 0.09/3) = 1.064 → 6.4%
EFFORT: 2 days.
ALTERNATIVE FOUND: average running batch is 8, memory allows 60.
FIX: investigate why (traffic? max_num_seqs? preemption?).
GAIN: potentially 3-5x.
EFFORT: hours.
→ Do the second one first.Always compute the Amdahl ceiling before starting. It takes 30 seconds and reorders your priorities regularly.
8. The performance journal#
Keep a running record. Format:
DATE 2026-03-14
SYMPTOM p95 TTFT 3.2 s, target 800 ms
WORKLOAD 1,800 avg prompt tokens, 250 output, 45 req/s
BASELINE TTFT p50 340 ms, p95 3,200 ms; queue_wait p95 2,700 ms
DIAGNOSIS Queue-bound. Running batch 58/175 — not memory limited.
Chunked prefill was disabled; long prompts (p99 = 22k tokens)
were blocking decode for 400+ ms each.
CHANGE --enable-chunked-prefill, --max-num-batched-tokens 4096
RESULT TTFT p95 780 ms; ITL p99 improved 145→52 ms; throughput +8%
QUALITY unchanged (verified: 200 paired generations, identical)
EFFORT 30 minutesThis journal becomes your team’s institutional memory and is the artifact that makes you credible in capacity discussions.
9. Common mistakes#
Starting with kernel analysis. The most common and most expensive.
Not defining the metric. “Faster” is not a target.
Changing multiple things. You learn nothing.
Not computing the Amdahl ceiling. Weeks on a 4% gain.
Optimizing p50 when the SLO is on p99. Different causes.
Not measuring quality alongside performance. Half of inference optimizations trade accuracy.
Benchmarking with an unrealistic workload (file 07).
Not writing it down. You will investigate the same thing again in six months.
10. Hands-on exercise#
A. Build the instrumentation. Add the per-request phase timing from section 4 to a real server. Emit it as structured logs or metrics. Write the query that decomposes p95 TTFT.
B. Run the full procedure. Take a system that’s slower than you’d like and execute all seven steps. Write it up in the journal format from section 8.
C. Practice the ceiling calculation. For five hypothetical findings (each a component with a known time share and a known possible speedup), compute the Amdahl ceiling. Rank them by gain-per-effort.
D. The three-way split. For a decode step, measure cpu_total, gpu_total, and
wall_clock. Compute the gaps. What are you waiting on?
E. Start the journal. Record your current baseline for a system you work on: workload, configuration, all key metrics. You need a baseline to detect regressions.
11. Interview questions#
- Walk me through your methodology for a slow inference system.
- Why start with phase decomposition rather than kernel profiling?
- What per-request telemetry would you emit, and what would each field diagnose?
gpu_timeandcpu_timeare both far below wall clock. What’s happening?- How do you decide whether an optimization is worth doing before doing it?
- Your p95 TTFT is 3 s. Walk me through the first ten minutes of investigation.
- Why change one thing at a time?
12. Further reading#
- [FUNDAMENTAL] Brendan Gregg, Systems Performance — the USE method and the general methodology
- [FUNDAMENTAL] Amdahl’s law
- [ESTABLISHED] Google SRE Book, chapters on debugging and postmortems
- Next: 02 — The roofline model in practice