1. What is it?#
A repeatable procedure for making an inference system faster, that doesn’t waste months on the wrong thing.
1. DEFINE what "faster" means (which metric, what target)
2. MEASURE where the time goes (phase, then kernel)
3. CLASSIFY the bottleneck (memory / compute / launch / scheduling / queueing)
4. SELECT an optimization that attacks THAT bottleneck
5. ESTIMATE the expected gain before doing the work
6. APPLY one change
7. VERIFY the gain, and the quality
8. REPEATSteps 3 and 5 are the ones people skip, and skipping them is why optimization projects fail.
Diagram — The optimization loop#
flowchart LR M["Measure<br/>baseline + SLO"] --> F["Find the bottleneck<br/>profile, do not guess"] F --> C["Classify<br/>memory / compute / host / queue"] C --> X["Apply ONE change"] X --> V["Verify<br/>speed AND output quality"] V -->|"still short of goal"| F V -->|"goal met"| S["Stop"] class M,F neutral class C queue class X compute class V memory class S io
2. Why does it exist?#
Because inference performance is counterintuitive, and the intuitions people bring from other domains actively mislead:
"We reduced FLOPs by 40%" → zero gain if memory-bound
"We got 100% GPU utilization" → means nothing (Section X.04)
"The kernel is 2x faster" → 5% end-to-end if it was 10% of time
"Quantization made it 4x faster" → for decode; prefill may be unchanged
"More GPUs will help" → only with the right parallelism, and not alwaysEvery one of those is a real statement someone has made in a real project, and every one wasted real time.
3. Simple analogy#
Amdahl’s law, restated as a rule of thumb: if a component is 10% of your time, making it infinitely fast gives you 11% overall. Optimizing anything below 20% of total time is rarely worth it until the big pieces are done.
speedup_overall = 1 / ((1 - p) + p/s)
p = fraction of time in the optimized part
s = speedup of that part
p=0.10, s=∞ → 1.11x
p=0.50, s=2 → 1.33x
p=0.80, s=2 → 1.67x
p=0.80, s=10 → 3.57xAlways compute this before starting work. It takes 30 seconds and frequently kills a bad idea.
4. Tiny example — the estimation habit#
PROPOSAL: "Let's write a fused RMSNorm kernel."
ESTIMATE FIRST:
Profile says: rmsnorm kernels = 8% of decode step time.
Fused version: ~3x faster on that op (6 passes → 2).
Amdahl: 1/((1-0.08) + 0.08/3) = 1/(0.92+0.027) = 1.056x = 5.6% gain.
Effort: 2 days.
Compare: enabling CUDA graphs = 25% gain, 1 hour.
Enabling FP8 = 80% gain, 3 days + validation.
DECISION: do CUDA graphs first, then FP8, then maybe the kernel.That five-minute calculation reorders a quarter’s worth of work.
5. The methodology in detail#
Step 1 — Define the target#
BAD: "make it faster"
GOOD: "reduce p95 ITL from 45 ms to 30 ms at batch 32, without increasing
p95 TTFT above 500 ms or degrading MMLU by more than 0.5 points"Without a target you can’t tell when you’re done, and you can’t trade off correctly. LLM optimization always involves a tradeoff; naming the constraint names the tradeoff.
Step 2 — Measure, at three levels#
LEVEL 1 — Request phases (application metrics)
queue_wait, tokenize, prefill, decode, detokenize, network
→ tells you WHICH PHASE
LEVEL 2 — Timeline (Nsight Systems)
kernel time vs gaps, sync points, transfers
→ tells you GPU vs CPU vs waiting
LEVEL 3 — Kernel (Nsight Compute)
sm%, dram%, occupancy, sectors/request, stall reasons
→ tells you WHY a kernel is slowAlways go 1 → 2 → 3. Starting at level 3 optimizes kernels that don’t matter.
Step 3 — Classify#
Symptom Bottleneck Section
────────────────────────────────────────────────────────────────────
Low batch despite available memory scheduling V.09, VIII.03
High queue_wait capacity XI.02
Gaps between kernels, CPU busy launch/dispatch VI.05, 08
dram% high, sm% low memory bandwidth 02-06, 11, 12
sm% high, dram% low compute 05, 09, 10
Both low latency/occupancy VI.10
Long prefill blocking decode scheduling XIII.05
Client TTFT >> server TTFT network/proxy II.08Step 4-5 — Select and estimate#
Use the decision tree in this section’s README. Then estimate:
expected_gain = (fraction of time in the bottleneck)
× (improvement factor for that bottleneck)
adjusted by AmdahlIf the estimate is under 10%, ask whether there’s a bigger lever available first.
Step 6-7 — Apply and verify#
One change at a time. Two simultaneous changes with a 15% net gain could be +40% and -25%, and you’d never know.
Verify both performance and quality. For any precision change, run the Level 1-4 validation from Section IV.12. A 2x speedup with a 5% quality regression may be a bad trade — but you can only decide that if you measured both.
The optimization ranking (typical, not universal)#
| Rank | Optimization | Typical gain | Effort | Risk |
|---|---|---|---|---|
| 1 | Use an engine with continuous batching | 3-10x | hours | low |
| 2 | Fix scheduling (batch size, admission) | 1.5-4x | days | low |
| 3 | Prefix caching (if prefixes are shared) | 1.3-5x | hours | low |
| 4 | Abort on client disconnect | 1.1-1.4x | hours | low |
| 5 | FP8/INT8 quantization | 1.5-2x | days | medium (quality) |
| 6 | CUDA graphs | 1.1-1.4x | hours | low |
| 7 | Right-size the model | 2-10x | weeks | high (quality) |
| 8 | Chunked prefill | 1.1-1.5x ITL | hours | low |
| 9 | INT4 weight-only (decode-heavy only) | 1.5-2.5x | days | medium |
| 10 | Speculative decoding | 1.3-2.5x | weeks | medium |
| 11 | Tensor parallelism (for latency) | 1.5-3x latency | days | low |
| 12 | Custom kernels | 1.05-1.3x | weeks | medium |
Items 1-6 are config and hours of work. Most teams jump to 10-12. Do them in order.
6. Under the hood: the estimation formulas#
Keep these to hand:
Decode floor: T = bytes_read / bandwidth
Prefill floor: T = flops / peak_flops
Quantization gain: bytes_ratio (for memory-bound), flops_ratio (for compute-bound)
Batching gain: min(target_batch/current_batch, ridge_point/current_intensity)
Fusion gain: (passes_before / passes_after) on that op's time
Graph gain: launch_overhead_total / step_time
Amdahl: 1/((1-p) + p/s)With these you can estimate any optimization’s payoff in under five minutes.
7-9. Performance, production, mistakes#
Performance: the methodology itself has a performance benefit — it prevents you from spending three weeks on a 4% gain when a config flag gives 40%.
Production:
- Keep a performance journal: baseline, change, measured gain, quality delta.
- Automate the benchmark so anyone can run it.
- Track achieved-vs-roofline as a regression signal.
- Re-measure after every dependency upgrade; library heuristics change.
Mistakes:
- Optimizing without measuring. The cardinal sin.
- Multiple simultaneous changes. You learn nothing.
- Not measuring quality. Half of inference optimizations trade accuracy.
- Benchmarking with unrealistic workloads. Fixed lengths, no concurrency.
- Ignoring Amdahl. Optimizing 5% of the time.
- Optimizing the median when the SLO is on the tail.
- Not re-measuring after the change. “It should be faster” is not data.
10. Hands-on exercise#
A. Build the benchmark. Write a reproducible benchmark for your system: realistic length distribution, realistic arrival pattern, reporting TTFT/ITL percentiles, throughput, and cost per million tokens. Everything else in this section depends on having this.
B. Full diagnosis. Take a running system and produce a complete diagnosis: phase breakdown, timeline analysis, top-5 kernel analysis, and a classification. Write it up in one page.
C. Estimate before doing. Pick three optimizations. For each, estimate the gain using the formulas in section 6. Then implement one and compare the actual gain to your estimate. How close were you? Adjust your estimation model.
D. Amdahl in practice. For your system, list every component with its share of time. For each, compute the maximum possible end-to-end gain from making it infinitely fast. Which are worth working on?
E. The journal. Start a performance-journal.md: date, baseline numbers, change made,
measured result, quality delta, decision. Keep it for the rest of this section.
11. Interview questions#
- Walk me through your methodology for optimizing an inference system.
- A colleague reduced FLOPs by 40% with no speedup. What do you tell them?
- What is Amdahl’s law and how do you apply it before starting work?
- Rank the top five optimizations for a typical LLM service and justify the order.
- Why do you change one thing at a time?
- How would you decide between quantization and speculative decoding?
- What would you measure to know whether an optimization is safe to ship?
12. Further reading#
- [FUNDAMENTAL] Amdahl, “Validity of the single processor approach” (1967)
- [FUNDAMENTAL] Brendan Gregg, the USE method
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” — analysis before implementation, done well
- Next: 02 — Quantization overview