Goal: a catalogue of the techniques that make inference faster, each analyzed the same way:
Problem → Why it happens → Optimization → How it works
→ Trade-offs → When to use → When NOT to useEvery file follows that structure inside the standard 12-part template. The “when NOT to use” sections are the ones worth rereading.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Optimization methodology | Intermediate | 60 min |
| 02 | Quantization overview and decision guide | Intermediate | 90 min |
| 03 | Weight-only quantization: GPTQ and AWQ | Advanced | 90 min |
| 04 | Activation quantization and SmoothQuant | Advanced | 75 min |
| 05 | FP8 and INT8 inference | Advanced | 75 min |
| 06 | INT4 and low-bit inference | Advanced | 75 min |
| 07 | Kernel and operator fusion | Advanced | 60 min |
| 08 | CUDA graphs in serving | Advanced | 45 min |
| 09 | TensorRT and TensorRT-LLM | Advanced | 75 min |
| 10 | FlashAttention ★ | Advanced | 120 min |
| 11 | KV cache optimization | Advanced | 90 min |
| 12 | KV cache quantization | Advanced | 60 min |
| 13 | Speculative decoding in practice | Advanced | 75 min |
| 14 | Pruning, distillation, architecture | Advanced | 60 min |
The decision tree#
flowchart TB
S["System is slow"] --> M["Measure first - 01<br/>which phase, which regime?"]
M --> R{"What binds you?"}
R -->|"scheduling: low batch, idle slots"| SC["Fix this FIRST<br/>continuous batching - V.09, VIII"]
R -->|"decode: memory bandwidth"| DEC["Move fewer bytes per token"]
R -->|"prefill: compute"| PRE["Do less or cheaper math"]
R -->|"launch / CPU"| CPU["Cut host overhead"]
DEC --> D1["Quantize weights - 02 to 06"]
DEC --> D2["Shrink or quantize KV - 11, 12"]
DEC --> D3["Speculative decoding - 13"]
PRE --> P1["Prefix caching - V.11"]
PRE --> P2["FlashAttention - 10"]
PRE --> P3["FP8 compute - 05"]
PRE --> P4["TensorRT, fusion - 09, 07"]
CPU --> C1["CUDA graphs - 08"]
CPU --> C2["Fusion - 07"]
class S,M neutral
class R,SC queue
class DEC,D1,D2,D3 memory
class PRE,P1,P2,P3,P4 compute
class CPU,C1,C2 ioThe same tree with every option listed:
Is your system slow?
├─ Measure first (01). Which phase? Which regime?
│
├─ DECODE-BOUND (memory bandwidth)
│ ├─ Reduce weight bytes ....... quantization (02-06)
│ ├─ Reduce KV bytes ........... GQA/MLA (XIII), KV quantization (12)
│ ├─ More tokens per read ...... speculative decoding (13)
│ ├─ More sequences per read ... batching (V.09)
│ └─ More aggregate bandwidth .. tensor parallelism (IX)
│
├─ PREFILL-BOUND (compute)
│ ├─ Skip redundant work ....... prefix caching (V.11)
│ ├─ Faster attention .......... FlashAttention (10)
│ ├─ Lower-precision compute ... FP8 (05), not INT4 weight-only
│ └─ Better kernels ............ TensorRT (09), fusion (07)
│
├─ LAUNCH/CPU-BOUND
│ ├─ CUDA graphs ............... (08)
│ ├─ Fusion .................... (07)
│ └─ Move logic off the hot path
│
└─ SCHEDULING-BOUND (low batch, idle slots)
└─ This is Section V.09 and VIII, not this section.
Fix it FIRST — it's a bigger lever than anything here.Order matters. A 20% kernel improvement on a system running at batch 4 when it could run at batch 64 is worth almost nothing. Fix scheduling, then precision, then kernels.