Below the API

Inference Optimization

IntermediateModule14 topics~17h 30m
completed

Topics, in order

About this module

Goal: a catalogue of the techniques that make inference faster, each analyzed the same way:

Problem  →  Why it happens  →  Optimization  →  How it works
         →  Trade-offs  →  When to use  →  When NOT to use

Every file follows that structure inside the standard 12-part template. The “when NOT to use” sections are the ones worth rereading.

Files#

#FileLevelTime
01Optimization methodologyIntermediate60 min
02Quantization overview and decision guideIntermediate90 min
03Weight-only quantization: GPTQ and AWQAdvanced90 min
04Activation quantization and SmoothQuantAdvanced75 min
05FP8 and INT8 inferenceAdvanced75 min
06INT4 and low-bit inferenceAdvanced75 min
07Kernel and operator fusionAdvanced60 min
08CUDA graphs in servingAdvanced45 min
09TensorRT and TensorRT-LLMAdvanced75 min
10FlashAttention ★Advanced120 min
11KV cache optimizationAdvanced90 min
12KV cache quantizationAdvanced60 min
13Speculative decoding in practiceAdvanced75 min
14Pruning, distillation, architectureAdvanced60 min

The decision tree#

flowchart TB
  S["System is slow"] --> M["Measure first - 01<br/>which phase, which regime?"]
  M --> R{"What binds you?"}
  R -->|"scheduling: low batch, idle slots"| SC["Fix this FIRST<br/>continuous batching - V.09, VIII"]
  R -->|"decode: memory bandwidth"| DEC["Move fewer bytes per token"]
  R -->|"prefill: compute"| PRE["Do less or cheaper math"]
  R -->|"launch / CPU"| CPU["Cut host overhead"]
  DEC --> D1["Quantize weights - 02 to 06"]
  DEC --> D2["Shrink or quantize KV - 11, 12"]
  DEC --> D3["Speculative decoding - 13"]
  PRE --> P1["Prefix caching - V.11"]
  PRE --> P2["FlashAttention - 10"]
  PRE --> P3["FP8 compute - 05"]
  PRE --> P4["TensorRT, fusion - 09, 07"]
  CPU --> C1["CUDA graphs - 08"]
  CPU --> C2["Fusion - 07"]

  class S,M neutral
  class R,SC queue
  class DEC,D1,D2,D3 memory
  class PRE,P1,P2,P3,P4 compute
  class CPU,C1,C2 io

The same tree with every option listed:

Is your system slow?
├─ Measure first (01). Which phase? Which regime?
│
├─ DECODE-BOUND (memory bandwidth)
│   ├─ Reduce weight bytes ....... quantization (02-06)
│   ├─ Reduce KV bytes ........... GQA/MLA (XIII), KV quantization (12)
│   ├─ More tokens per read ...... speculative decoding (13)
│   ├─ More sequences per read ... batching (V.09)
│   └─ More aggregate bandwidth .. tensor parallelism (IX)
│
├─ PREFILL-BOUND (compute)
│   ├─ Skip redundant work ....... prefix caching (V.11)
│   ├─ Faster attention .......... FlashAttention (10)
│   ├─ Lower-precision compute ... FP8 (05), not INT4 weight-only
│   └─ Better kernels ............ TensorRT (09), fusion (07)
│
├─ LAUNCH/CPU-BOUND
│   ├─ CUDA graphs ............... (08)
│   ├─ Fusion .................... (07)
│   └─ Move logic off the hot path
│
└─ SCHEDULING-BOUND (low batch, idle slots)
    └─ This is Section V.09 and VIII, not this section.
       Fix it FIRST — it's a bigger lever than anything here.

Order matters. A 20% kernel improvement on a system running at batch 4 when it could run at batch 64 is worth almost nothing. Fix scheduling, then precision, then kernels.

↑↓ navigate ↵ open