Below the API

Advanced LLM Inference

ExpertModule12 topics~15h 15m
completed

Topics, in order

About this module

Goal: the techniques that define the current frontier of production inference — and an honest assessment of which are ready.

Everything here is labelled:

  • [ESTABLISHED] — deployed at scale by multiple organizations
  • [EMERGING] — promising, deployed by some, still maturing
  • [RESEARCH] — interesting, not production-ready

Read the labels. Several techniques in this section are frequently presented as ready when they aren’t.

Files#

#FileLevelStatusTime
01MQA, GQA, MLAAdvancedESTABLISHED75 min
02Mixture of ExpertsAdvancedESTABLISHED90 min
03MoE serving and expert parallelismAdvancedESTABLISHED75 min
04Long-context inferenceAdvancedmixed90 min
05Chunked prefillAdvancedESTABLISHED60 min
06Disaggregated prefill/decodeAdvancedEMERGING90 min
07KV offloading and transferAdvancedEMERGING75 min
08Request scheduling and prioritiesAdvancedmixed60 min
09Speculative decoding variantsAdvancedmixed75 min
10Low-bit inference: FP4 and beyondAdvancedEMERGING60 min
11Triton kernelsAdvancedESTABLISHED90 min
12Compilers: torch.compile and beyondAdvancedESTABLISHED75 min

The thread#

flowchart TD
  N0["Shrink the KV cache architecturally<br/><b>01</b>"]
  N1["Or shrink the ACTIVE parameters per token<br/><b>02, 03</b>"]
  N2["Which matters most at long context<br/><b>04</b>"]
  N3["Where prefill must be chunked to be usable<br/><b>05</b>"]
  N4["Or separated entirely onto different hardware<br/><b>06</b>"]
  N5["With KV moving between tiers<br/><b>07</b>"]
  N6["Under a scheduler that knows about priorities<br/><b>08</b>"]
  N7["Producing multiple tokens per pass where possible<br/><b>09</b>"]
  N8["In as few bits as the hardware supports<br/><b>10</b>"]
  N9["Using kernels you can now write yourself<br/><b>11</b>"]
  N10["Or that a compiler writes for you<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10

  class N0,N1,N2 neutral
  class N3,N4 io
  class N5,N6 queue
  class N7,N8 compute
  class N9,N10 memory

How to read this section#

If you’re operating a system: files 01, 02, 05, and 08 affect decisions you make today. Files 04 and 10 affect model selection. The rest are for when you hit their specific bottleneck.

If you’re building: files 11 and 12 are the practical skills. 06 and 07 are the architectures to understand before you need them.

↑↓ navigate ↵ open