Goal: the techniques that define the current frontier of production inference — and an honest assessment of which are ready.
Everything here is labelled:
- [ESTABLISHED] — deployed at scale by multiple organizations
- [EMERGING] — promising, deployed by some, still maturing
- [RESEARCH] — interesting, not production-ready
Read the labels. Several techniques in this section are frequently presented as ready when they aren’t.
Files#
| # | File | Level | Status | Time |
|---|---|---|---|---|
| 01 | MQA, GQA, MLA | Advanced | ESTABLISHED | 75 min |
| 02 | Mixture of Experts | Advanced | ESTABLISHED | 90 min |
| 03 | MoE serving and expert parallelism | Advanced | ESTABLISHED | 75 min |
| 04 | Long-context inference | Advanced | mixed | 90 min |
| 05 | Chunked prefill | Advanced | ESTABLISHED | 60 min |
| 06 | Disaggregated prefill/decode | Advanced | EMERGING | 90 min |
| 07 | KV offloading and transfer | Advanced | EMERGING | 75 min |
| 08 | Request scheduling and priorities | Advanced | mixed | 60 min |
| 09 | Speculative decoding variants | Advanced | mixed | 75 min |
| 10 | Low-bit inference: FP4 and beyond | Advanced | EMERGING | 60 min |
| 11 | Triton kernels | Advanced | ESTABLISHED | 90 min |
| 12 | Compilers: torch.compile and beyond | Advanced | ESTABLISHED | 75 min |
The thread#
flowchart TD N0["Shrink the KV cache architecturally<br/><b>01</b>"] N1["Or shrink the ACTIVE parameters per token<br/><b>02, 03</b>"] N2["Which matters most at long context<br/><b>04</b>"] N3["Where prefill must be chunked to be usable<br/><b>05</b>"] N4["Or separated entirely onto different hardware<br/><b>06</b>"] N5["With KV moving between tiers<br/><b>07</b>"] N6["Under a scheduler that knows about priorities<br/><b>08</b>"] N7["Producing multiple tokens per pass where possible<br/><b>09</b>"] N8["In as few bits as the hardware supports<br/><b>10</b>"] N9["Using kernels you can now write yourself<br/><b>11</b>"] N10["Or that a compiler writes for you<br/><b>12</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 class N0,N1,N2 neutral class N3,N4 io class N5,N6 queue class N7,N8 compute class N9,N10 memory
How to read this section#
If you’re operating a system: files 01, 02, 05, and 08 affect decisions you make today. Files 04 and 10 affect model selection. The rest are for when you hit their specific bottleneck.
If you’re building: files 11 and 12 are the practical skills. 06 and 07 are the architectures to understand before you need them.