Below the API

ROADMAP — The Complete Learning Path

This is the dependency graph of the entire curriculum, plus time estimates and checkpoints. Read it once now, then come back whenever you want to know “what am I allowed to skip?”


The one-line version#

I -> II -> III -> IV -> V -> VI -> VII -> VIII -> IX -> X -> XI -> XII -> XIII -> XIV

The honest version#

The sections are not a straight line; they are a lattice. Here is what actually depends on what:

flowchart TB
  I["I. Fundamentals<br/>vocabulary + the physics"] --> II["II. Computer Systems<br/>CPU, memory, OS, PCIe"]
  I --> III["III. ML Fundamentals<br/>linear algebra to transformer"]
  II --> IV["IV. Neural Network Inference<br/>the forward pass as engineering"]
  III --> IV
  IV --> V["V. LLM Inference<br/>prefill/decode, KV cache, batching"]
  IV --> VI["VI. GPU Computing<br/>SMs, warps, HBM, CUDA, profiling"]
  V --> VII["VII. Inference Optimization<br/>quantization, fusion, FlashAttention"]
  VI --> VII
  VII --> VIII["VIII. Serving Systems<br/>vLLM, SGLang, TGI, Triton"]
  VIII --> IX["IX. Distributed Inference<br/>TP/PP/EP, NCCL, NVLink"]
  VIII --> X["X. Memory & Performance<br/>roofline, profiling"]
  IX --> XI["XI. Production Inference<br/>SLOs, capacity, cost, observability"]
  X --> XI
  XI --> XII["XII. Platform Engineering<br/>registry, K8s, routing"]
  XII --> XIII["XIII. Advanced LLM Inference<br/>MoE, MLA, disaggregation"]
  XIII --> XIV["XIV. Research & Frontier"]
  VI -.->|"soft: readable any time after VI"| X
  VIII -.->|"soft: if impatient for production"| XI

  class I,II,III,IV neutral
  class V,VI memory
  class VII,VIII compute
  class IX,X queue
  class XI,XII io
  class XIII,XIV warn

Green is the core (V, VI). Solid arrows are the intended order; dotted arrows are shortcuts you may take.

Hard dependencies (do not skip): I → III → IV → V. Everything else in the curriculum leans on those four.

Soft dependencies: you can read VI before V if you prefer hardware-first. You can read X any time after VI. XI and XII are readable after VIII if you are impatient for production material.


Time estimates#

Estimates assume ~1 hour of focused study per file including its exercise, and that you actually do the exercises. Skimming is 4-5x faster and roughly 4-5x less useful.

SectionFilesReadingWith exercisesCumulative
I — Fundamentals116 h12 h12 h
II — Computer Systems138 h16 h28 h
III — ML Fundamentals128 h18 h46 h
IV — NN Inference128 h16 h62 h
V — LLM Inference1512 h26 h88 h
VI — GPU Computing1210 h24 h112 h
VII — Optimization1410 h22 h134 h
VIII — Serving1410 h20 h154 h
IX — Distributed129 h18 h172 h
X — Memory & Perf118 h18 h190 h
XI — Production118 h14 h204 h
XII — Platform118 h14 h218 h
XIII — Advanced LLM1210 h18 h236 h
XIV — Research118 h12 h248 h
Projects 01-1515—120-200 h~400 h

Total: roughly 250 hours of study + 150 hours of building.

At 10 h/week that is about 10 months. At 20 h/week, 5 months. This is a real curriculum, not a weekend. Treat the estimates as a planning tool, not a promise.


Checkpoints#

After each checkpoint, you should be able to answer the listed questions without looking anything up. If you cannot, go back — do not push forward. Inference engineering compounds: gaps early become confusion later.

Checkpoint A — after Section I#

  1. What is the difference between latency and throughput, and why can improving one hurt the other?
  2. A 7B model in FP16 — how much GPU memory just for weights? Show the arithmetic.
  3. Why is generating token #500 of a response no more expensive than token #5 in FLOPs, but more expensive in memory traffic?
  4. What is arithmetic intensity and why is batch size the main lever on it?

Checkpoint B — after Section III#

  1. Multiply a 2x3 by a 3x2 matrix by hand.
  2. Explain attention in three sentences with no equations.
  3. Why is softmax computed with a max subtraction?
  4. What does BF16 trade away relative to FP16, and why does that trade matter for inference?

Checkpoint C — after Section V (the big one)#

  1. Draw the timeline of a single request through prefill and decode.
  2. Compute the KV cache size for Llama-3-8B at 8k context, batch 32. Show every term.
  3. Explain continuous batching to a backend engineer in 60 seconds.
  4. Why does PagedAttention exist? What exactly was wasteful before it?
  5. At what batch size does a decode step stop being memory-bound on your hardware?

Checkpoint D — after Section VII#

  1. Explain FlashAttention’s core trick without saying “it’s faster.”
  2. Give three reasons INT4 weight-only quantization helps decode more than prefill.
  3. When does speculative decoding lose?

Checkpoint E — after Section X#

  1. Given a slow inference service, list the first five measurements you take, in order.
  2. Draw a roofline and place decode, prefill, and a fused elementwise kernel on it.
  3. nvidia-smi says 100% GPU utilization. Why might the GPU still be mostly idle?

Checkpoint F — after Section XII#

  1. Design the serving stack for a 70B model, 10k concurrent users, p95 TTFT < 500 ms, fixed budget of 64 H100s. Justify every choice.
  2. How do you attribute cost per tenant in a shared inference cluster?

Where the projects fit#

Section I    ......................................
Section II   ................. [P02 CPU matmul benchmark]
Section III  ................. [P01 NumPy inference engine]
Section IV   ..............................
Section V    ................. [P04 tiny transformer] [P05 KV cache]
Section VI   ................. [P03 first GPU kernel]
Section VII  ................. [P10 quantized inference]
Section VIII ................. [P06 server] [P07 dynamic batching] [P08 continuous batching]
Section IX   ................. [P13 multi-GPU] [P14 distributed]
Section X    ................. [P09 KV cache manager] [P11 GPU benchmark suite]
Section XI   ................. [P12 inference gateway]
Section XII  ................. [P15 mini platform]

Projects 01-05 are the “understand it” projects. 06-11 are the “build it” projects. 12-15 are the “operate it” projects.


Study patterns that work#

The 3-pass read. Pass 1: read sections 1-4 of a file (what/why/analogy/example) and stop. Pass 2 (same day): read 5-8 (technical/under the hood/performance/production). Pass 3 (next day): do the exercise and the interview questions. Spacing beats cramming, especially for material this dense.

Measure everything. Every claim in this curriculum about performance can be verified on your own hardware. Do it. The engineer who has personally measured that a memcpy over PCIe Gen4 x16 tops out near 25 GB/s will never again propose an architecture that ships KV caches over PCIe every token.

Explain it to someone. The 12-part file structure exists partly so you can teach from it. If you cannot do part 3 (analogy) in your own words, you have not understood part 5.

Keep a numbers.md. Your own file of measured constants: your GPU’s HBM bandwidth, your peak FP16 TFLOPs, your PCIe bandwidth, your model’s tokens/sec at batch 1 and batch 64. This file becomes your intuition.


After XIV#

You will have finished the curriculum but not the field. Sensible next moves:

  • Contribute to vLLM or SGLang. Start with issues labeled good first issue; the schedulers and the attention backends are where the education is.
  • Reproduce a recent paper’s benchmark and try to break it.
  • Take one production system and cut its cost per million tokens in half. Write down what you did. That document is your career.

↑↓ navigate ↵ open