Below the API

LLM Inference

BasicModule15 topics~22h 15m
completed

Topics, in order

About this module

This is the core section of the curriculum. Everything before it was preparation; everything after it is elaboration.

If you learn one section properly, make it this one. The concepts here — prefill vs decode, the KV cache, continuous batching, paged attention — are what separate people who can operate an LLM serving system from people who can only start one.

Files#

#FileLevelTime
01Transformer inference overviewIntermediate60 min
02Tokenization and its consequencesBeginner60 min
03Prefill vs decode ★Intermediate120 min
04Autoregressive generationIntermediate60 min
05The KV cache ★Intermediate120 min
06KV cache math ★Intermediate90 min
07Context length and how it scalesAdvanced75 min
08Batching: static and dynamicIntermediate75 min
09Continuous batching ★Advanced120 min
10PagedAttention ★Advanced120 min
11Prefix and prompt cachingAdvanced90 min
12Speculative decodingAdvanced90 min
13Sampling and decoding strategiesIntermediate75 min
14Streaming inferenceIntermediate60 min
15Capacity math: worked examples ★Advanced120 min

★ = do not skip.

The thread#

flowchart TD
  N0["Text becomes tokens<br/><b>02</b>"]
  N1["The prompt is processed in one parallel pass — PREFILL<br/><b>03</b>"]
  N2["Then tokens are generated one at a time — DECODE<br/><b>03, 04</b>"]
  N3["Which is only affordable because we cache K and V<br/><b>05</b>"]
  N4["And that cache costs memory that we can compute exactly<br/><b>06</b>"]
  N5["Memory that grows with context length<br/><b>07</b>"]
  N6["So we batch to amortize weight reads<br/><b>08</b>"]
  N7["But static batching wastes 60-80% of slots<br/><b>08</b>"]
  N8["So we schedule at every step — CONTINUOUS BATCHING<br/><b>09</b>"]
  N9["And manage KV memory in pages to avoid fragmentation<br/><b>10</b>"]
  N10["And reuse KV across requests when prefixes match<br/><b>11</b>"]
  N11["And generate several tokens per weight read when we can<br/><b>12</b>"]
  N12["Choosing each token from a distribution<br/><b>13</b>"]
  N13["Streaming them out as they appear<br/><b>14</b>"]
  N14["All of which we can size on paper before we build it<br/><b>15</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14

  class N0,N1,N2 neutral
  class N3,N4,N5 io
  class N6,N7,N8 queue
  class N9,N10,N11 compute
  class N12,N13,N14 memory

Checkpoint C — the big one#

Before moving on, you should be able to answer without notes:

  1. Draw the timeline of a single request through prefill and decode.
  2. Compute the KV cache size for Llama-3-8B at 8k context, batch 32. Show every term.
  3. Explain continuous batching to a backend engineer in 60 seconds.
  4. Why does PagedAttention exist? What exactly was wasteful before it?
  5. At what batch size does a decode step stop being memory-bound on your hardware?
  6. A user reports 4-second TTFT. List the five things you check, in order.
  7. Given 8×H100 and Llama-3-70B, how many concurrent users at 8k context? Show the arithmetic.

↑↓ navigate ↵ open