Goal of this section: give you the vocabulary and the physics of inference. By the end you will be able to read any inference discussion without getting lost, and — more importantly — you will be able to estimate whether a proposed system can possibly work, using nothing but arithmetic.
This section is deliberately light on ML. We are not yet asking “what is a transformer.” We are asking “what does it mean to run any trained model, and what makes that hard?”
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | What is inference? | Beginner | 45 min |
| 02 | Training vs inference | Beginner | 45 min |
| 03 | Model anatomy: parameters, weights, activations | Beginner | 60 min |
| 04 | Tensors and the forward pass | Beginner | 60 min |
| 05 | Latency, throughput, and the metrics that matter | Beginner | 75 min |
| 06 | Batching: the central tradeoff | Beginner | 60 min |
| 07 | Compute vs memory | Intermediate | 75 min |
| 08 | FLOPs, bandwidth, arithmetic intensity | Intermediate | 90 min |
| 09 | CPU vs GPU | Beginner | 60 min |
| 10 | Why inference is not normal backend engineering | Intermediate | 60 min |
| 11 | Why inference gets expensive at scale | Intermediate | 60 min |
The thread running through this section#
flowchart TD N0["A model is a big pile of numbers<br/><b>03</b>"] N1["Running it means moving those numbers through arithmetic<br/><b>04</b>"] N2["Moving numbers costs time; arithmetic costs time<br/><b>07, 08</b>"] N3["Which one dominates depends on how much arithmetic per byte you do<br/><b>08</b>"] N4["Batching is the main lever on that ratio<br/><b>06</b>"] N5["Which is why every serving system is, underneath, a batching scheduler<br/><b>10, 11</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 class N0,N1 neutral class N2 io class N3 queue class N4 compute class N5 memory
Checkpoint A#
Before moving to Section II, you should be able to answer, without looking anything up:
- Latency vs throughput — define both and give an example where improving one hurts the other.
- A 7B model in FP16: how much memory for weights? Show the arithmetic.
- Why is generating token #500 no more expensive in FLOPs than token #5, but more expensive in memory traffic?
- What is arithmetic intensity, and why is batch size the main lever on it?
- Why does a GPU beat a CPU at inference — is it the FLOPs, the bandwidth, or both?