Below the API

Fundamentals

FoundationsModule11 topics~11h 30m
completed

Topics, in order

About this module

Goal of this section: give you the vocabulary and the physics of inference. By the end you will be able to read any inference discussion without getting lost, and — more importantly — you will be able to estimate whether a proposed system can possibly work, using nothing but arithmetic.

This section is deliberately light on ML. We are not yet asking “what is a transformer.” We are asking “what does it mean to run any trained model, and what makes that hard?”

Files#

#FileLevelTime
01What is inference?Beginner45 min
02Training vs inferenceBeginner45 min
03Model anatomy: parameters, weights, activationsBeginner60 min
04Tensors and the forward passBeginner60 min
05Latency, throughput, and the metrics that matterBeginner75 min
06Batching: the central tradeoffBeginner60 min
07Compute vs memoryIntermediate75 min
08FLOPs, bandwidth, arithmetic intensityIntermediate90 min
09CPU vs GPUBeginner60 min
10Why inference is not normal backend engineeringIntermediate60 min
11Why inference gets expensive at scaleIntermediate60 min

The thread running through this section#

flowchart TD
  N0["A model is a big pile of numbers<br/><b>03</b>"]
  N1["Running it means moving those numbers through arithmetic<br/><b>04</b>"]
  N2["Moving numbers costs time; arithmetic costs time<br/><b>07, 08</b>"]
  N3["Which one dominates depends on how much arithmetic per byte you do<br/><b>08</b>"]
  N4["Batching is the main lever on that ratio<br/><b>06</b>"]
  N5["Which is why every serving system is, underneath, a batching scheduler<br/><b>10, 11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5

  class N0,N1 neutral
  class N2 io
  class N3 queue
  class N4 compute
  class N5 memory

Checkpoint A#

Before moving to Section II, you should be able to answer, without looking anything up:

  1. Latency vs throughput — define both and give an example where improving one hurts the other.
  2. A 7B model in FP16: how much memory for weights? Show the arithmetic.
  3. Why is generating token #500 no more expensive in FLOPs than token #5, but more expensive in memory traffic?
  4. What is arithmetic intensity, and why is batch size the main lever on it?
  5. Why does a GPU beat a CPU at inference — is it the FLOPs, the bandwidth, or both?

↑↓ navigate ↵ open