Below the API

Distributed Inference

AdvancedModule12 topics~14h 45m
completed

Topics, in order

About this module

Goal: understand what happens when one GPU isn’t enough — and, equally important, when adding GPUs makes things worse.

The central lesson of this section: parallelism is not free. Every form of it trades communication for computation, and whether that trade is good depends on your interconnect, your model, and your batch size. The arithmetic in file 09 is what separates people who can size a distributed deployment from people who guess.

Files#

#FileLevelTime
01Why distribute (and when not to)Intermediate60 min
02Data parallelismIntermediate45 min
03Tensor parallelism ★Advanced120 min
04Pipeline parallelismAdvanced75 min
05Expert parallelismAdvanced75 min
06Sequence and context parallelismAdvanced60 min
07Collectives and NCCLAdvanced75 min
08Interconnects: NVLink, PCIe, InfiniBand, RDMAAdvanced75 min
09Communication cost math ★Advanced90 min
10Multi-node inferenceAdvanced75 min
11Distributed KV cacheAdvanced60 min
12When more GPUs make things worse ★Advanced75 min

The thread#

flowchart TD
  N0["The model doesn't fit, or is too slow on one GPU<br/><b>01</b>"]
  N1["Replicate it (02) — helps throughput, not latency"]
  N2["Split each layer's weights (03) — helps latency, costs an AllReduce per layer"]
  N3["Or split by layers (04) — cheap communication, but bubbles"]
  N4["Or split by experts (05) — for MoE, with All-to-All"]
  N5["Or split by sequence (06) — for very long context"]
  N6["All of which run on collectives (07) over an interconnect<br/><b>08</b>"]
  N7["Whose cost you can compute<br/><b>09</b>"]
  N8["And which gets much worse across nodes<br/><b>10</b>"]
  N9["Where the KV cache also has to live somewhere<br/><b>11</b>"]
  N10["And where, past a point, more GPUs make everything slower<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10

  class N0,N1,N2 neutral
  class N3,N4 io
  class N5,N6 queue
  class N7,N8 compute
  class N9,N10 memory

↑↓ navigate ↵ open