Goal: understand what happens when one GPU isn’t enough — and, equally important, when adding GPUs makes things worse.
The central lesson of this section: parallelism is not free. Every form of it trades communication for computation, and whether that trade is good depends on your interconnect, your model, and your batch size. The arithmetic in file 09 is what separates people who can size a distributed deployment from people who guess.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Why distribute (and when not to) | Intermediate | 60 min |
| 02 | Data parallelism | Intermediate | 45 min |
| 03 | Tensor parallelism ★ | Advanced | 120 min |
| 04 | Pipeline parallelism | Advanced | 75 min |
| 05 | Expert parallelism | Advanced | 75 min |
| 06 | Sequence and context parallelism | Advanced | 60 min |
| 07 | Collectives and NCCL | Advanced | 75 min |
| 08 | Interconnects: NVLink, PCIe, InfiniBand, RDMA | Advanced | 75 min |
| 09 | Communication cost math ★ | Advanced | 90 min |
| 10 | Multi-node inference | Advanced | 75 min |
| 11 | Distributed KV cache | Advanced | 60 min |
| 12 | When more GPUs make things worse ★ | Advanced | 75 min |
The thread#
flowchart TD N0["The model doesn't fit, or is too slow on one GPU<br/><b>01</b>"] N1["Replicate it (02) — helps throughput, not latency"] N2["Split each layer's weights (03) — helps latency, costs an AllReduce per layer"] N3["Or split by layers (04) — cheap communication, but bubbles"] N4["Or split by experts (05) — for MoE, with All-to-All"] N5["Or split by sequence (06) — for very long context"] N6["All of which run on collectives (07) over an interconnect<br/><b>08</b>"] N7["Whose cost you can compute<br/><b>09</b>"] N8["And which gets much worse across nodes<br/><b>10</b>"] N9["Where the KV cache also has to live somewhere<br/><b>11</b>"] N10["And where, past a point, more GPUs make everything slower<br/><b>12</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 class N0,N1,N2 neutral class N3,N4 io class N5,N6 queue class N7,N8 compute class N9,N10 memory