Goal: build the system that lets many teams serve many models on shared hardware, safely and efficiently.
Section XI was about running a service well. This section is about running a platform — which is a different problem, with the central insight from Section IX.12: consolidating traffic is often worth more than any kernel optimization.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | Platform overview: control plane vs data plane | Advanced | 60 min |
| 02 | Model registry | Intermediate | 60 min |
| 03 | Kubernetes for GPUs ★ | Advanced | 90 min |
| 04 | GPU scheduling and model placement | Advanced | 90 min |
| 05 | Multi-tenancy and quotas | Advanced | 75 min |
| 06 | Routing strategies | Advanced | 75 min |
| 07 | Prefix-aware routing | Advanced | 60 min |
| 08 | Inference gateways | Advanced | 75 min |
| 09 | Cascades and fallbacks | Advanced | 60 min |
| 10 | Cluster and capacity management | Advanced | 60 min |
| 11 | Building the platform: a roadmap ★ | Advanced | 75 min |
The thread#
flowchart TD N0["A platform separates what changes slowly from what happens per request<br/><b>01</b>"] N1["It needs to know what models exist<br/><b>02</b>"] N2["And where to run them, on scarce hardware<br/><b>03, 04</b>"] N3["Shared fairly between teams<br/><b>05</b>"] N4["With requests routed intelligently<br/><b>06, 07</b>"] N5["Through a front door that handles everything common<br/><b>08</b>"] N6["With fallbacks when the primary path fails or is too expensive<br/><b>09</b>"] N7["On a cluster whose capacity is actively managed<br/><b>10</b>"] N8["Built incrementally, in the right order<br/><b>11</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 class N0,N1 neutral class N2,N3 io class N4,N5 queue class N6,N7 compute class N8 memory
Checkpoint F (part 2)#
- What belongs in the control plane and what in the data plane? Why does the distinction matter?
- Design the model placement algorithm for a cluster with 200 GPUs and 40 models.
- How do you attribute cost to tenants, and what’s the term everyone forgets?
- Why does prefix-aware routing matter more than load balance for some workloads?
- What’s the first thing you’d build, and what’s the last?