Below the API

Inference Platform Engineering

ExpertModule11 topics~13h
completed

Topics, in order

About this module

Goal: build the system that lets many teams serve many models on shared hardware, safely and efficiently.

Section XI was about running a service well. This section is about running a platform — which is a different problem, with the central insight from Section IX.12: consolidating traffic is often worth more than any kernel optimization.

Files#

#FileLevelTime
01Platform overview: control plane vs data planeAdvanced60 min
02Model registryIntermediate60 min
03Kubernetes for GPUs ★Advanced90 min
04GPU scheduling and model placementAdvanced90 min
05Multi-tenancy and quotasAdvanced75 min
06Routing strategiesAdvanced75 min
07Prefix-aware routingAdvanced60 min
08Inference gatewaysAdvanced75 min
09Cascades and fallbacksAdvanced60 min
10Cluster and capacity managementAdvanced60 min
11Building the platform: a roadmap ★Advanced75 min

The thread#

flowchart TD
  N0["A platform separates what changes slowly from what happens per request<br/><b>01</b>"]
  N1["It needs to know what models exist<br/><b>02</b>"]
  N2["And where to run them, on scarce hardware<br/><b>03, 04</b>"]
  N3["Shared fairly between teams<br/><b>05</b>"]
  N4["With requests routed intelligently<br/><b>06, 07</b>"]
  N5["Through a front door that handles everything common<br/><b>08</b>"]
  N6["With fallbacks when the primary path fails or is too expensive<br/><b>09</b>"]
  N7["On a cluster whose capacity is actively managed<br/><b>10</b>"]
  N8["Built incrementally, in the right order<br/><b>11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint F (part 2)#

  1. What belongs in the control plane and what in the data plane? Why does the distinction matter?
  2. Design the model placement algorithm for a cluster with 200 GPUs and 40 models.
  3. How do you attribute cost to tenants, and what’s the term everyone forgets?
  4. Why does prefix-aware routing matter more than load balance for some workloads?
  5. What’s the first thing you’d build, and what’s the last?

↑↓ navigate ↵ open