Below the API

Production Inference Engineering

AdvancedModule11 topics~13h 30m
completed

Topics, in order

About this module

Goal: run an inference service that meets commitments, costs what you planned, and doesn’t wake you at 3 a.m.

Sections I-X taught you how inference works and how to make it fast. This section is about running it: SLOs, capacity, cost, reliability, observability, and security. It is less about GPUs and more about judgment.

Files#

#FileLevelTime
01SLOs, SLIs, and latency budgets ★Intermediate75 min
02Capacity planning ★Advanced90 min
03Cost per token ★Advanced75 min
04Autoscaling in productionAdvanced60 min
05Fault tolerance and failure modesAdvanced75 min
06Multi-region and disaster recoveryAdvanced60 min
07Model rolloutsIntermediate60 min
08Observability for LLM servicesIntermediate60 min
09Security and isolationAdvanced75 min
10Abuse prevention, rate limiting, quotasAdvanced60 min
11Scenario: 70B, 10,000 users, fixed budget ★Expert120 min

The thread#

flowchart TD
  N0["Promise something specific<br/><b>01</b>"]
  N1["Size the fleet to keep the promise<br/><b>02</b>"]
  N2["Know what it costs and why<br/><b>03</b>"]
  N3["Adjust as load changes — slowly<br/><b>04</b>"]
  N4["Survive failures (05) and regions going away<br/><b>06</b>"]
  N5["Change the model without breaking anything<br/><b>07</b>"]
  N6["See what's happening<br/><b>08</b>"]
  N7["Keep tenants isolated (09) and abusers out<br/><b>10</b>"]
  N8["Then do all of it at once, under constraints<br/><b>11</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint F (part 1)#

  1. Write an SLO for an LLM chat service. Why those specific metrics and thresholds?
  2. Size a fleet for 10,000 concurrent users on a 70B model. Show every step.
  3. Derive cost per million tokens and name the five terms that dominate it.
  4. Your service is over capacity and cannot scale in time. What happens?
  5. How do you attribute cost to individual tenants in a shared cluster?

↑↓ navigate ↵ open