Goal: run an inference service that meets commitments, costs what you planned, and doesn’t wake you at 3 a.m.
Sections I-X taught you how inference works and how to make it fast. This section is about running it: SLOs, capacity, cost, reliability, observability, and security. It is less about GPUs and more about judgment.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | SLOs, SLIs, and latency budgets ★ | Intermediate | 75 min |
| 02 | Capacity planning ★ | Advanced | 90 min |
| 03 | Cost per token ★ | Advanced | 75 min |
| 04 | Autoscaling in production | Advanced | 60 min |
| 05 | Fault tolerance and failure modes | Advanced | 75 min |
| 06 | Multi-region and disaster recovery | Advanced | 60 min |
| 07 | Model rollouts | Intermediate | 60 min |
| 08 | Observability for LLM services | Intermediate | 60 min |
| 09 | Security and isolation | Advanced | 75 min |
| 10 | Abuse prevention, rate limiting, quotas | Advanced | 60 min |
| 11 | Scenario: 70B, 10,000 users, fixed budget ★ | Expert | 120 min |
The thread#
flowchart TD N0["Promise something specific<br/><b>01</b>"] N1["Size the fleet to keep the promise<br/><b>02</b>"] N2["Know what it costs and why<br/><b>03</b>"] N3["Adjust as load changes — slowly<br/><b>04</b>"] N4["Survive failures (05) and regions going away<br/><b>06</b>"] N5["Change the model without breaking anything<br/><b>07</b>"] N6["See what's happening<br/><b>08</b>"] N7["Keep tenants isolated (09) and abusers out<br/><b>10</b>"] N8["Then do all of it at once, under constraints<br/><b>11</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 class N0,N1 neutral class N2,N3 io class N4,N5 queue class N6,N7 compute class N8 memory
Checkpoint F (part 1)#
- Write an SLO for an LLM chat service. Why those specific metrics and thresholds?
- Size a fleet for 10,000 concurrent users on a 70B model. Show every step.
- Derive cost per million tokens and name the five terms that dominate it.
- Your service is over capacity and cannot scale in time. What happens?
- How do you attribute cost to individual tenants in a shared cluster?