★ The order matters more than the components. This is the file to read before you start building.
1. The principle#
Build in the order of value delivered per unit of effort, and never build ahead of demand.
Most inference platforms fail not because a component was built badly, but because the team built a sophisticated scheduler before they had a second model, or a multi-region control plane before they had a second region.
The wrong question: "what does a complete platform look like?"
The right question: "what is the most valuable thing I don't have?"2. Stage 0 — before you have a platform#
YOU HAVE: one model, one team, one deployment.
DO
□ use vLLM or SGLang directly (Section VIII.14)
□ OpenAI-compatible API
□ basic metrics (Section X.10) and structured per-request logs
□ a capacity calculation on paper (Section V.15)
□ tune the engine properly (Section VIII.04)
DON'T
□ build a registry (a YAML file is fine)
□ build a router (one deployment doesn't need one)
□ build a control plane
TIME: days
VALUE: this is 80% of what most single-model deployments need.Many organizations should stop here. A well-tuned single deployment serves a lot of traffic.
3. Stage 1 — the essentials (weeks 1-6)#
Trigger: a second team wants to use the model, or a second model appears.
BUILD, IN THIS ORDER
1. GATEWAY (Section XII.08) ~2 weeks
auth, token-based rate limiting, validation,
tokenization, streaming, usage recording
→ WHY FIRST: it's where policy lives, and everything else
assumes it exists. Without it, every new tenant is a
bespoke integration.
2. PER-REQUEST LOGGING AND COST ATTRIBUTION ~3 days
(Sections X.10, XI.03)
→ WHY EARLY: you cannot make good decisions without this data,
and the data must accumulate before you need it.
3. MODEL REGISTRY, minimal (Section XII.02) ~1 week
YAML in Git + an object store + checksums + a promotion gate
→ WHY: you now have more than one thing to deploy.
→ DON'T build the service yet.
4. BASIC ROUTER (Section XII.06) ~1 week
least-KV-loaded with power-of-two-choices,
plus session affinity
→ WHY: as soon as you have multiple replicas, round-robin
is costing you 20-40%.
5. SLOs AND DASHBOARDS (Sections XI.01, X.10) ~1 week
→ WHY: you now have users with expectations.
TOTAL: ~6 weeks. This covers 90% of platform value.Notice what’s not here: no scheduler, no control plane, no multi-tenancy machinery beyond rate limiting. Those come when the pain arrives.
4. Stage 2 — efficiency (weeks 7-16)#
Trigger: cost is a concern, or you have 5+ models.
6. PREFIX-AWARE ROUTING (Section XII.07) ~1 week
→ measure the value first (Section XII.07 §7)
→ typically 15-25% capacity for chat workloads
7. QUANTIZATION PIPELINE (Section VII) ~2 weeks
offline quantization, validation harness,
registry integration
→ typically 1.5-2x
8. PLACEMENT CONTROLLER (Section XII.04) ~3 weeks
demand-proportional replica counts,
hot/warm/cold tiering
→ typically 20-40% GPU reduction with many models
9. MULTI-TENANCY (Section XII.05) ~2 weeks
weighted fair queueing, per-tenant quotas
including concurrency and KV, showback dashboards
→ enables consolidation, which is 2-4x
10. AUTOSCALING (Sections VIII.07, XI.04) ~1 week
predictive (cron) + reactive on queue wait
→ typically 15-25% for diurnal workloads
TOTAL: ~10 weeks. Combined effect: 3-6x better economics.Item 9 is the one that unlocks the largest win (consolidation, Section IX.12/XII.05), but it requires the trust that comes from items 2 and 5.
5. Stage 3 — scale and reliability (months 5-12)#
Trigger: a second region, or an availability commitment, or 20+ models.
11. CONTROL PLANE PROPER (Section XII.01) ~4 weeks
registry as a service, config distribution with
last-known-good, deployment controller
12. FAULT TOLERANCE (Section XI.05) ~3 weeks
watchdogs, automatic cordoning, game days,
circuit breakers
13. MODEL ROLLOUT AUTOMATION (Section XI.07) ~3 weeks
progressive delivery with quality guards
14. MULTI-REGION (Section XI.06) ~6 weeks
regional weight distribution, cross-region quota,
failover, DR plan
15. CASCADES AND FALLBACKS (Section XII.09) ~3 weeks
16. MULTI-LORA (Section VIII.09) ~2 weeks
if you serve fine-tunes — a 50x economic change
TOTAL: ~5 months.6. Stage 4 — advanced (year 2+)#
Only if the measurements justify it.
17. Disaggregated prefill/decode (Section XIII.06)
→ if your prefill/decode ratio is extreme and you're at scale
18. Distributed / tiered KV cache (Section IX.11)
→ if you have very long shared prefixes and a fast interconnect
19. Custom kernels (Section VI.08, XIII.11)
→ if you've exhausted everything else and have a specific gap
20. Speculative decoding at scale (Section VII.13)
→ if you're decode-bound at low batch
21. Learned routing / cascades (Section XII.09)
→ if heuristic routing is leaving measurable valueEach of these should be justified by a measurement showing the specific bottleneck it addresses. They are all significant engineering investments with modest returns compared to stages 1-2.
7. The value curve#
Cumulative value delivered
▲
│ ┌─── stage 4 (marginal)
│ ┌─────────┘
│ ┌─────────┘ stage 3 (reliability)
│ ┌───────────┘
│ │ stage 2 (efficiency: 3-6x)
│ ┌─────────┘
│ │ stage 1 (essentials: enables everything)
│ ┌──┘
│ │ stage 0
└─┴───────────────────────────────────────────────────► effort
days 6 weeks 16 weeks 12 months year 2+The curve flattens hard. Stage 1 and 2 together are ~4 months of work and deliver most of the value. Stage 4 is a year of work for incremental gains.
Teams that build stage 4 first fail. They have a sophisticated disaggregated serving architecture and no cost attribution, so they can’t tell whether it helped.
8. The anti-roadmap: what not to build#
✗ YOUR OWN INFERENCE ENGINE
Section VIII.14. Hundreds of engineer-years of existing work.
Contribute to vLLM/SGLang instead.
✗ A GENERIC MODEL-SERVING ABSTRACTION
"We'll support any model type through a plugin interface."
You'll spend six months on the abstraction and serve one model type.
✗ A SCHEDULER BEFORE YOU HAVE MULTIPLE MODELS
A YAML file naming which model runs where is fine for 5 models.
✗ MULTI-REGION BEFORE YOU HAVE A SECOND REGION'S TRAFFIC
✗ A CUSTOM UI
Grafana. Every time.
✗ YOUR OWN METRICS SYSTEM
Prometheus.
✗ PERFECT COST ATTRIBUTION BEFORE APPROXIMATE COST ATTRIBUTION
weighted tokens gets you 90% of the way; ship it and refine.9. Staffing and ownership#
TEAM SIZE vs SCOPE
1-2 engineers stage 0-1. Use managed services aggressively.
3-5 engineers stage 2. This is where most platforms should be.
6-10 engineers stage 3. Multi-region, high reliability.
10+ stage 4, or you're serving at very large scale.
SKILLS NEEDED
systems / distributed (the majority of the work)
ML familiarity (enough to reason about quality)
SRE / operations
performance engineering (at least one person deep on GPUs)
WHAT TO OUTSOURCE
the engine (vLLM/SGLang/TRT-LLM)
observability infrastructure (Prometheus/Grafana/OTel)
Kubernetes itself
the base model (unless you train)The most common staffing mistake: hiring ML engineers for what is mostly a distributed systems job. The platform needs systems engineers who understand enough ML, not ML researchers who tolerate systems.
10. Deciding what to build next#
Use this every quarter:
FOR EACH CANDIDATE:
1. What measurement shows this is a problem?
(no measurement → don't build it)
2. What's the expected gain?
Use Amdahl (Section VII.01). Below 15%, deprioritize.
3. What's the effort?
4. What's the ongoing operational cost?
(a component you must maintain forever)
5. Is there a cheaper alternative?
(a config flag? an existing tool? doing nothing?)
6. gain / (effort + ongoing_cost) → rank
THEN: build the top item. Only the top item. Measure. Repeat.Step 1 is the discipline. “It would be nice to have a scheduler” is not a justification; “we’re at 34% average utilization because placement is static and demand shifted” is.
11. Hands-on exercise#
A. Assess where you are. For a platform you work on (or your organization’s situation), mark which items from stages 0-4 exist. What’s the highest-value missing item?
B. Build the roadmap. Using the framework in section 10, rank the missing items by gain/effort. Write the next-two-quarters plan with justifications.
C. Justify with measurement. For your top item, produce the measurement that proves it’s a problem. If you can’t measure it, that’s your real next task.
D. The anti-roadmap check. Is anything on your current roadmap in section 8’s list? Why?
E. Estimate the value curve. For your situation, estimate the cumulative value at each stage. Where does your curve flatten?
F. Write the one-pager. Produce a one-page platform strategy: current state, the next three things to build, the measurement justifying each, and what you’re explicitly not building. This is a real deliverable.
12. Interview questions#
- What would you build first for an inference platform, and why that?
- What would you not build, and why?
- How do you decide what to build next?
- Why does building a scheduler before having multiple models fail?
- What’s the largest single economic lever in a multi-team platform?
- What skills does an inference platform team need?
- Where does the value curve flatten, and what does that imply?
13. Further reading#
- [FUNDAMENTAL] Google SRE Book and The Site Reliability Workbook
- [ESTABLISHED] Sculley et al., “Hidden Technical Debt in Machine Learning Systems” (2015)
- [REFERENCE] Public writeups of internal ML platforms (Uber Michelangelo, Meta FBLearner, and successors) — read for the organizational lessons, not the specifics
- Next: Section XIII — Advanced LLM Inference