Below the API

Building the Platform: A Roadmap

Expert Advanced 1h 15m Difficulty 4/5

Prerequisites all of Section XII

★ The order matters more than the components. This is the file to read before you start building.


1. The principle#

Build in the order of value delivered per unit of effort, and never build ahead of demand.

Most inference platforms fail not because a component was built badly, but because the team built a sophisticated scheduler before they had a second model, or a multi-region control plane before they had a second region.

The wrong question: "what does a complete platform look like?"
The right question: "what is the most valuable thing I don't have?"

2. Stage 0 — before you have a platform#

YOU HAVE: one model, one team, one deployment.

DO
  □ use vLLM or SGLang directly (Section VIII.14)
  □ OpenAI-compatible API
  □ basic metrics (Section X.10) and structured per-request logs
  □ a capacity calculation on paper (Section V.15)
  □ tune the engine properly (Section VIII.04)

DON'T
  □ build a registry (a YAML file is fine)
  □ build a router (one deployment doesn't need one)
  □ build a control plane

TIME: days
VALUE: this is 80% of what most single-model deployments need.

Many organizations should stop here. A well-tuned single deployment serves a lot of traffic.


3. Stage 1 — the essentials (weeks 1-6)#

Trigger: a second team wants to use the model, or a second model appears.

BUILD, IN THIS ORDER

1. GATEWAY (Section XII.08)                          ~2 weeks
   auth, token-based rate limiting, validation,
   tokenization, streaming, usage recording
   → WHY FIRST: it's where policy lives, and everything else
     assumes it exists. Without it, every new tenant is a
     bespoke integration.

2. PER-REQUEST LOGGING AND COST ATTRIBUTION          ~3 days
   (Sections X.10, XI.03)
   → WHY EARLY: you cannot make good decisions without this data,
     and the data must accumulate before you need it.

3. MODEL REGISTRY, minimal (Section XII.02)          ~1 week
   YAML in Git + an object store + checksums + a promotion gate
   → WHY: you now have more than one thing to deploy.
   → DON'T build the service yet.

4. BASIC ROUTER (Section XII.06)                     ~1 week
   least-KV-loaded with power-of-two-choices,
   plus session affinity
   → WHY: as soon as you have multiple replicas, round-robin
     is costing you 20-40%.

5. SLOs AND DASHBOARDS (Sections XI.01, X.10)        ~1 week
   → WHY: you now have users with expectations.

TOTAL: ~6 weeks. This covers 90% of platform value.

Notice what’s not here: no scheduler, no control plane, no multi-tenancy machinery beyond rate limiting. Those come when the pain arrives.


4. Stage 2 — efficiency (weeks 7-16)#

Trigger: cost is a concern, or you have 5+ models.

6. PREFIX-AWARE ROUTING (Section XII.07)             ~1 week
   → measure the value first (Section XII.07 §7)
   → typically 15-25% capacity for chat workloads

7. QUANTIZATION PIPELINE (Section VII)               ~2 weeks
   offline quantization, validation harness,
   registry integration
   → typically 1.5-2x

8. PLACEMENT CONTROLLER (Section XII.04)             ~3 weeks
   demand-proportional replica counts,
   hot/warm/cold tiering
   → typically 20-40% GPU reduction with many models

9. MULTI-TENANCY (Section XII.05)                    ~2 weeks
   weighted fair queueing, per-tenant quotas
   including concurrency and KV, showback dashboards
   → enables consolidation, which is 2-4x

10. AUTOSCALING (Sections VIII.07, XI.04)            ~1 week
    predictive (cron) + reactive on queue wait
    → typically 15-25% for diurnal workloads

TOTAL: ~10 weeks. Combined effect: 3-6x better economics.

Item 9 is the one that unlocks the largest win (consolidation, Section IX.12/XII.05), but it requires the trust that comes from items 2 and 5.


5. Stage 3 — scale and reliability (months 5-12)#

Trigger: a second region, or an availability commitment, or 20+ models.

11. CONTROL PLANE PROPER (Section XII.01)            ~4 weeks
    registry as a service, config distribution with
    last-known-good, deployment controller

12. FAULT TOLERANCE (Section XI.05)                  ~3 weeks
    watchdogs, automatic cordoning, game days,
    circuit breakers

13. MODEL ROLLOUT AUTOMATION (Section XI.07)         ~3 weeks
    progressive delivery with quality guards

14. MULTI-REGION (Section XI.06)                     ~6 weeks
    regional weight distribution, cross-region quota,
    failover, DR plan

15. CASCADES AND FALLBACKS (Section XII.09)          ~3 weeks

16. MULTI-LORA (Section VIII.09)                     ~2 weeks
    if you serve fine-tunes — a 50x economic change

TOTAL: ~5 months.

6. Stage 4 — advanced (year 2+)#

Only if the measurements justify it.

17. Disaggregated prefill/decode (Section XIII.06)
    → if your prefill/decode ratio is extreme and you're at scale

18. Distributed / tiered KV cache (Section IX.11)
    → if you have very long shared prefixes and a fast interconnect

19. Custom kernels (Section VI.08, XIII.11)
    → if you've exhausted everything else and have a specific gap

20. Speculative decoding at scale (Section VII.13)
    → if you're decode-bound at low batch

21. Learned routing / cascades (Section XII.09)
    → if heuristic routing is leaving measurable value

Each of these should be justified by a measurement showing the specific bottleneck it addresses. They are all significant engineering investments with modest returns compared to stages 1-2.


7. The value curve#

Cumulative value delivered
   ▲
   │                                              ┌─── stage 4 (marginal)
   │                                    ┌─────────┘
   │                          ┌─────────┘  stage 3 (reliability)
   │              ┌───────────┘
   │              │  stage 2 (efficiency: 3-6x)
   │    ┌─────────┘
   │    │  stage 1 (essentials: enables everything)
   │ ┌──┘
   │ │ stage 0
   └─┴───────────────────────────────────────────────────► effort
     days   6 weeks      16 weeks        12 months    year 2+

The curve flattens hard. Stage 1 and 2 together are ~4 months of work and deliver most of the value. Stage 4 is a year of work for incremental gains.

Teams that build stage 4 first fail. They have a sophisticated disaggregated serving architecture and no cost attribution, so they can’t tell whether it helped.


8. The anti-roadmap: what not to build#

✗ YOUR OWN INFERENCE ENGINE
  Section VIII.14. Hundreds of engineer-years of existing work.
  Contribute to vLLM/SGLang instead.

✗ A GENERIC MODEL-SERVING ABSTRACTION
  "We'll support any model type through a plugin interface."
  You'll spend six months on the abstraction and serve one model type.

✗ A SCHEDULER BEFORE YOU HAVE MULTIPLE MODELS
  A YAML file naming which model runs where is fine for 5 models.

✗ MULTI-REGION BEFORE YOU HAVE A SECOND REGION'S TRAFFIC

✗ A CUSTOM UI
  Grafana. Every time.

✗ YOUR OWN METRICS SYSTEM
  Prometheus.

✗ PERFECT COST ATTRIBUTION BEFORE APPROXIMATE COST ATTRIBUTION
  weighted tokens gets you 90% of the way; ship it and refine.

9. Staffing and ownership#

TEAM SIZE vs SCOPE

  1-2 engineers   stage 0-1. Use managed services aggressively.
  3-5 engineers   stage 2. This is where most platforms should be.
  6-10 engineers  stage 3. Multi-region, high reliability.
  10+             stage 4, or you're serving at very large scale.

SKILLS NEEDED
  systems / distributed (the majority of the work)
  ML familiarity (enough to reason about quality)
  SRE / operations
  performance engineering (at least one person deep on GPUs)

WHAT TO OUTSOURCE
  the engine (vLLM/SGLang/TRT-LLM)
  observability infrastructure (Prometheus/Grafana/OTel)
  Kubernetes itself
  the base model (unless you train)

The most common staffing mistake: hiring ML engineers for what is mostly a distributed systems job. The platform needs systems engineers who understand enough ML, not ML researchers who tolerate systems.


10. Deciding what to build next#

Use this every quarter:

FOR EACH CANDIDATE:

  1. What measurement shows this is a problem?
     (no measurement → don't build it)

  2. What's the expected gain?
     Use Amdahl (Section VII.01). Below 15%, deprioritize.

  3. What's the effort?

  4. What's the ongoing operational cost?
     (a component you must maintain forever)

  5. Is there a cheaper alternative?
     (a config flag? an existing tool? doing nothing?)

  6. gain / (effort + ongoing_cost) → rank

THEN: build the top item. Only the top item. Measure. Repeat.

Step 1 is the discipline. “It would be nice to have a scheduler” is not a justification; “we’re at 34% average utilization because placement is static and demand shifted” is.


11. Hands-on exercise#

A. Assess where you are. For a platform you work on (or your organization’s situation), mark which items from stages 0-4 exist. What’s the highest-value missing item?

B. Build the roadmap. Using the framework in section 10, rank the missing items by gain/effort. Write the next-two-quarters plan with justifications.

C. Justify with measurement. For your top item, produce the measurement that proves it’s a problem. If you can’t measure it, that’s your real next task.

D. The anti-roadmap check. Is anything on your current roadmap in section 8’s list? Why?

E. Estimate the value curve. For your situation, estimate the cumulative value at each stage. Where does your curve flatten?

F. Write the one-pager. Produce a one-page platform strategy: current state, the next three things to build, the measurement justifying each, and what you’re explicitly not building. This is a real deliverable.


12. Interview questions#

  1. What would you build first for an inference platform, and why that?
  2. What would you not build, and why?
  3. How do you decide what to build next?
  4. Why does building a scheduler before having multiple models fail?
  5. What’s the largest single economic lever in a multi-team platform?
  6. What skills does an inference platform team need?
  7. Where does the value curve flatten, and what does that imply?

13. Further reading#

  • [FUNDAMENTAL] Google SRE Book and The Site Reliability Workbook
  • [ESTABLISHED] Sculley et al., “Hidden Technical Debt in Machine Learning Systems” (2015)
  • [REFERENCE] Public writeups of internal ML platforms (Uber Michelangelo, Meta FBLearner, and successors) — read for the organizational lessons, not the specifics
  • Next: Section XIII — Advanced LLM Inference

↑↓ navigate ↵ open