The capstone. Put a control plane around everything you have built: declare a model, and the platform deploys it, routes to it, scales it, rolls it, and bills for it.
1. What you build#
A small but complete multi-tenant inference platform on Kubernetes:
CONTROL PLANE
┌───────────────┐ ┌────────────────┐ ┌──────────────┐
│ Model registry│ → │ Controller / │ → │ Autoscaler │
│ (CRD or YAML) │ │ reconciler │ │ (queue-based)│
└───────────────┘ └────────────────┘ └──────────────┘
│ creates / updates
▼
DATA PLANE
client → Gateway (P12) → Engine pods (P08 or vLLM) → GPUs / CPUs
│ │
└──── metrics ───────┴──→ Prometheus → Grafana, alerts, cost reportOne kubectl apply -f model.yaml should take a model from a registry entry to a served,
routed, autoscaled, observable endpoint.
Diagram — The platform#
flowchart TB
DEV["kubectl apply ModelDeployment"] --> REG
subgraph CP["Control plane"]
REG["Model registry"] --> CTL["Controller / reconciler"]
CTL --> AS["Autoscaler<br/>queue-based"]
CTL --> RO["Rollout manager<br/>canary + gates"]
end
subgraph DP["Data plane"]
GW["Gateway - Project 12"] --> ENG["Engine pods - Project 08 or vLLM"]
ENG --> HW["GPUs / CPUs"]
end
U["Tenants"] --> GW
CTL -.->|"creates, updates"| ENG
RO -.->|"traffic weights"| GW
AS -.->|"replicas"| ENG
GW -.-> PROM["Prometheus"]
ENG -.-> PROM
PROM --> AS
PROM --> GRAF["Grafana, alerts, cost report"]
class REG,CTL,AS,RO queue
class GW,ENG,HW compute
class PROM,GRAF io
class DEV,U neutral2. Why it matters#
This is the job. Platform and infrastructure roles in this field are not asked to invent attention kernels; they are asked to make fleets of engines dependable and cheap. Everything in Sections XI and XII becomes one working system here — and one portfolio piece that speaks directly to that role.
3. Read first#
- XII.01 — Control vs data plane
- XII.02 — Model registry
- XII.03 — Kubernetes for GPUs
- XII.05 — Multi-tenancy
- XI.04 — Autoscaling
- XI.07 — Model rollouts
- XI.08 — Observability
- VIII.07 — Autoscaling and cold starts
4. Spec#
apiVersion: lab.inference/v1
kind: ModelDeployment
metadata: { name: chat-small }
spec:
model: { uri: "hf://Qwen/Qwen2.5-0.5B-Instruct", revision: "<commit sha>", sha256: "..." }
engine: { image: "ghcr.io/you/p08-engine:1.4", maxNumSeqs: 32, maxLen: 4096 }
resources:{ gpu: 0, cpu: "4", memory: "8Gi" }
scaling: { min: 1, max: 6, metric: queue_wait_p95_ms, target: 200,
scaleDownStabilizationSeconds: 300 }
rollout: { strategy: canary, steps: [5, 25, 100], gate: { ttft_p95_ms: 800, error_rate: 0.01 } }
tenants: [ { name: team-a, tpm: 200000, priority: high },
{ name: team-b, tpm: 50000, priority: batch } ]Components:
Registry immutable, content-addressed model versions (URI + revision + checksum)
Controller reconciles ModelDeployment → Deployment, Service, gateway route, ServiceMonitor
(kopf / controller-runtime / a plain loop over the API — your choice)
Model cache weights on a PVC or node-local cache; init container verifies the checksum
Probes startup (model loading can take minutes), readiness (loaded AND not saturated),
liveness (process only)
Gateway Project 12, configured by the controller
Autoscaler scales on queue wait / running sequences per replica — NOT on GPU utilization
Rollout weighted canary through the gateway, automatic promote or rollback on the gate
Observability dashboards: TTFT, ITL, queue wait, running batch size, tokens/s, KV usage,
errors, per-tenant usage ; SLO burn-rate alert
Cost per-tenant $ = tokens-share-weighted GPU-seconds × price (XI.03)5. Milestones#
- Cluster.
kindork3d. If you have a GPU node: NVIDIA GPU Operator, and request devices via the device plugin or DRA. - Hand-written manifests first. Deploy one engine + gateway manually. Get probes right.
- Registry + controller.
kubectl applyaModelDeployment; the controller creates everything. Deleting it cleans everything up. ChangingmaxNumSeqstriggers a rolling update. - Observability. Prometheus scraping engines and gateway; one Grafana dashboard per model; an SLO alert.
- Autoscaling. Drive load with your generator. Watch replicas follow queue wait. Measure
cold-start time and decide your
minfrom it. - Canary rollout. Ship a new engine image at 5% → 25% → 100%. Then ship a deliberately slow one and watch the gate roll it back.
- Multi-tenancy. Quotas, priority, and a noisy-neighbour test.
- Cost report. Per-tenant cost for a one-hour replay. Idle capacity must be accounted for, not hidden.
- Game day. Kill a pod mid-stream, drain a node, corrupt a weight file, exhaust a quota. Write a short runbook entry for each.
6. Starter skeleton#
import kopf, kubernetes as k8s
@kopf.on.create("lab.inference", "v1", "modeldeployments")
@kopf.on.update("lab.inference", "v1", "modeldeployments")
def reconcile(spec, name, namespace, patch, **_):
desired = render_deployment(name, spec) # pure function: spec → manifests
kopf.adopt(desired) # owner refs → garbage collection on delete
apply(desired) # server-side apply; idempotent
gateway.upsert_route(model=name, service=f"{name}.{namespace}.svc", tenants=spec["tenants"])
patch.status["observedRevision"] = spec["model"]["revision"]
@kopf.timer("lab.inference", "v1", "modeldeployments", interval=15)
def autoscale(spec, name, namespace, status, **_):
s = spec["scaling"]
cur = current_replicas(name, namespace)
val = prom(f'histogram_quantile(0.95, sum(rate(queue_wait_seconds_bucket{{model="{name}"}}[1m])) by (le))') * 1000
want = min(s["max"], max(s["min"], math.ceil(cur * val / s["target"])))
if want > cur:
scale(name, namespace, want) # scale up immediately
elif want < cur and stable_for(name, s["scaleDownStabilizationSeconds"]):
scale(name, namespace, cur - 1) # scale down slowly, one at a time7. What to measure#
| Measurement | Expectation to write down first |
|---|---|
Time from kubectl apply to first successful token | Dominated by image pull + weight load |
| Cold start breakdown: schedule / pull / load / warmup | Know each term |
| Scale-up reaction time vs load ramp | Is the SLO violated before capacity arrives? |
| TTFT p95 during a rolling update | No visible blip if draining works |
| Canary: time to detect and roll back a bad version | Minutes, automatically |
| Noisy tenant at 10× quota: effect on others’ p95 | None |
| Utilization: served tokens / capacity tokens | Probably far lower than you hoped |
| Cost per million tokens, per tenant | With idle capacity included |
8. Done when#
- One manifest deploys a model end to end; deleting it removes everything.
- Replicas follow load, with a recorded cold-start number justifying
min. - A bad canary is rolled back without human action.
- Quotas and priorities hold under a noisy-neighbour test.
- A dashboard answers “is it healthy, is it fast, who is using it, what does it cost.”
- You wrote a two-page design doc: architecture, SLOs, capacity model, failure modes, what you would change for 100 GPUs. That document is the real deliverable.
- You can answer Checkpoint F from ROADMAP.md with this system as your worked example.
9. Common pitfalls#
Autoscaling on GPU utilization. It reads ~100% from one request to saturation (X.04). Scale on queue wait or running sequences.
Liveness probe that depends on the model. Slow load or a busy engine → restart loop.
Readiness that only checks “loaded”. A saturated pod keeps receiving traffic.
Scaling down as fast as up. With multi-minute cold starts, flapping is very expensive.
No graceful drain. SIGTERM kills in-flight streams on every rollout. Use a preStop hook
and a long enough termination grace period.
Mutable model references (main, latest). Two replicas end up serving different weights.
Pin the revision and verify the checksum.
Downloading weights on every pod start. Cache them on the node or a shared volume.
Control plane on the request path. If the controller is down, existing traffic must keep flowing.
10. Stretch goals#
- Swap your engine for vLLM and your gateway’s routing for the Gateway API Inference Extension or llm-d; compare what they give you against what you built.
- Scale to zero with request buffering at the gateway; measure the first-request penalty.
- Multi-LoRA: one base model, per-tenant adapters loaded on demand.
- Bin-packing: several small models on one GPU (MPS / time-slicing / MIG), with interference measured (XII.04).
- A second “region” (second cluster) with failover and a DR drill (XI.06).
- GitOps:
ModelDeploymentmanifests in a repo, reconciled by Argo CD or Flux.
11. Interview questions this project answers#
- Design a multi-tenant LLM serving platform. What is in the control plane, what is in the data plane, and why does the split matter?
- Which metric do you autoscale on, and why not GPU utilization?
- How do you roll out a new model version safely?
- How do you attribute cost per tenant on shared GPUs?
- How do you keep cold starts from violating your SLO?
- What happens to in-flight streams during a deploy?
12. After this#
You have built the stack from a NumPy matmul to a platform. See “After XIV” in ROADMAP.md: contribute upstream, reproduce a paper, or halve a real system’s cost per million tokens and write it up.