1. The distinction#
DATA PLANE handles every request. Must be fast, simple, and always up.
CONTROL PLANE decides what the data plane does. Can be slower, more complex,
and briefly unavailable. CONTROL PLANE DATA PLANE
(changes per minute/hour) (per request)
┌──────────────────────┐ ┌──────────────────┐
│ model registry │ │ gateway │
│ placement scheduler │───────▶│ router │
│ capacity manager │ config │ engines │
│ autoscaler │ │ │
│ quota manager │ │ │
│ deployment controller│ │ │
└──────────────────────┘ └──────────────────┘
if this is down for 5 min: if this is down for 5 s:
nothing changes, but users see errors
everything keeps workingThe critical property: the data plane must keep working when the control plane is unavailable. This is the difference between a platform that degrades and one that fails.
Diagram — Two planes, different jobs#
flowchart TB
U["Users"] --> GW
subgraph CP["CONTROL PLANE - slow path, decides"]
REG["Model registry"] --> CTL["Controllers / reconcilers"]
CTL --> SCALE["Autoscaler"]
CTL --> CFG["Routing + quota config"]
end
subgraph DP["DATA PLANE - hot path, serves"]
GW["Gateway"] --> ENG["Engine replicas"] --> GPU["GPUs"]
end
CFG -.->|"config push"| GW
SCALE -.->|"replica count"| ENG
ENG -.->|"metrics"| SCALE
class REG,CTL,SCALE,CFG queue
class GW,ENG,GPU compute
class U neutral2. Why this matters specifically for inference platforms#
The control plane makes EXPENSIVE, SLOW decisions:
- which models run on which GPUs (minutes to enact — model loading)
- how many replicas of each (minutes)
- quota allocations (policy)
The data plane makes CHEAP, FAST decisions:
- which replica gets this request (microseconds)
- is this tenant over quota (microseconds, from cached state)
- admit or reject (microseconds)
Conflating them means either:
✗ the data plane is slow (calling the control plane per request), or
✗ the control plane is in the request path (a single point of failure)The common architectural mistake: a per-request lookup to a central service. It adds latency to every request and makes that service a hard dependency. Push the state to the data plane and let it cache.
3. Simple analogy#
Air traffic control versus the aircraft.
Control (control plane): assigns routes, altitudes, and slots. Works on a timescale of minutes. If the tower loses radio for two minutes, aircraft continue on their assigned routes safely.
The aircraft (data plane): flies, second by second. It has its clearance cached and can complete the flight without further instruction.
An aircraft that needed continuous instruction to stay airborne would be a bad design. So is a data plane that needs the control plane per request.
4. What goes where#
CONTROL PLANE
Model registry what models exist, their configs and artifacts
Placement scheduler which model runs on which nodes
Capacity manager how many replicas; scaling decisions
Deployment controller rollouts, canaries, rollbacks
Quota manager tenant allocations (the POLICY)
Configuration distribution pushing config to the data plane
Cost accounting aggregating usage into bills
Fleet health node status, cordoning, draining
DATA PLANE
Gateway auth, rate limit ENFORCEMENT, validation
Router replica selection
Engines the actual inference
Local caches model→replica map, quota state, auth tokens
SHARED STATE (the interface)
the routing table control → data, pushed
quota counters data → control, aggregated
health and load metrics data → control, scraped
usage records data → control, streamed5. The interface between them#
CONTROL → DATA (configuration push)
routing table: model_id → [replica endpoints, weights, capabilities]
quota policy: tenant → limits
feature flags
DELIVERY: push (watch/subscribe) with local caching and a TTL.
The data plane uses its last-known-good config if the
control plane is unreachable.
DATA → CONTROL (telemetry)
per-replica load (KV usage, batch size, queue depth)
usage records (tokens by tenant)
health status
DELIVERY: scrape (Prometheus) or push (streaming). Lossy is acceptable.The “last-known-good config” property is the whole point. If the control plane is down, the data plane keeps routing using the config it has. New deployments don’t happen, but existing traffic is served.
// RoutingTable is the data plane's view of the control plane's decisions.
type RoutingTable struct {
table atomic.Pointer[map[string][]string] // model id -> replica addresses
lastUpdate atomic.Int64
}
func NewRoutingTable(ctx context.Context, controlPlaneURL string) *RoutingTable {
rt := &RoutingTable{}
rt.table.Store(&map[string][]string{})
go func() {
for {
if t, err := fetch(ctx, controlPlaneURL); err == nil {
rt.table.Store(&t) // swap the whole table atomically
rt.lastUpdate.Store(time.Now().Unix())
} else {
slog.Warn("control plane unreachable; using cached table",
"age", time.Since(time.Unix(rt.lastUpdate.Load(), 0)))
}
select {
case <-ctx.Done():
return
case <-time.After(10 * time.Second):
}
}
}()
return rt
}
// Lookup NEVER blocks on the control plane. It always answers from the cached table.
func (rt *RoutingTable) Lookup(modelID string) []string { return (*rt.table.Load())[modelID] }
6. Failure modes#
CONTROL PLANE DOWN
✓ existing traffic continues (cached routing table)
✗ no new deployments
✗ no autoscaling
✗ quota state not reconciled (data plane enforces locally)
→ acceptable for minutes to hours
DATA PLANE COMPONENT DOWN
gateway down → total outage for that region
router down → total outage (unless the gateway can route directly)
one engine down → 1/N capacity lost
→ must be highly available; run multiple instances of each
CONFIG PUSH BROKEN (control plane up, but pushing bad config)
→ THE DANGEROUS ONE. A bad routing table can take everything down.
→ mitigations: validate config before pushing, canary config changes,
version the config, allow instant rollback, and have the data plane
reject obviously-invalid config (e.g. an empty routing table)The “empty routing table” check is worth implementing explicitly. A control plane bug that pushes an empty or malformed table is a total outage; a data plane that refuses to accept it is a degraded deployment.
7. Build vs adopt#
COMPONENT BUILD? OR USE
Model registry maybe MLflow, an artifact store + metadata DB
Placement scheduler maybe Kubernetes scheduler + custom scoring
Capacity manager maybe KEDA + custom metrics
Deployment controller no Argo Rollouts, Flagger
Gateway maybe Envoy, Kong, or custom
Router usually this is where the LLM-specific logic lives
Engines NO vLLM, SGLang, TGI, TRT-LLM
Quota manager maybe Redis + custom logicThe router is the component most worth building yourself, because prefix-aware, load-aware, cost-aware routing is LLM-specific and no off-the-shelf load balancer does it (Section VIII.06, XII.06-07).
The engine is the component least worth building. (Section VIII.14.)
8. Production implications#
- Draw the line explicitly. Document what’s control plane and what’s data plane.
- The data plane must survive control plane outage. Test this: kill the control plane and verify traffic continues.
- Never call the control plane per request. Cache and push.
- Validate configuration before pushing, and version it for rollback.
- Run the data plane components redundantly. The gateway and router are single points of failure if you don’t.
- Keep the data plane simple. Complexity there costs latency on every request and availability for everyone.
9. Common mistakes#
Per-request calls to a central service. Latency and a hard dependency.
No local cache of the routing table. Control plane outage = total outage.
Control plane and data plane in the same process. They have different availability requirements.
Unvalidated config push. One bad push takes everything down.
Building the engine. (Section VIII.14.)
Not building the router. Generic load balancers are wrong for LLM traffic.
Too much logic in the data plane. Every millisecond is multiplied by every request.
10. Hands-on exercise#
A. Draw the boundary. For a platform you know or are designing, list every component and classify it. Which are in the request path?
B. Test control plane failure. Kill your control plane (or its equivalent) and verify traffic continues. How long until something breaks? What breaks first?
C. Implement the cached routing table. Build the RoutingTable class from section 5, with
last-known-good behavior and an age metric. Verify it survives a control plane outage.
D. Config validation. Implement validation on the data plane side: reject empty tables, tables with unreachable endpoints, or tables missing required models. Test with deliberately bad config.
E. Measure the latency. Compare a per-request control plane lookup against a cached lookup. Quantify the latency added per request and the availability impact.
11. Interview questions#
- What’s the difference between control plane and data plane, and why does it matter here?
- What must happen when the control plane is unavailable?
- Why should you never call a central service per request?
- What’s the dangerous failure mode of a healthy control plane?
- Which platform components would you build and which would you adopt?
- Why is the router worth building yourself?
- How do you protect against a bad configuration push?
12. Further reading#
- [FUNDAMENTAL] The control/data plane distinction in networking (SDN literature)
- [REFERENCE] Kubernetes architecture — a well-designed example of the separation
- [REFERENCE] Envoy’s xDS protocol — config distribution done well
- Next: 02 — Model registry