Section VIII.06 covered load balancing between replicas of one model. This file covers the platform’s routing layer, which also decides which model and which pool.
1. The routing decisions#
A platform router makes several decisions per request, in order:
1. WHICH MODEL? alias resolution, cascades, fallbacks
2. WHICH POOL? general / long-context / premium / batch
3. WHICH REPLICA? load-aware, prefix-aware, session-affine
4. ADMIT OR REJECT? quota, capacity, priorityEach has different inputs and different failure modes.
2. Decision 1 — which model#
ALIAS RESOLUTION
client asks for "chat-large" → registry says llama-3-70b-instruct v1.5.0
→ decouples client-facing names from deployments (Section XII.02)
VERSION SELECTION
canary: 5% → v1.6.0, 95% → v1.5.0, sticky by conversation_id
(Section VIII.10)
COST-AWARE / CASCADE ROUTING
route by request characteristics:
short, simple query → the 8B model
complex reasoning → the 70B model
code → the code-specialized model
→ Section XII.09
TENANT-SPECIFIC
tenant has a fine-tuned adapter → route to a replica with that LoRA
tenant is restricted to certain models → enforce here3. Decision 2 — which pool#
POOL SEGREGATION (Section VIII.06 — one of the highest-value decisions)
general the bulk of traffic; tuned for typical request shapes
long-context requests > 16k tokens; different max_model_len,
different max_num_seqs, isolated so they don't
poison the general pool's batch
premium reserved capacity, lower batch size, better ITL
batch no latency SLO; huge batch; may use spot instances
canary the version under test
ROUTING RULES
if request.prompt_tokens > 16000: → long-context pool
elif tenant.tier == "premium": → premium pool
elif request.priority == "batch": → batch pool
elif canary_selected(request): → canary pool
else: → general poolSegregating long-context requests is worth calling out again. A single 128k request in a pool tuned for 4k consumes 32 sequences’ worth of KV and blocks prefill for everyone. It’s a one-line routing rule that protects your p99.
4. Decision 3 — which replica#
Covered in Section VIII.06; the platform version adds the multi-model dimension:
def choose_replica(request, model, pool):
candidates = routing_table[model][pool]
candidates = [r for r in candidates if r.ready and not r.draining]
# 1. Session affinity (for prefix cache locality)
if request.conversation_id:
preferred = consistent_hash(request.conversation_id, candidates)
if preferred.kv_usage < AFFINITY_RELEASE_THRESHOLD: # e.g. 0.85
return preferred
# 2. Prefix affinity (for shared system prompts)
prefix_key = hash(request.token_ids[:PREFIX_WINDOW])
warm = [r for r in candidates if prefix_key in r.known_prefixes
and r.kv_usage < AFFINITY_RELEASE_THRESHOLD]
if warm:
return min(warm, key=lambda r: r.kv_usage)
# 3. Load-aware: power of two choices
a, b = random.sample(candidates, min(2, len(candidates)))
chosen = a if a.load_score() < b.load_score() else b
remember_prefix(prefix_key, chosen)
return chosen
def load_score(r):
# combine the signals; KV usage dominates
return (0.6 * r.kv_usage
+ 0.3 * (r.queue_depth / r.max_queue)
+ 0.1 * (r.running_batch / r.max_batch))The AFFINITY_RELEASE_THRESHOLD is the crucial tuning knob. Too high and a popular prefix
overloads one replica; too low and you lose cache locality. 0.80-0.90 is a reasonable range;
tune by measuring end-to-end TTFT.
5. Decision 4 — admit or reject#
The router's admission decision is DIFFERENT from the engine's:
ROUTER-LEVEL
is the tenant over quota? → 429
is EVERY replica of this model at capacity? → 503
is the estimated wait > the client's deadline? → 503
is the request malformed / over limits? → 400
ENGINE-LEVEL (Section VIII.03)
is there KV space right now?
does this fit in the token budget for this step?The router should reject early (cheap) but the engine has the authoritative state. Both are needed.
The router's view of capacity is STALE (metrics scraped every 1-5 s).
→ it should be conservative, and the engine must still be able to reject.
→ never let the router promise what the engine can't deliver.6. The routing table#
STRUCTURE
model_alias → model_id, version(s), weights
model_id + version → pools
pool → [replica endpoints, capabilities, current load]
DISTRIBUTION (Section XII.01)
control plane computes it, pushes to routers
routers cache it with a TTL and last-known-good behavior
load metrics scraped separately, more frequently
FRESHNESS REQUIREMENTS
routing table (which replicas exist): seconds to minutes — slow is fine
load metrics (how busy): 1-5 seconds — matters
prefix location hints: seconds; approximate is fineLoad metrics need to be fresher than the routing table, and they’re the thing that must keep flowing. A router with a 5-minute-old routing table but current load data works fine; the reverse does not.
7. Routing for multi-LoRA#
With multi-LoRA (Section VIII.09), routing gains a dimension:
request specifies adapter "team-x-v3"
→ prefer a replica that already has it loaded (avoids a load)
→ and concentrate requests for the same adapter on the same replica
(fewer distinct adapters per batch → lower grouped-GEMM overhead)
def choose_replica_lora(request, candidates):
adapter = request.lora_id
loaded = [r for r in candidates if adapter in r.loaded_adapters]
if loaded:
return min(loaded, key=load_score)
# not loaded anywhere: pick the least loaded replica with room
# for another adapter
room = [r for r in candidates if len(r.loaded_adapters) < r.max_loras]
return min(room or candidates, key=load_score)Adapter affinity matters more than you’d expect: the grouped-GEMM overhead grows with the number of distinct adapters in a batch, so concentrating them is worth some load imbalance.
8. Failure handling in the router#
REPLICA UNHEALTHY
remove from candidates. But: distinguish
- not ready (loading, warming) → exclude
- saturated (KV at 100%) → DEPRIORITIZE, don't exclude
(excluding cascades — Section VIII.06)
- failing health checks → exclude, alert
- draining → exclude for new requests
ALL REPLICAS SATURATED
→ 503 with Retry-After, don't queue at the router
(queueing at the router hides the problem from the engine's
scheduler, which has better information)
MODEL NOT DEPLOYED
→ if it's a cold-tier model: trigger a load, queue the request,
tell the client to expect latency (or 503 with a long Retry-After)
→ else: 404
CONTROL PLANE UNREACHABLE
→ keep using the cached routing table (Section XII.01)
→ alert, but keep serving“Deprioritize, don’t exclude” for saturated replicas is the rule that prevents cascading failure: excluding them concentrates load on the remaining replicas, which then saturate.
9. Production implications#
- The router is where LLM-specific logic lives. Build it; don’t use a generic LB (Section XII.01).
- Segregate long-context requests into their own pool. Single highest-value routing rule.
- Session affinity by conversation ID, with a load release valve.
- Load metrics must be fresh (1-5 s). Stale load data makes routing worse than random.
- Never exclude saturated replicas; deprioritize them.
- Don’t queue at the router. Let the engine’s scheduler decide.
- Adapter affinity for multi-LoRA.
- Instrument routing decisions: which replica, why, and what the alternatives were. When routing goes wrong you need to see the decision.
10. Common mistakes#
Generic load balancer with round-robin. Wrong on every axis (Section VIII.06).
Excluding saturated replicas. Cascading failure.
Queueing at the router. Hides state from the scheduler and adds a second queue.
Stale load metrics. Routing on 60-second-old data is worse than random.
No session affinity with prefix caching enabled. 1/N hit rate.
Affinity without a release valve. One popular conversation overloads a replica.
No long-context segregation. One request poisons the pool.
Global least-loaded across router instances. Herding (Section VIII.06).
11. Hands-on exercise#
A. Build the router. Implement the four decisions from section 1. Support: aliases, pool selection, session affinity, prefix affinity, load-aware selection with power-of-two, and admission. This is a core piece of Project 15.
B. Measure the affinity threshold. Sweep AFFINITY_RELEASE_THRESHOLD from 0.5 to 1.0 in a
simulation with a popular shared prefix. Plot prefix hit rate and p99 TTFT. Find the optimum.
C. Demonstrate the cascade. Simulate excluding saturated replicas versus deprioritizing them under increasing load. Show the cascade in the first case.
D. Metric freshness. Simulate routing with load metrics that are 1 s, 5 s, 30 s, and 300 s stale. At what staleness does load-aware routing become worse than random?
E. Long-context segregation. Simulate a pool with 2% of requests at 64k context, with and without segregation. Measure p99 TTFT for the 98%.
F. LoRA affinity. Simulate 20 adapters across 4 replicas with and without adapter affinity. Measure the average number of distinct adapters per batch.
12. Interview questions#
- What decisions does a platform router make, and in what order?
- Why segregate long-context requests, and what does it protect?
- Why deprioritize rather than exclude a saturated replica?
- Why shouldn’t the router queue requests?
- How fresh must load metrics be, and what happens when they’re stale?
- How does multi-LoRA change replica selection?
- What’s the tension between cache affinity and load balance, and how do you resolve it?
13. Further reading#
- [FUNDAMENTAL] Mitzenmacher, “The Power of Two Choices”
- [ESTABLISHED] Kubernetes Gateway API Inference Extension —
InferencePoolis a stable (v1) API; the endpoint picker now lives in llm-d - [ESTABLISHED] SGLang’s router implementation
- [ESTABLISHED] Envoy load balancing documentation for the general patterns
- Next: 07 — Prefix-aware routing