1. What multi-tenancy has to provide#
1. ISOLATION one tenant cannot see another's data
2. FAIRNESS one tenant cannot starve another
3. ACCOUNTABILITY you know who used what
4. FLEXIBILITY tenants can have different SLAs and limitsSection XI.09 covered isolation (security) and XI.10 covered limits (enforcement). This file is about the platform-level design: how tenants share a cluster.
2. The economic case for sharing#
5 teams, each with a dedicated deployment:
each sized for their peak, at 60% target utilization
each gets ~15 req/s peak, needs 2 replicas → 10 replicas total
average batch per replica: 11
Shared platform:
combined peak 75 req/s (peaks don't coincide — this is the key)
needs 6 replicas at the same target utilization
average batch per replica: 28
→ 40% fewer GPUs, AND 2.5x better per-GPU throughput (Section IX.12)
→ combined effect: ~2.5-4x better economicsTwo effects compound: peak smoothing (peaks don’t coincide) and batch concentration. This is why a shared platform beats per-team deployments, and it’s the core value proposition of Section XII.
The organizational cost: teams give up control, and you must provide fairness guarantees convincing enough that they accept it. That is largely a trust problem, solved by transparency (section 6).
3. The isolation spectrum#
LEVEL ISOLATION EFFICIENCY USE
Shared model, shared weak highest internal, mutually
engine process trusting teams
Shared model, separate weak-med high different teams,
prefix cache namespaces same trust domain
Separate engine process medium medium untrusted tenants,
per tenant same model
MIG partition strong medium-low regulatory isolation
Separate node strongest low air-gapped
Separate cluster strongest lowest sovereign / contractualMost platforms should default to the second row: shared model, per-tenant prefix cache namespaces. It preserves nearly all the efficiency while eliminating the timing side channel (Section XI.09).
// Prefix cache namespacing: include the tenant in the block hash,
// so one tenant can never hit (or time) another tenant's cached prefix.
func blockHash(tenantID string, prevHash uint64, tokenIDs []int32) uint64 {
h := fnv.New64a()
h.Write([]byte(tenantID))
binary.Write(h, binary.LittleEndian, prevHash)
binary.Write(h, binary.LittleEndian, tokenIDs)
return h.Sum64()
}
Cost: a shared system prompt used by 5 tenants is cached 5 times. For a 2,000-token system prompt at 128 KiB/token, that’s 5 × 256 MB = 1.28 GB instead of 256 MB. Usually acceptable.
4. Fairness under contention#
Quotas bound the maximum; fairness governs what happens when everyone wants their maximum at once.
WITHOUT FAIRNESS
tenant A (quota 100k tok/min) and tenant B (quota 10k tok/min)
both submit continuously.
FCFS scheduling → A gets ~90% of capacity because it submits more.
B is effectively starved despite being within quota.
WITH WEIGHTED FAIR QUEUEING
weights proportional to committed capacity (or tier)
→ each gets its share regardless of submission rate// drr.go — deficit round robin over tenants, weighted by their share.
package main
import "fmt"
type Request struct {
Tenant string
Cost float64 // weighted tokens
}
type FairScheduler struct {
weights map[string]float64
deficit map[string]float64
queues map[string][]Request
order []string
idx int
}
func NewFairScheduler(order []string, weights map[string]float64) *FairScheduler {
return &FairScheduler{weights, map[string]float64{}, map[string][]Request{}, order, 0}
}
func (s *FairScheduler) Enqueue(r Request) { s.queues[r.Tenant] = append(s.queues[r.Tenant], r) }
func (s *FairScheduler) Dequeue() (Request, bool) {
pending := false
for _, q := range s.queues {
pending = pending || len(q) > 0
}
for pending {
t := s.order[s.idx]
q := s.queues[t]
if len(q) > 0 && s.deficit[t] >= q[0].Cost { // keep serving this tenant while its credit lasts
s.deficit[t] -= q[0].Cost
s.queues[t] = q[1:]
return q[0], true
}
if len(q) == 0 {
s.deficit[t] = 0 // don't accumulate credit while idle
}
// move on: the next tenant earns its weight in credit
s.idx = (s.idx + 1) % len(s.order)
s.deficit[s.order[s.idx]] += s.weights[s.order[s.idx]]
}
return Request{}, false
}
func main() {
s := NewFairScheduler([]string{"big", "small"}, map[string]float64{"big": 300, "small": 100})
for i := 0; i < 40; i++ { // both tenants flood the queue with identical requests
s.Enqueue(Request{"big", 100})
s.Enqueue(Request{"small", 100})
}
served := map[string]int{}
for i := 0; i < 40; i++ {
r, _ := s.Dequeue()
served[r.Tenant]++
}
fmt.Println(served) // map[big:30 small:10] — a 3:1 split, matching the weights
}
The deficit[t] = 0.0 reset when idle is important: without it, a tenant that was idle
accumulates credit and then bursts, starving others.
5. Quota models#
HARD PARTITION
tenant A gets nodes 0-3, tenant B gets 4-7.
✓ perfect isolation, trivially fair
✗ no statistical multiplexing — you lose the whole economic benefit
→ only when contractually required
RESERVED + BURST
tenant A: 20 req/s reserved (guaranteed), may burst to 50 if capacity is free
✓ guarantees for those who need them, efficiency from sharing
✗ needs admission control that distinguishes reserved from burst traffic
→ THE USUAL ANSWER
PURE SHARED WITH WEIGHTS
no guarantees; each tenant gets weight/Σweights of capacity under contention
✓ maximum efficiency
✗ no guarantees; hard to sell to teams with SLAs
→ good for internal platforms with cooperative teams
CREDIT / SPEND
tenants have a budget; usage draws it down; no rate limit until exhausted
✓ aligns with billing
✗ doesn't prevent instantaneous contention
→ combine with one of the aboveReserved + burst is what most platforms converge on, because it lets you make credible guarantees while capturing most of the sharing benefit.
IMPLEMENTATION
Σ reserved ≤ 70% of capacity (leave room for burst and headroom)
reserved traffic: admitted always, highest priority
burst traffic: admitted if capacity available, shed first under load6. Accountability and transparency#
This is what makes teams accept sharing. If a team can’t see what they’re getting, they’ll demand dedicated capacity.
PER-TENANT DASHBOARD
requests, tokens (input/output), and cost — daily and monthly
latency percentiles for THEIR traffic
quota consumption and headroom
rejection rate and reasons
their share of cluster capacity
comparison to their reserved allocation
PER-TENANT COST (Section XI.03)
input_tokens × w_in + output_tokens × w_out + kv_block_seconds × w_kv
Publish it. Internally, showback (visibility) changes behavior even
without chargeback (actual billing).Showback alone typically reduces waste by 10-30%, because teams discover they’re spending more than they thought on something they didn’t need.
7. Noisy neighbours#
WAYS ONE TENANT DEGRADES ANOTHER
1. Consuming all KV cache with long-context requests
→ per-tenant KV block limits (Section XI.10)
2. Long prefills blocking decode
→ chunked prefill (mandatory), and route long prompts to a
separate pool (Section VIII.06)
3. Submitting continuously and winning FCFS
→ weighted fair queueing (section 4)
4. Triggering preemption by pushing memory to the limit
→ per-tenant memory limits, and headroom
5. Pathological requests (huge n, degenerate grammars)
→ validation limits (Section VIII.02)
6. Polluting the prefix cache with unique prefixes
→ per-tenant cache namespaces with per-tenant size limitsItem 6 is subtle: a tenant sending millions of unique prompts fills the prefix cache with never-reused entries, evicting other tenants’ useful cached prefixes. Per-tenant cache quotas fix it.
8. Onboarding a tenant#
CHECKLIST
□ tenant ID, auth credentials (scoped API keys)
□ tier assignment (free / standard / premium / reserved)
□ quotas: token rate, concurrency, KV blocks, requests/min
□ model access list (which models can they call?)
□ capability scoping (can they use tools? structured output? logprobs?)
□ data handling agreement (logging, retention, training use)
□ cost center for showback/chargeback
□ SLO tier and any contractual commitments
□ contact for incidents affecting them
□ dashboard access
CAPACITY IMPACT
□ does their projected demand fit in existing headroom?
□ if reserved: does Σ reserved still ≤ 70%?
□ do they need a model that isn't currently deployed?The “does their demand fit” check is the one that gets skipped, and it’s how a platform ends up over-subscribed.
9. Production implications#
- Default to shared model with per-tenant prefix cache namespaces.
- Implement weighted fair queueing. Quotas alone don’t ensure fairness.
- Reserved + burst is the quota model that works for platforms with mixed SLA needs.
- Cap Σ reserved at ~70% of capacity.
- Publish per-tenant cost and usage. Showback changes behavior.
- Per-tenant KV block limits, not just token rate limits.
- Per-tenant prefix cache quotas to prevent cache pollution.
- Have an onboarding checklist including the capacity check.
- Give each tenant visibility into their own latency, not just aggregate.
10. Common mistakes#
Quotas without fair queueing. A tenant within quota starves others.
Rate limiting tokens but not concurrency or KV. The real contention is memory.
Shared prefix cache across untrusted tenants. Timing side channel (XI.09).
No per-tenant cache quota. One tenant pollutes the cache.
No showback. Teams have no incentive to be efficient.
Over-committing reserved capacity. No room for burst or headroom.
Onboarding without a capacity check.
Only aggregate dashboards. Teams can’t see their own experience and lose trust.
11. Hands-on exercise#
A. Quantify the sharing benefit. For 5 synthetic tenants with realistic (non-coincident) traffic patterns, compute GPUs needed for dedicated deployments versus a shared platform. Include the batch-size effect.
B. Implement fair queueing. Build the FairScheduler from section 4. Test with tenants
submitting at very different rates. Verify each gets its weighted share and that idle tenants
don’t accumulate credit.
C. Noisy neighbour. Simulate a tenant sending long-context requests in a shared pool. Measure the impact on other tenants’ latency. Then add per-tenant KV limits and re-measure.
D. Cache pollution. Simulate one tenant sending unique prompts and another sending a shared prefix. Measure the second tenant’s cache hit rate with and without per-tenant cache quotas.
E. Build the tenant dashboard. Implement per-tenant metrics: requests, tokens, cost, latency percentiles, quota consumption. Would a team trust this enough to give up dedicated capacity?
12. Interview questions#
- What’s the economic case for a shared inference platform? Quantify it.
- Why do quotas alone not ensure fairness?
- Explain weighted fair queueing and why idle tenants must not accumulate credit.
- What’s the reserved + burst quota model and why is it common?
- Name six ways one tenant can degrade another, and the mitigation for each.
- Why namespace the prefix cache per tenant, and what does it cost?
- Why does showback reduce waste even without chargeback?
13. Further reading#
- [ESTABLISHED] Weighted fair queueing and deficit round robin literature
- [FUNDAMENTAL] Google’s Borg paper — multi-tenancy at scale
- [REFERENCE] Kubernetes ResourceQuota and LimitRange
- Next: 06 — Routing strategies