1. Problem → Why → Optimization#
PROBLEM Standard autoscaling doesn't work for LLM serving: the scaling
signal is wrong, and the reaction time is 10-100x too slow.
WHY GPU utilization is a useless signal, and cold start is minutes
because you must load tens of GB of weights.
OPTIMIZE Scale on KV utilization / queue wait; make cold start fast;
and pre-provision, because reactive scaling cannot catch spikes.2. Why the standard approach fails#
STANDARD WEB SERVICE LLM SERVICE
scale on CPU% (meaningful) GPU% is ~100% whenever anything runs (useless)
scale out in 10-30 s 2-12 minutes
requests are stateless in-flight requests hold KV state
drain in seconds drain takes as long as the longest generationBoth halves of autoscaling are broken: the signal and the reaction time. Fixing the signal is easy; fixing the reaction time is a capacity-planning problem, not an engineering one.
3. Simple analogy#
Hiring versus opening more checkout lanes.
A supermarket opens another lane in 30 seconds. If the queue grows, you open a lane.
An airline cannot add a plane in 30 seconds. Aircraft are provisioned weeks ahead based on forecast demand, with spare capacity for irregular operations. When demand exceeds capacity, you don’t scale — you queue, reroute, or turn people away.
LLM serving is the airline. Plan capacity; use autoscaling for slow trends, not spikes.
4. The scaling signal#
SIGNAL QUALITY WHY
GPU utilization USELESS ~100% whenever a kernel runs
CPU utilization useless the GPU is the constraint
requests in flight poor costs vary 1000x
queue depth ok but length ≠ wait time
KV cache utilization GOOD directly measures the binding constraint
queue wait time p95 BEST directly measures the SLO you care about
tokens/sec vs capacity good requires knowing capacityScale on p95 queue wait time, with KV utilization as a secondary signal.
# Kubernetes HPA with custom metrics (via prometheus-adapter / KEDA)
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
minReplicas: 4 # never scale to zero for a large model
maxReplicas: 40
metrics:
- type: Pods
pods:
metric: {name: vllm_queue_wait_seconds_p95}
target: {type: AverageValue, averageValue: "2"}
- type: Pods
pods:
metric: {name: vllm_gpu_cache_usage_perc}
target: {type: AverageValue, averageValue: "70"}
behavior:
scaleUp:
stabilizationWindowSeconds: 60 # react reasonably fast
policies: [{type: Percent, value: 100, periodSeconds: 60}]
scaleDown:
stabilizationWindowSeconds: 900 # scale down SLOWLY — 15 min
policies: [{type: Pods, value: 1, periodSeconds: 300}]
Note the asymmetry: fast up, slow down. Scaling down is nearly free to delay and expensive to get wrong (you pay a 5-minute cold start to undo it). Fifteen-minute stabilization on scale-down is conservative and correct.
5. Cold start — the real problem#
Phase Typical Optimized
Pod scheduling 5-30 s 5 s (pre-provisioned nodes)
Image pull (if not cached) 30-180 s 0 s (pre-pulled / small image)
Container start + Python import 10-30 s 10 s
Weight load from object storage 60-600 s 20-60 s (local NVMe cache)
Weight load to GPU 10-60 s 7-20 s (parallel, pinned)
Quantization (if at load time) 10-120 s 0 s (pre-quantized)
torch.compile (if uncached) 60-600 s 0 s (cached artifacts)
CUDA graph capture 20-60 s 20-60 s
Warmup requests 5-20 s 5-20 s
─────────────────────────────────────────────────────────────
TOTAL 3-25 min 1-3 minEvery line of the “optimized” column is achievable and covered in Section II.07. Getting from 15 minutes to 2 minutes changes what autoscaling can do.
The techniques, ranked#
1. Cache weights on local NVMe (DaemonSet or init container populates a hostPath)
→ 600 s becomes 30 s. The biggest single win.
2. Ship pre-quantized checkpoints. → removes 10-120 s
3. Persist compile/engine artifacts. → removes 60-600 s
4. Pre-pull images; keep them small (weights NOT in the image). → removes 30-180 s
5. Parallel shard loading with pinned buffers. → 2-4x on the load itself
6. Pre-provisioned warm node pool. → removes scheduling and pull entirely
7. Reduce CUDA graph capture sizes if startup matters more than the last 5%.What autoscaling cannot do#
Spike duration 2 min, cold start 3 min:
→ autoscaling contributes NOTHING. The spike is over before capacity arrives.
Spike duration 30 min, cold start 3 min:
→ autoscaling handles the last 27 minutes. The first 3 are on your headroom.
Daily cycle (traffic rises over 2 hours):
→ autoscaling works well. This is its natural use case.Therefore: autoscaling handles daily and weekly patterns. Headroom handles spikes. You
must size for peak_predictable + spike_headroom, and the headroom is real, permanent cost.
6. The strategies that actually work#
1. PRE-PROVISIONED HEADROOM
Run at 60-70% utilization. The spare 30-40% absorbs spikes.
Cost: 40-60% more GPUs than peak-average would suggest.
This is the baseline strategy and it is not optional.
2. PREDICTIVE SCALING
Scale ahead of known patterns (time of day, day of week, product launches).
Kubernetes: a CronJob adjusting minReplicas, or KEDA cron triggers.
Cheap and effective for predictable cycles.
3. WARM POOL
Keep N replicas loaded but not receiving traffic; promote them instantly.
Cost: N × GPU cost. Benefit: zero-latency scale-up for N replicas.
4. GRACEFUL DEGRADATION
When over capacity: shed load (503), reduce max_tokens, route to a
smaller model, or increase batch size (accepting worse ITL).
This is what you do INSTEAD of scaling, in the 3 minutes it takes.
5. MULTI-TIER FALLBACK
Overflow to a different model, a different region, or an external API
provider. Expensive per token, but far cheaper than an outage.Strategies 1 and 4 are mandatory. 2 is cheap and worth doing. 3 and 5 depend on your economics.
7. Scale-down and draining#
Scaling down an LLM replica is harder than scaling up:
1. Remove from the LB pool (stop new requests)
2. Wait for in-flight generations to finish
← could be 10 minutes for a long generation
3. Terminate
Kubernetes: preStop hook + terminationGracePeriodSecondslifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "curl -X POST localhost:8000/drain && sleep 5"]
terminationGracePeriodSeconds: 900 # 15 minutes
Where /drain marks readiness false and stops admitting new requests.
With a short grace period, you kill in-flight generations — visible to users as truncated responses. Set it to at least your p99 generation time.
Alternative: a deadline-based drain — refuse new requests, and after a bounded time abort whatever remains with a clear error, so a stuck request doesn’t block the rollout forever.
8. Production implications#
- Never scale to zero for large models unless you accept minutes of cold start on the first request. For small models on shared GPUs it can be viable.
- Set
minReplicasfrom your baseline traffic, not from zero. - Fast up, slow down. Asymmetric stabilization windows.
- Measure and alarm on cold-start duration. It regresses silently when someone adds a step to the startup path.
- Test the drain path. Roll a deployment under load and verify no generations are truncated.
- Budget headroom explicitly in your capacity model (Section XI.02) — typically 30-40%.
- Have a documented degradation plan for when you’re over capacity.
9. Common mistakes#
Autoscaling on GPU utilization. The classic. It’s always ~100%.
Expecting autoscaling to handle spikes. It cannot; the physics don’t allow it.
Scaling to zero for a large model. First request waits 5 minutes.
Short terminationGracePeriodSeconds. Truncated generations during every deploy.
Aggressive scale-down. Thrashing: scale down, traffic returns, pay a cold start.
Not optimizing cold start. It determines what autoscaling can do at all.
No degradation plan. When you’re over capacity and can’t scale, what happens? Decide in advance.
10. Hands-on exercise#
A. Measure cold start, phase by phase. Instrument a real model server’s startup. Produce a bar chart of the phases from section 5. Which dominates?
B. Optimize it. Apply the top three techniques for your setup (local weight cache, pre-quantized checkpoint, cached compile artifacts). Measure the improvement. How close to the “optimized” column do you get?
C. Simulate spike response. Build a simulator: traffic spike of duration D, cold start C, headroom H. Compute the fraction of requests that fail or exceed SLO. Sweep D, C, H. Plot the headroom needed for a given spike shape.
D. Test the drain. Roll a deployment while streaming requests are in flight. Are any truncated? Fix the grace period and repeat.
E. Build the scaling policy. Write the HPA/KEDA configuration for your service, with the right signal and asymmetric windows. Load-test to verify it scales when you expect.
11. Interview questions#
- Why is GPU utilization a bad autoscaling signal? What would you use instead?
- Break down cold start for a 70B model. Which phase dominates and how do you fix it?
- Why can’t autoscaling handle traffic spikes for LLM services?
- Why are scale-up and scale-down policies asymmetric?
- What happens if
terminationGracePeriodSecondsis too short? - What’s your plan when you’re over capacity and cannot scale in time?
- When is scaling to zero acceptable?
12. Further reading#
- [REFERENCE] Kubernetes HPA and KEDA documentation
- [REFERENCE] vLLM production metrics reference
- [FUNDAMENTAL] Google SRE Book, capacity planning chapters
- Next: 08 — Model lifecycle