Below the API

Autoscaling in Production

Advanced Advanced 1h Difficulty 3/5

Prerequisites VIII.07, 02

Section VIII.07 covered the mechanics. This file covers the operational policy: what to actually configure, and what autoscaling can and cannot do for you.


1. The honest framing#

Autoscaling for LLM inference handles slow trends. It does not handle spikes. Plan accordingly.

Cold start 2-10 minutes  →  autoscaling reacts on a 5-15 minute timescale

WORKS FOR:                        DOESN'T WORK FOR:
daily traffic cycles              a product launch tweet
weekly patterns                   a 90-second burst
gradual growth                    a retry storm
predictable events (scheduled)    an upstream service recovering and
                                    replaying its queue

The mechanism that handles spikes is headroom (Section XI.02), which you pay for permanently. Autoscaling reduces how much headroom you need for the predictable variation, not for the unpredictable.

Diagram — The autoscaling loop#

flowchart LR
  M["Signals<br/>queue wait, running sequences,<br/>KV usage - NOT GPU utilization"] --> D{"Above target?"}
  D -->|"yes"| UP["Scale up now"]
  UP --> CS["Cold start<br/>schedule, pull image, load weights, warm up"]
  CS --> RDY["Ready: takes traffic"]
  D -->|"below, for a long window"| DN["Scale down slowly<br/>drain in-flight streams first"]
  RDY --> M
  DN --> M

  class M neutral
  class D queue
  class UP,RDY compute
  class CS warn
  class DN io

2. The policy#

# KEDA ScaledObject — better than raw HPA for custom metrics
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llama-70b
spec:
  scaleTargetRef:
    name: llama-70b-deployment
  minReplicaCount: 6            # baseline traffic, never below
  maxReplicaCount: 40
  pollingInterval: 15
  cooldownPeriod: 900           # 15 min before scaling down
  advanced:
    horizontalPodAutoscalerConfig:
      behavior:
        scaleUp:
          stabilizationWindowSeconds: 60
          policies:
            - type: Percent
              value: 50          # add up to 50% more replicas per minute
              periodSeconds: 60
            - type: Pods
              value: 4
              periodSeconds: 60
          selectPolicy: Max
        scaleDown:
          stabilizationWindowSeconds: 900     # 15 minutes of consistently
                                              # low load before removing
          policies:
            - type: Pods
              value: 1
              periodSeconds: 300              # at most 1 pod per 5 minutes
  triggers:
    - type: prometheus
      metadata:
        query: |
          avg(histogram_quantile(0.95,
            sum(rate(vllm_queue_wait_seconds_bucket[3m])) by (le)))
        threshold: "1.5"                      # scale up above 1.5 s queue wait
    - type: prometheus
      metadata:
        query: avg(vllm_gpu_cache_usage_perc)
        threshold: "0.75"
    - type: cron                              # predictive
      metadata:
        timezone: America/New_York
        start: 0 8 * * 1-5                    # weekday morning ramp
        end: 0 20 * * 1-5
        desiredReplicas: "14"

The four things that make this a production policy rather than a default:

  1. minReplicaCount: 6 — never scale to zero for a large model.
  2. Asymmetric windows — 60 s up, 900 s down.
  3. Queue wait as the primary trigger — it’s the SLO-relevant signal.
  4. A cron trigger — predictive scaling for known patterns costs nothing and works.

3. Why scale-down must be slow#

Scaling down is nearly free to delay and expensive to get wrong:

  scale down at t=0  →  traffic returns at t=120s
                     →  scale up at t=180s (after stabilization)
                     →  new replica ready at t=480s
                     →  6 minutes of degraded service

  vs. keeping the replica for 15 minutes: cost of 15 GPU-minutes.

At $2.50/GPU-hour × 8 GPUs × 0.25 hours = $5.
Six minutes of SLO violation across all users: worth far more than $5.

Thrashing is the failure mode. A too-eager scale-down policy plus normal traffic variance produces continuous scale up/down cycles, each paying a cold start.


4. The scaling signal, revisited#

SIGNAL                    QUALITY   NOTES
GPU utilization           USELESS   ~100% always (Section X.04)
CPU utilization           useless   the GPU is the constraint
requests in flight        poor      costs vary 1000x
queue depth               ok        length ≠ wait time
KV cache utilization      GOOD      the binding constraint
queue wait p95            BEST      directly measures the SLO
tokens/sec vs capacity    good      needs a capacity model
pending tokens            good      queue_depth × avg_tokens_per_request

Use queue wait as primary, KV utilization as secondary. Queue wait tells you about SLO violation; KV utilization tells you about impending capacity exhaustion. Together they cover both failure modes.


5. Predictive scaling — cheap and effective#

Most LLM traffic is highly predictable at the daily and weekly scale:

Typical consumer product:
  trough at 03:00-05:00 local     = 20% of peak
  ramp 07:00-10:00
  plateau 10:00-18:00             = 85-100% of peak
  peak 19:00-22:00                = 100%
  decline 22:00-02:00

Typical enterprise product:
  near-zero on weekends
  sharp 09:00 ramp on weekdays
  lunch dip
  hard stop at 18:00

A cron-based schedule beats reactive autoscaling for these patterns, because it scales before the load arrives rather than 8 minutes after.

Implementation:
  1. Extract the hourly profile from 4+ weeks of request logs
  2. Set minReplicas by hour of day and day of week
  3. Keep the reactive triggers on top, for the unpredictable part
  4. Review monthly as the profile drifts

This is the highest-value autoscaling work for most services and it takes an afternoon.


6. What to do when you can’t scale in time#

You will be over capacity sometimes. Decide in advance what happens.

DEGRADATION LADDER (apply in order as load exceeds capacity)

1. INCREASE BATCH SIZE
   Accept worse ITL for more throughput. Automatic in most engines
   (the scheduler fills the batch), but you can raise max_num_seqs.
   Cost: ITL degrades. Bounded, graceful.

2. REDUCE max_tokens for low-priority tiers
   A free-tier user gets 512 tokens instead of 2,048.
   Cost: truncated responses for some users.

3. SHED LOW-PRIORITY TRAFFIC (429 with Retry-After)
   Free tier first, then standard, never premium.
   Cost: some users rejected. HONEST and fast.

4. ROUTE TO A SMALLER MODEL
   Overflow to an 8B model instead of 70B.
   Cost: quality degradation for overflow traffic.

5. ROUTE TO AN EXTERNAL PROVIDER
   Expensive per token, but far cheaper than an outage.
   Cost: money, and data-residency implications.

6. QUEUE WITH A LONG DEADLINE (batch tier only)
   For workloads that tolerate it.

Steps 1-3 should be automatic. Steps 4-6 are business decisions to make in advance and implement as feature flags.

The alternative to a degradation ladder is a metastable collapse (Section VIII.05). Choose deliberately.


7. Scale-to-zero — when it’s viable#

VIABLE:
  small models (< 3B) with fast loading
  internal tools with tolerant users
  development environments
  models with genuinely bursty, infrequent use

NOT VIABLE:
  large models (cold start minutes)
  anything user-facing with a latency SLO
  models with expensive warmup (CUDA graphs, compile)

MIDDLE GROUND:
  scale to a small non-zero minimum (1-2 replicas)
  keep a "warm pool" of loaded-but-idle replicas that can be
  promoted instantly

The warm pool is the useful pattern for large models: pay for N idle replicas, get instant scale-up for N. It’s headroom with a different name, but it’s headroom you can share across models if the pool is generic.


8. Production implications#

  • Set minReplicas from baseline traffic, never zero for large models.
  • Asymmetric windows: fast up (60 s), slow down (900 s).
  • Queue wait as the primary signal.
  • Add predictive (cron) scaling for your daily and weekly patterns. Highest value per hour of work.
  • Implement the degradation ladder and test it.
  • Alarm on scaling events, especially thrashing (multiple up/down cycles in an hour).
  • Track “time spent over capacity” as an SLI — it’s what headroom buys you.
  • Optimize cold start (Section II.07). It determines what autoscaling can do at all.

9. Common mistakes#

Autoscaling on GPU utilization. Never varies.

Symmetric scale-up and scale-down windows. Thrashing.

Scale to zero for a large model. First request waits minutes.

Expecting autoscaling to handle spikes. It cannot.

No degradation plan. Collapse instead of graceful shedding.

No predictive scaling when the pattern is obviously predictable.

Not testing the scale-down path. Truncated generations during scale-down.

maxReplicas set without checking quota. The autoscaler asks for pods that can never be scheduled.


10. Hands-on exercise#

A. Extract the profile. From request logs, build the hourly traffic profile for 4 weeks. Plot it. How predictable is it? What fraction of variance is explained by hour-of-day and day-of-week?

B. Design the policy. Write the full KEDA/HPA configuration for a service you know, with predictive and reactive triggers, asymmetric windows, and justified thresholds.

C. Simulate. Build a simulator: traffic profile + cold start time + scaling policy. Measure the fraction of time over capacity, and the GPU-hours consumed. Compare policies: reactive only, predictive only, both.

D. The degradation ladder. Implement steps 1-3 for a real service. Load-test past capacity and verify each step engages in order.

E. Thrashing. Deliberately configure a symmetric scale-down window and generate variable load. Observe the thrashing. Fix it and quantify the cold-start cost you avoided.


11. Interview questions#

  1. Why can’t autoscaling handle traffic spikes for LLM services?
  2. What signal would you scale on, and why not GPU utilization?
  3. Why are scale-up and scale-down policies asymmetric?
  4. What is predictive scaling and when is it better than reactive?
  5. Describe a degradation ladder. What’s automatic and what’s a business decision?
  6. When is scale-to-zero viable?
  7. What is scaling thrash and how do you prevent it?

12. Further reading#

  • [REFERENCE] KEDA and Kubernetes HPA documentation
  • [FUNDAMENTAL] Google SRE Book, “Handling Overload”
  • [ESTABLISHED] AWS Builders’ Library on load shedding
  • Next: 05 — Fault tolerance and failure modes

↑↓ navigate ↵ open