1. Why multi-region, for LLM services specifically#
Three distinct reasons, with different implications:
1. LATENCY
The speed of light. A user in Singapore talking to a US-East endpoint
pays 200+ ms of RTT before any inference happens.
→ for a 800 ms TTFT SLO, that's 25% of the budget, spent on physics.
2. AVAILABILITY
Survive the loss of a region or an availability zone.
3. DATA RESIDENCY
Legal or contractual requirements that data not leave a jurisdiction.
→ often the actual driver, and it's non-negotiable.For LLM services, reason 1 is stronger than for most services, because TTFT is the metric users feel and network RTT is a fixed tax on it.
2. What makes it harder than for stateless services#
STATELESS WEB SERVICE LLM SERVICE
replicas are cheap and small each replica is 8 GPUs and 140 GB
scale up in seconds minutes
capacity is fungible GPUs are scarce and region-constrained
state is in a shared database KV cache is per-replica and ephemeral
failover = DNS change failover = DNS change + capacity in the
other region, which you must PAY FORThe cost of multi-region is the duplicated GPU capacity, and GPUs are the expensive part. A 2-region active-active deployment with full failover capacity costs roughly 2x.
3. The deployment patterns#
PATTERN COST FAILOVER LATENCY WHEN
Single region 1.0x none poor for small scale,
distant one market
users
Active-passive 1.5-2x minutes poor compliance-driven
(warm standby) to hours (until fo) DR requirement
Active-active, 2.0x seconds good the usual answer
each sized for full load
Active-active, 1.3-1.5x degraded good cost-conscious;
each sized for 65% accept degradation
during failover
Multi-region with varies seconds good large scale
overflow routing“Active-active, each sized for 65%” is the pragmatic middle ground: both regions serve traffic normally at 65% utilization; if one fails, the other runs at 130% of its comfortable load — degraded (higher latency, some shedding) but not down.
Document that degradation explicitly as part of your DR plan.
4. Routing#
LATENCY-BASED DNS / ANYCAST
Route users to the nearest healthy region.
✓ simple, works with standard infrastructure
✗ DNS TTL means failover takes minutes
✗ doesn't account for regional capacity
GLOBAL LOAD BALANCER (application layer)
✓ health-aware, capacity-aware, instant failover
✓ can implement overflow ("if EU is at 90%, send 10% to US")
✗ another component to run
HYBRID (the usual answer)
latency-based DNS for the primary route
+ application-layer overflow between regions
+ health checks that remove a region quicklyLLM-specific consideration: session affinity. A multi-turn conversation should stay in one region, both for prefix cache locality (Section V.11) and for consistency. Route by conversation ID within a region, and only move sessions on failover.
5. What has to be replicated#
MUST BE IN EVERY REGION
model weights (in regional object storage + node caches)
the serving stack (container images)
configuration
the model registry (or a regional read replica)
SHOULD BE REGIONAL, NOT REPLICATED
KV cache (ephemeral; never replicate)
in-flight request state
local caches
MUST BE GLOBALLY CONSISTENT (or carefully partitioned)
rate limit / quota state ← the hard one
billing records
auth tokens
audit logsQuota state across regions is genuinely hard. Options:
1. Partition quota by region (user's quota is 1/N per region)
✓ simple, no coordination
✗ a user hitting one region gets 1/N of their quota
2. Global quota store with eventual consistency
✓ correct enough; a user can briefly exceed by the sync window
✗ requires a cross-region store; adds latency
3. Local enforcement with periodic reconciliation
✓ fast, no cross-region latency in the request path
✗ over-consumption is detected after the fact
→ THE USUAL ANSWER: enforce locally with a generous local budget,
reconcile centrally, and correct on the next window.6. Model weight distribution#
Getting 140 GB of weights into every region:
1. Push to regional object storage (S3 cross-region replication,
or an explicit pipeline)
2. Pre-populate node-local caches (DaemonSet, or bake into an AMI/image)
3. Verify checksums per region
TIMING
A model release must complete in all regions before you shift traffic.
For a 140 GB model across 4 regions: hours, mostly transfer time.
→ build this into your release timeline.
COST
Cross-region egress is expensive. Push once to each region's storage
and pull locally, rather than pulling cross-region per node.A common and expensive mistake: nodes in region B pulling weights from region A’s object storage. Egress charges for 140 GB × 50 nodes × every deployment adds up quickly.
7. The DR plan#
Write it down, with numbers:
SCENARIO: loss of region us-east-1
DETECTION
health checks fail for > 60 s, or error rate > 50% for > 30 s
→ automated: global LB removes the region
IMPACT
45% of traffic must move to us-west-2 and eu-west-1
us-west-2 goes from 65% → 105% utilization → degraded
eu-west-1 goes from 60% → 88% utilization → acceptable
DEGRADATION (expected and accepted)
p95 TTFT: 780 ms → 1,900 ms for affected users
free-tier traffic shed: ~15%
premium tier: unaffected (reserved capacity)
RECOVERY
RTO (time to serve degraded): 90 seconds (DNS + LB health check)
RTO (time to serve normally): 45-90 minutes (spin up additional
capacity in the surviving regions, IF quota is available)
RPO: not applicable — inference is stateless; in-flight requests
are lost
ACTIONS
1. automated: LB removes the region
2. automated: degradation ladder engages (Section XI.04)
3. manual: request additional capacity in surviving regions
4. manual: communicate degraded status
5. on recovery: gradual traffic return (10% steps) to avoid
thundering herd on cold replicasStep 5 matters. Returning 45% of traffic instantly to a region whose replicas just started produces a cold-start stampede. Ramp it.
8. Cost management#
Multi-region is expensive. Ways to reduce the cost:
1. ASYMMETRIC SIZING
Size regions by their traffic, not equally.
2. SHARED OVERFLOW CAPACITY
One region has extra capacity that any region can overflow into,
rather than every region having its own N+1.
3. SPOT / PREEMPTIBLE for the overflow tier
Cheaper, and you only need it during failover.
4. TIER-AWARE FAILOVER
Only premium traffic fails over; free tier is shed.
→ dramatically reduces the capacity you must hold.
5. EXTERNAL PROVIDER AS THE DR TARGET
Instead of holding idle GPUs, contract with an API provider for
burst capacity. Expensive per token, cheap as insurance.Option 4 is the most under-used. Holding full failover capacity for free-tier traffic is usually not worth it; shedding it during a regional failure is an acceptable degradation.
9. Production implications#
- Decide the driver first: latency, availability, or compliance. They imply different architectures.
- Size for degraded operation, not full failover, unless you can afford 2x.
- Document the expected degradation numerically as part of the DR plan.
- Keep sessions regional for prefix cache locality.
- Enforce quota locally with central reconciliation.
- Distribute weights regionally; never pull cross-region per node.
- Test failover. A DR plan that has never been executed is a document, not a capability.
- Ramp traffic back gradually after recovery.
10. Common mistakes#
Multi-region for availability when latency is the actual need (or vice versa) — leading to the wrong architecture.
Sizing every region for 100% of total load. 3x cost for a 3-region deployment.
Cross-region weight pulls. Large egress bills.
Global quota checks in the request path. Adds cross-region latency to every request.
Sessions moving between regions. Cold prefix cache, inconsistent behavior.
Instant traffic return after recovery. Cold-start stampede.
Never testing failover.
Not documenting the expected degradation, so an incident becomes a surprise.
11. Hands-on exercise#
A. Compute the latency benefit. For your user distribution, compute the p50 and p95 network RTT to a single region versus to the nearest of N regions. What fraction of your TTFT budget does each save?
B. Size the regions. For a given traffic distribution and a “survive one region loss with degradation” requirement, compute the capacity needed in each region. Compare to full-failover sizing.
C. Write the DR plan. Produce the document from section 7 for a service you know, with real numbers.
D. Quota design. Design cross-region quota enforcement. Implement local enforcement with periodic reconciliation and measure the over-consumption window.
E. Test it. In a test environment, simulate a regional failure. Measure: detection time, traffic shift time, degradation experienced, and recovery time. Compare to your plan’s RTO.
12. Interview questions#
- What are the three reasons for multi-region, and how do they differ in architecture?
- Why is multi-region more expensive for LLM services than for stateless ones?
- How would you size regions to survive one region’s loss without 2x cost?
- How do you handle rate limit quota across regions?
- Why should sessions stay in one region?
- What’s your RTO and RPO for an inference service, and why is RPO unusual here?
- Why ramp traffic back gradually after a region recovers?
13. Further reading#
- [FUNDAMENTAL] Google SRE Book, chapters on reliability and disaster recovery
- [REFERENCE] Cloud provider multi-region architecture guidance
- [FUNDAMENTAL] Grigorik, High Performance Browser Networking — the latency argument
- Next: 07 — Model rollouts