1. The failure modes, catalogued#
FAILURE FREQUENCY BLAST RADIUS DETECTION
GPU Xid error / fault weeks 1 GPU → 1 TP group dmesg, DCGM
GPU falls off the bus months 1 node nvidia-smi fails
ECC uncorrectable error months 1 GPU DCGM, dmesg
Driver hang months 1 node all CUDA calls hang
NCCL hang (one slow rank) weeks 1 TP/PP group timeout
OOM days-weeks 1 replica exception or SIGKILL
Model load failure per deploy 1 replica readiness never passes
Network partition (multi-node) months 1 instance NCCL error
Node preemption (spot) hours 1 node cloud signal
Deployment error per deploy everything canary should catch
Quality regression per deploy everything hard (Section VIII.10)
Upstream dependency failure varies varies circuit breaker
Traffic spike beyond capacity weekly everything queue waitThe two that are LLM-specific and most damaging: NCCL hangs and quality regressions. Both are hard to detect and both hold resources while broken.
2. The design principles#
1. FAIL FAST, NOT SLOW.
A hung process holding 8 GPUs is worse than a crashed one.
→ timeouts everywhere, especially NCCL.
2. LIMIT THE BLAST RADIUS.
DP replicas fail independently; TP groups fail together.
→ prefer more, smaller instances where latency permits.
3. DEGRADE, DON'T COLLAPSE.
Shed load explicitly rather than queueing until failure.
4. MAKE FAILURE VISIBLE.
A silently degraded replica serving slowly is worse than a dead one,
because the load balancer keeps sending it traffic.
5. RECOVER AUTOMATICALLY, BUT BOUNDED.
Restart, but with backoff and a crashloop limit.
6. PRESERVE WHAT YOU CAN.
In-flight requests should fail cleanly with a clear error,
not hang until the client times out.3. GPU failures#
DETECTION
dmesg -T | grep -i xid
nvidia-smi --query-gpu=ecc.errors.uncorrected.aggregate.total --format=csv
DCGM_FI_DEV_XID_ERRORS
COMMON XID CODES
13, 31 illegal memory access (usually a software bug, not hardware)
43 stopped processing (software)
48 double-bit ECC error (HARDWARE — replace)
63, 64 ECC page retirement (hardware degrading)
74 NVLink error (hardware or cable)
79 GPU has fallen off the bus (HARDWARE — node is unusable)
94, 95 contained/uncontained ECC error
RESPONSE
software Xids (13, 31, 43) → restart the process; investigate the bug
hardware Xids (48, 79, 94) → CORDON THE NODE, drain, replace the GPU
degrading (63, 64) → schedule replacement; monitorAutomate the cordon. A node with an uncorrectable ECC error will keep failing; leaving it in the pool produces repeated incidents.
# NVIDIA GPU Operator's node-problem-detector can taint nodes automatically
# Or: a DaemonSet that watches DCGM_FI_DEV_XID_ERRORS and applies a taint4. NCCL hangs — the LLM-specific one#
SYMPTOM
A TP or PP group stops making progress. No error. No crash.
All ranks are in a collective, waiting for one that never arrives.
8 or 16 GPUs are held indefinitely.
CAUSES
one rank crashed (the others wait forever)
one rank is descheduled or throttled (Section II.12)
a network partition
a driver issue on one node
a deadlock from mismatched collective ordering (a code bug)
DETECTION
NCCL_TIMEOUT (set it! default is very long)
TORCH_NCCL_ASYNC_ERROR_HANDLING=1
a watchdog: if no step completed in N seconds, kill the process
externally: queue depth growing while throughput is zero
RESPONSE
1. Kill ALL ranks of the group (NCCL has no partial recovery)
2. Restart the group
3. If it recurs on the same node, cordon that node// StepWatchdog is an application-level watchdog — necessary because NCCL timeouts
// don't always fire, and a hang is much worse than a crash.
type StepWatchdog struct {
last atomic.Int64 // unix nanoseconds of the last engine step
timeout time.Duration
}
func NewStepWatchdog(timeout time.Duration) *StepWatchdog {
w := &StepWatchdog{timeout: timeout}
w.Beat()
go func() {
for range time.Tick(10 * time.Second) {
if idle := time.Since(time.Unix(0, w.last.Load())); idle > w.timeout {
slog.Error("watchdog: no engine step, aborting", "idle", idle)
os.Exit(1) // hard exit; let the supervisor restart us
}
}
}()
return w
}
// Beat is called by the engine loop after every step.
func (w *StepWatchdog) Beat() { w.last.Store(time.Now().UnixNano()) }os._exit(1) rather than a graceful shutdown is correct here: a hung process may not be
able to shut down gracefully, and the supervisor’s restart is the recovery path.
5. Handling in-flight requests during failure#
WHEN A REPLICA DIES
its in-flight requests are lost. The clients see:
- a connection reset (non-streaming)
- a truncated stream (streaming) ← worse; the client got partial output
MITIGATIONS
1. Fail cleanly: send an in-band error before dying, if possible
data: {"error": {"message": "...", "type": "server_error"}}
2. Client-side: detect truncation (no [DONE] received) and surface it
3. Do NOT auto-retry a streamed request — the user has partial output
4. For non-streaming, a retry is safe IF the original is confirmed deadThere is no way to migrate an in-flight generation to another replica without transferring its KV cache — which is possible in principle (Section IX.11) but not implemented in production systems for failure recovery. Accept the loss and handle it cleanly.
6. Circuit breakers and dependency failures#
An inference gateway typically depends on:
auth service
rate limit store (Redis)
model registry
the engines themselves
(sometimes) a safety classifier, a retrieval service
FOR EACH: what happens if it's down?
auth down → fail closed (reject) or fail open (allow)?
SECURITY DECISION. Usually fail closed, with a
short-lived cache to ride out brief outages.
rate limiter → fail OPEN (allow), with a conservative local limit.
Rejecting all traffic because Redis is down is worse
than briefly allowing over-limit traffic.
registry → cache aggressively; it changes rarely.
safety filter → fail closed for high-risk content, or degrade to a
cheaper local check. A policy decision.
an engine → circuit-break it, route elsewhere.var ErrCircuitOpen = errors.New("circuit open")
type CircuitBreaker struct {
mu sync.Mutex
failures int
threshold int // consecutive failures before opening, e.g. 5
cooldown time.Duration // how long to stay open, e.g. 30s
openedAt time.Time
}
func (cb *CircuitBreaker) Call(fn func() error) error {
cb.mu.Lock()
if !cb.openedAt.IsZero() && time.Since(cb.openedAt) < cb.cooldown {
cb.mu.Unlock()
return ErrCircuitOpen // fail fast: do not even try the unhealthy replica
}
cb.mu.Unlock()
err := fn()
cb.mu.Lock()
defer cb.mu.Unlock()
if err == nil {
cb.failures, cb.openedAt = 0, time.Time{}
return nil
}
if cb.failures++; cb.failures >= cb.threshold {
cb.openedAt = time.Now()
}
return err
}Write down the fail-open/fail-closed decision for every dependency. It is a decision, and if you don’t make it deliberately the code makes it accidentally.
7. Testing failure#
GAME DAYS — practice these deliberately:
1. Kill one replica under load.
Expect: LB removes it, in-flight requests fail cleanly, no cascade.
2. Kill one rank of a TP group.
Expect: the group hangs, the watchdog fires within N seconds,
the supervisor restarts all ranks, readiness gates traffic.
3. Fill the KV cache (send max-length requests).
Expect: preemption, then admission rejection with 503. No OOM.
4. Saturate the network between nodes.
Expect: NCCL slows, the watchdog eventually fires, or throughput
degrades gracefully.
5. Stop the rate-limit store.
Expect: fail-open with a local fallback limit, alarm fires.
6. Deploy a deliberately broken model version.
Expect: readiness never passes, the rollout halts, no traffic shifted.
7. Exceed capacity by 3x.
Expect: 503s with Retry-After, served requests keep their SLO,
recovery when load drops (no metastable state).Most teams have never tested #2 or #7, and both are the ones that cause real incidents.
8. Production implications#
- Set
NCCL_TIMEOUTand an application watchdog. Hangs are worse than crashes. - Automate node cordoning on hardware Xid errors.
- Prefer more, smaller instances where latency permits — smaller blast radius.
- Write down fail-open/fail-closed for every dependency.
- Handle in-flight request loss cleanly. Don’t leave clients hanging.
- Run game days. Quarterly, on the seven scenarios above.
- Track MTTR per failure mode, not just availability.
- Crashloop limits: a replica that has restarted 5 times in 10 minutes should stay down and alert, not keep cycling.
9. Common mistakes#
No NCCL timeout. A hang holds 8-16 GPUs indefinitely.
Restarting one rank of a TP group. The others still hang.
No watchdog. Relying on NCCL’s timeout alone is insufficient.
Leaving failed nodes in the pool. Repeated incidents from the same hardware.
Auto-retrying streamed requests. Duplicate partial output.
Fail-closed on the rate limiter. Total outage because Redis blipped.
Never testing failure. The first time you exercise the path is during an incident.
Unbounded crashloops. A replica cycling forever, consuming GPUs and generating noise.
10. Hands-on exercise#
A. Implement the watchdog. Add the step watchdog from section 4 to an engine. Test it by
suspending a rank (kill -STOP) and verifying the watchdog fires.
B. The Xid response. Write a DaemonSet or script that watches for hardware Xid errors and cordons the node. Test with a simulated error.
C. Dependency matrix. For a service you know, list every dependency and write the fail-open/fail-closed decision with justification. Implement circuit breakers.
D. Game day. Run all seven scenarios from section 7 on a test environment. Document what actually happened versus what you expected. Fix the gaps.
E. Clean failure. Implement in-band error reporting for a streaming endpoint. Kill the engine mid-generation and verify the client receives a structured error rather than a hang.
F. Blast radius. Compare the impact of a single GPU failure in a TP=8 deployment versus 8 independent replicas. Quantify in requests affected.
11. Interview questions#
- What happens when one rank of a TP group fails? How do you handle it?
- Why is a hang worse than a crash for a GPU workload?
- What is a hardware Xid error and how should you respond?
- Should the rate limiter fail open or closed? Justify.
- Can you migrate an in-flight generation to another replica? Why or why not?
- Design a game day for an LLM inference service.
- How does parallelism strategy affect blast radius?
12. Further reading#
- [REFERENCE] NVIDIA Xid error documentation
- [REFERENCE] NVIDIA GPU Operator and node-problem-detector
- [FUNDAMENTAL] Google SRE Book, “Addressing Cascading Failures”
- [FUNDAMENTAL] Netflix’s chaos engineering principles
- Next: 06 — Multi-region and disaster recovery