1. The taxonomy#
Not all OOMs are the same. The first job is to classify.
TYPE SYMPTOM CAUSE
1. GPU OOM at startup fails during model load weights + pool don't fit
2. GPU OOM on first long fails on a long prompt activation spike
request
3. GPU OOM under load fails intermittently KV exhaustion or
fragmentation
4. GPU OOM after hours fails eventually leak or fragmentation
accumulation
5. HOST OOM (exit 137) no traceback, SIGKILL cgroup memory limit
6. "OOM" that isn't CUDA error, not memory a different error
reported confusinglyType 5 is the one that confuses people most, because there’s no Python traceback — the process is simply killed.
Diagram — When it OOMs tells you why#
flowchart TB
O["CUDA out of memory"] --> T{"When?"}
T -->|"at load"| L["Weights + context do not fit<br/>quantize, larger GPU, or TP"]
T -->|"on the first requests"| P["Peak activations, long-prompt prefill<br/>cap batched tokens, chunk prefill"]
T -->|"under load"| K["KV cache growth<br/>cap concurrency and context, preallocate a pool"]
T -->|"after hours"| F["Fragmentation or a leak<br/>reserved far above allocated"]
class O warn
class T queue
class L,K,F memory
class P compute2. Diagnosis, by type#
Type 1 — startup#
Compute what should fit (Section V.06):
weights + framework overhead + activation headroom + KV pool ≤ GPU memory
Check:
- Is the model the size you think? (check the checkpoint)
- Is the precision what you think? (a "quantized" model loaded as BF16?)
- Is gpu_memory_utilization too high?
- Is max_model_len too large (bigger block tables, bigger profiling batch)?
- Is another process using the GPU? (nvidia-smi --query-compute-apps)
- With TP: is the sharding correct? (each rank should hold 1/N)Type 2 — activation spike on a long prompt#
The engine's memory profiling pass sizes the KV pool based on a
worst-case batch it constructs. If your real worst case is larger,
you OOM later.
DIAGNOSE:
Does it fail at a specific prompt length? Bisect to find it.
Compare to max_num_batched_tokens — the profiling uses that.
FIX:
- lower gpu_memory_utilization (0.90 → 0.85)
- lower max_num_batched_tokens (smaller prefill chunks)
- enable chunked prefill (bounds the activation size)
- cap max input length at the gatewayThis is the most common production OOM and it’s usually fixed by chunked prefill, which
bounds the prefill activation tensor to max_num_batched_tokens regardless of prompt length.
Type 3 — KV exhaustion under load#
Not really an OOM if the engine handles it — it preempts (Section V.09).
If it OOMs instead, the pool sizing or accounting is wrong.
DIAGNOSE:
KV utilization at the time of failure (should hit 100% and preempt)
preemption counter (should be nonzero before any OOM)
max_num_seqs vs what memory allows
FIX:
- reduce max_num_seqs (leave 20-30% headroom)
- reduce max_model_len
- quantize the KV cache
- add capacityType 4 — accumulation over hours#
DIAGNOSE:
Plot torch.cuda.memory_reserved() and memory_allocated() over time.
allocated grows → a genuine leak (tensors retained)
reserved-allocated grows → fragmentation
both flat, still OOM → something outside PyTorch (NCCL buffers,
cuBLAS workspaces, a library)
COMMON LEAKS:
- a request object retained after completion (holding its tensors)
- a growing cache without eviction
- exception paths that skip cleanup
- CUDA graphs captured per request instead of once
- missing torch.inference_mode() (Section I.02)Type 5 — host OOM#
SYMPTOM: process disappears, exit code 137, no Python traceback.
CONFIRM:
dmesg -T | grep -i "killed process"
cat /sys/fs/cgroup/memory.events # oom_kill counter
kubectl describe pod ... # OOMKilled reason
CAUSES:
- loading the model into host memory before moving to GPU
(a 140 GB model needs 140 GB of host RAM unless you stream shard by shard)
- page cache from mmap'd weights counting against the cgroup limit
- pinned buffer pools
- the tokenizer's memory on very long prompts
- a leak in the API layer
FIX:
- raise the container memory limit
- stream weights shard by shard to GPU
- bound pinned buffer pools
- check that page cache is being reclaimed (it should be, but verify)3. The prevention checklist#
[ ] Compute the memory budget explicitly before deploying (Section V.06)
[ ] Set gpu_memory_utilization with headroom (0.88-0.92)
[ ] Set max_model_len from traffic, not the model maximum
[ ] Set max_num_seqs to 70-80% of what memory allows
[ ] Enable chunked prefill (bounds activation spikes)
[ ] Cap max input length at the gateway
[ ] Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
[ ] Test with the longest prompt you'll accept, at max batch
[ ] Set the container memory limit with headroom for host-side buffers
[ ] Monitor: KV utilization, memory_reserved, memory_allocated, preemptions
[ ] Alert on preemption rate > 1% (precursor to OOM)
[ ] Have a runbook entry for exit code 137The “test with the longest prompt at max batch” line catches most type-2 OOMs before production. It is routinely skipped.
4. Worked example#
INCIDENT: server OOMs 3-4 times per day, always around peak traffic.
STEP 1 — CLASSIFY
Not at startup (type 1 ruled out).
Not gradual (memory_reserved is flat between incidents — type 4 ruled out).
Python traceback present, "CUDA out of memory" (type 5 ruled out).
→ Type 2 or 3.
STEP 2 — CORRELATE
Every OOM coincides with a request whose prompt > 28,000 tokens.
→ Type 2: activation spike.
STEP 3 — CONFIRM
Reproduce: send a 30,000-token prompt while 40 sequences are decoding.
OOM reproduced.
Instrument: peak memory during that prefill is 11.2 GB of activations,
versus 2.1 GB during the profiling pass (which used
max_num_batched_tokens = 2048 but no chunked prefill, so it
processed the whole 30k prompt at once).
STEP 4 — FIX
Enable chunked prefill. Now the prefill is processed in 2,048-token
chunks; peak activation is bounded at 2.1 GB regardless of prompt length.
STEP 5 — ALSO
Cap max input at 32,000 tokens at the gateway with a clear 400 error.
Lower gpu_memory_utilization from 0.94 to 0.90 for headroom.
RESULT: no OOMs in 30 days. Throughput unchanged. p99 ITL improved 40%
(chunked prefill also fixed the head-of-line blocking).Note the bonus in step 5’s result: the fix for the OOM also fixed a latency problem, because both had the same root cause. That’s common — long unchunked prefills cause several symptoms.
5. What to do when it happens in production#
IMMEDIATE
1. Is it crashlooping? If so, reduce load or lower max_num_seqs and restart.
2. Capture: memory_summary(), the failing request's shape, KV utilization
at the time, recent preemption counts.
3. If it's a specific request shape, block that shape at the gateway
temporarily.
SHORT TERM
4. Lower gpu_memory_utilization by 0.03-0.05.
5. Lower max_num_seqs by 20%.
6. Enable chunked prefill if not already on.
7. Cap input length.
ROOT CAUSE
8. Classify (section 1), diagnose (section 2), fix properly.
9. Add a test that would have caught it.
10. Add monitoring for the precursor (KV utilization, preemption rate).Step 9 matters. An OOM that recurs is a testing failure, not just an operational one.
6. Production implications#
- Preemption is the healthy version of running out of KV memory. OOM means the engine failed to preempt in time or the accounting was wrong. Monitor preemption as a leading indicator.
- Headroom is not waste. 10% unused GPU memory is cheap insurance against a class of incidents.
- Bound every input. Prompt length, output length, batch size,
n. Each is a memory control. - Test the extremes in CI. Longest prompt × max batch × longest generation.
- Exit code 137 belongs in your runbook with the
dmesgcommand to confirm. - A crashlooping OOM is worse than a slow service. Prefer rejecting requests to crashing.
7. Common mistakes#
Raising gpu_memory_utilization to squeeze out capacity. Trades a small gain for incidents.
Not testing with the longest realistic prompt.
Setting max_num_seqs to the memory maximum. No headroom for variability.
Confusing host OOM with GPU OOM. Different fixes entirely.
Not reading memory_summary(). It usually names the cause.
No input length cap. One request can OOM the server.
Treating OOM as random. It isn’t; it correlates with something. Find the correlation.
8. Hands-on exercise#
A. Induce each type. Deliberately cause a type-1, type-2, and type-5 OOM on a test system. Observe the different symptoms and error messages. Record what each looks like.
B. Find the activation cliff. Bisect to find the prompt length at which your configuration OOMs at a given batch size. Compare to the theoretical activation memory. Then enable chunked prefill and confirm the cliff disappears.
C. Build the memory budget. For your deployment, write out every consumer with its size: weights, KV pool, CUDA graphs, activations (peak), framework overhead, NCCL buffers. Sum it. Compare to actual GPU memory. How much headroom do you have?
D. Leak detection. Run a server for an hour under load, sampling memory_allocated() and
memory_reserved() every 10 seconds. Plot both. Is either growing?
E. The runbook. Write the OOM runbook entry for your service: symptoms, classification steps, immediate mitigations, and diagnosis procedure.
9. Interview questions#
- Name the types of OOM in an inference service and how you’d distinguish them.
- Your server OOMs only on long prompts. What’s happening and how do you fix it?
- What does exit code 137 mean and how do you confirm it?
- Why is preemption the healthy version of KV exhaustion?
- Your process OOMs after 6 hours but not at startup. Diagnose.
- What would you cap at the gateway, and why is each a memory control?
- Walk through your immediate response to an OOM crashloop in production.
10. Further reading#
- [REFERENCE] PyTorch CUDA memory documentation and
memory_summary() - [REFERENCE] Linux cgroup v2 memory documentation
- [ESTABLISHED] vLLM’s memory profiling implementation (how it sizes the KV pool)
- Next: 07 — Benchmarking inference correctly