Section VI.07 covered CUDA graphs mechanically. This file covers the serving-specific decisions: capture lists, memory budgets, and feature interactions.
1. Problem → Why → Optimization#
PROBLEM At batch 1-32, 30-50% of decode step time is CPU launch overhead.
WHY ~350-2000 kernels per token, ~5 µs each, while each kernel does
only microseconds of work.
OPTIMIZE Capture the decode step as a graph; replay with one launch.
TRADE-OFFS
✓ 10-45% end-to-end, largest at small batch
✗ 20-60 s startup capture time
✗ 1-3 GB of GPU memory (= KV cache you don't have)
✗ static shapes → batch padding waste
✗ conflicts with some dynamic features
WHEN TO USE Always for decode in production.
WHEN NOT TO Debugging (obscures stack traces); prefill (shapes vary too much);
when an incompatible feature is more valuable.2. The serving-specific problem#
Decode shapes are almost static:
Fixed: model architecture, hidden size, number of layers
the KV cache pool's base address
the block table buffer's address
Variable: batch size (changes every step as requests join and leave)
context length per sequence (grows every step)
sampling parameters per requestTwo of those three are handled elegantly:
- Context length is data (in the block table and seq_lens buffers), not shape. The paged attention kernel loops based on it. No recapture needed.
- Sampling parameters live in fixed-size per-slot buffers.
Only batch size genuinely changes the shapes. That’s why engines capture a list of batch sizes.
3. Capture list design#
vLLM default (approximately):
[1, 2, 4, 8] + [16, 24, 32, 40, ..., 256] (step 8 above 8)
Trade-off:
more sizes → less padding waste, more memory and startup time
fewer sizes → less memory, more padding wasteComputing the padding cost#
Given a batch size distribution P(b) and a capture list C:
expected_waste = Σ_b P(b) × (next_capture(b) − b) / b × KV_fraction
where KV_fraction is the share of step bytes that scale with batch
(i.e. KV reads, not weights).Worked example:
70B, context 4k, batch distribution centered at 40:
weights: 17.6 GB/GPU; KV: 40 KiB/token/GPU × 4096 = 164 MB/sequence
batch 40, capture list has 40 → no padding.
batch 41, next capture is 48 → 7 extra sequences × 164 MB = 1.15 GB extra
of 17.6 + 40×0.164 = 24.2 GB → 4.7% waste on that step.
Average over the distribution: typically 3-8% with a good capture list.
Compare to the 25-40% launch overhead saved. Clear win.Tuning the list#
1. Measure your batch size distribution over a day of real traffic.
2. Place capture points densely where the distribution is dense.
3. Verify total graph memory stays within your budget.
Example: if 80% of steps are batch 20-60, capture
[1,2,4,8,12,16,20,24,28,32,36,40,44,48,56,64,96,128,192,256]
rather than powers of two.Most engines expose this (--cuda-graph-sizes or equivalent). Few teams tune it; it’s worth
30 minutes.
4. Memory budget#
Per captured graph:
- static input/output buffers
- intermediate activation buffers (from the graph memory pool)
- the graph structure
For a 70B model with TP=8, per GPU:
~50-150 MB per captured batch size
× 20 captured sizes = 1-3 GBThat’s 1-3 GB of KV cache you don’t have — roughly 6-18 fewer concurrent sequences at 4k context. Factor it into capacity planning; the engine’s reported block count already accounts for it, but your own estimates should too.
If memory is tight, reduce the capture list rather than disabling graphs entirely.
5. Feature interactions#
This is where production surprises live:
FEATURE GRAPH COMPATIBILITY
Paged attention ✓ (context length is data, not shape)
Continuous batching ✓ (with a capture list per batch size)
Prefix caching ✓ (block tables are data)
Chunked prefill ✗ for the prefill part; ✓ for the decode part
Multi-LoRA ⚠ requires careful handling; some engines disable graphs
Speculative decoding ⚠ variable accepted-token counts complicate capture
Structured/guided decoding ⚠ the mask computation is often outside the graph
Custom logits processors ✗ if implemented in Python
Sampling with per-request ✓ if parameters live in fixed buffers
parameters
Tensor parallelism ✓ (NCCL ops are capturable)
Pipeline parallelism ⚠ cross-stage communication complicates itThe dangerous case is silent disabling. You enable multi-LoRA, the engine quietly falls back to eager for graph-incompatible paths, and you lose 25% with no error. Check the logs after enabling any new feature, and re-benchmark.
6. Verification#
# vLLM logs at startup:
# "Capturing CUDA graphs (decode, ...): 100%|...| 35/35"
# "Graph capturing finished in 24 secs, took 1.42 GiB"
# If you see:
# "CUDA graphs are not supported with <feature>, falling back to eager"
# you have lost the optimization.Also verify by profiling: an Nsight Systems timeline of a graph-replayed decode step shows one
cudaGraphLaunch and then back-to-back kernels with no gaps. Eager shows hundreds of
cudaLaunchKernel calls interleaved with gaps.
7. Performance#
Model / batch Speedup from graphs
7B, batch 1 1.45-1.60x
7B, batch 32 1.15-1.25x
70B (TP=8), batch 1 1.30-1.40x
70B (TP=8), batch 32 1.15-1.22x
70B (TP=8), batch 128 1.06-1.12x
Prefill ~1.00xTwo patterns: smaller models benefit more (more kernels per unit of useful work), and smaller batches benefit more (less work per kernel).
For a latency-critical service running at low batch, this is one of the largest single-config wins available.
8-9. Production and mistakes#
Production:
- Leave graphs enabled. Default in vLLM and TensorRT-LLM.
- Budget the startup time. 20-60 s adds to your cold start (Section II.07). If cold start is critical, reduce the capture list.
- Budget the memory. 1-3 GB per GPU.
- Tune the capture list to your measured batch distribution.
- Re-benchmark after enabling any new feature, in case graphs were silently disabled.
- Use
--enforce-eageronly for debugging, and make sure it’s not in your production config. This has happened to many teams. - Watch the logs for capture failures.
Mistakes:
--enforce-eagerleft in production. Free 25% loss.- Not noticing silent fallback after enabling a feature.
- Capture list mismatched to the traffic distribution.
- Forgetting the memory cost in capacity planning.
- Expecting graphs to help prefill.
10. Hands-on exercise#
A. Measure the win. Run a real engine with and without graphs (--enforce-eager) at batch
1, 8, 32, 128. Plot the speedup vs batch. Compare to section 7.
B. Measure the cost. Record startup time and GPU memory (via the reported block count) with and without graph capture. How many concurrent sequences did you trade away?
C. Tune the list. Collect the batch-size distribution from a realistic load test. Design a capture list that minimizes expected padding waste within a fixed memory budget. Compare to the default.
D. Break it. Enable a feature known to conflict (multi-LoRA or speculative decoding, in a version where it conflicts). Check the logs and re-benchmark. Quantify what you lost.
E. Verify in a profile. Capture Nsight Systems timelines with and without graphs. Count the
cudaLaunchKernel calls and measure the total gap time in each.
11. Interview questions#
- Why do CUDA graphs help LLM decode so much at small batch?
- Why doesn’t varying context length require recapture?
- How do engines handle varying batch size, and what does it cost?
- How would you tune the capture list?
- What is the memory cost of graph capture and how does it affect capacity?
- Name three features that can conflict with graph capture.
- How would you detect that graphs were silently disabled?
12. Further reading#
- [REFERENCE] CUDA Programming Guide, “CUDA Graphs”
- [REFERENCE] vLLM’s
capture_modelimplementation and thecuda_graph_sizesoption - [REFERENCE] TensorRT-LLM in-flight batching documentation
- Next: 09 — TensorRT and TensorRT-LLM