1. The classification, and why it’s the first question#
COMPUTE BOUND the arithmetic units are the constraint
MEMORY BOUND the memory system is the constraint
LATENCY BOUND neither is saturated; you're waiting on dependenciesEvery optimization helps exactly one of these and does nothing for the other two. Applying the wrong one is the single largest source of wasted performance work.
2. How to classify, three ways#
Method 1 — Nsight Compute (definitive)#
ncu --metrics \
sm__throughput.avg.pct_of_peak_sustained_elapsed,\
dram__throughput.avg.pct_of_peak_sustained_elapsed \
-k regex:"kernel" ./appsm% dram% classification
> 70 < 40 COMPUTE BOUND
< 40 > 70 MEMORY BOUND
> 60 > 60 BALANCED (well optimized)
< 40 < 40 LATENCY BOUND ← the interesting caseMethod 2 — arithmetic (fast, no tools)#
Compute FLOPs and bytes for the operation. Compare intensity to the ridge point.
I = FLOPs / Bytes
I_ridge = peak_FLOPs / peak_bandwidth
I < I_ridge → memory bound
I > I_ridge → compute boundWorks for anything you can count. For an LLM decode step:
FLOPs ≈ 2 × P × B
Bytes ≈ P × bytes_per_weight + KV_bytes
I ≈ 2B / bytes_per_weight ≈ B (for BF16)So decode intensity ≈ batch size. Below the ridge point (296 on H100), memory bound.
Method 3 — the perturbation test (empirical, no tools)#
If you can vary the clocks:
nvidia-smi -lgc <low>,<low> # lock SM clocks low
→ if the kernel slows proportionally, it's COMPUTE bound
→ if it barely changes, it's MEMORY bound
Memory clocks are harder to vary, but the SM-clock test alone is
usually conclusive.This is a genuinely useful trick when you don’t have profiler access — for example on a
production system where ncu is too invasive.
3. What helps each regime#
MEMORY BOUND (at the roofline)
✓ Quantization (fewer bytes per value)
✓ Kernel fusion (fewer round trips)
✓ Batching (more work per byte read)
✓ Algorithmic changes that read less (GQA, MoE, speculative decoding)
✓ Faster memory (H200 over H100)
✗ Faster arithmetic
✗ Tensor cores (for the FLOPs; they help via precision → bytes)
✗ More SMs
MEMORY BOUND (below the roofline)
✓ Coalescing, better layout
✓ Vectorized loads
✓ More occupancy (to have more requests in flight)
✓ Avoiding bank conflicts
Then the above.
COMPUTE BOUND (at the roofline)
✓ Lower precision (FP8, INT8 — 2x tensor core throughput)
✓ Fewer FLOPs (smaller model, sparsity, shorter sequences)
✓ Better hardware
✗ Quantizing weights only (compute stays the same)
✗ More bandwidth
✗ Fusion of surrounding elementwise ops (marginal)
COMPUTE BOUND (below the roofline)
✓ Use tensor cores (check you actually are — Section VI.11)
✓ Better tiling, avoid tile/wave quantization
✓ Reduce warp divergence
✓ Increase ILP
LATENCY BOUND
✓ More occupancy (more warps to hide latency)
✓ More instruction-level parallelism per thread
✓ Reduce dependency chains
✓ Prefetching / async copy
✓ Fewer, larger kernels (launch overhead is a form of this)
✗ Quantization (you're not bandwidth-limited)
✗ More FLOPs capacityPrint the “✗” lines and read them before your next optimization. They are the things that sound helpful and aren’t.
4. The LLM inference map#
PHASE / OPERATION REGIME WHY
Prefill GEMM (S > 512) compute intensity ~S
Prefill attention compute intensity ~S/4
Decode GEMM (B < 128) memory intensity ~B
Decode GEMM (B > 256) approaching compute intensity ~B
Decode attention (MHA) memory intensity ~1
Decode attention (GQA-8) memory intensity ~8
RMSNorm, residual, activation memory intensity ~1
Softmax memory intensity ~1
Sampling memory + latency small tensors
Embedding lookup memory (random) intensity ~0
LM head (decode) memory intensity ~B
Small-batch anything + latency bound launch overheadThe summary: LLM inference is memory-bound almost everywhere, with prefill as the exception.
This one fact explains: why quantization is the dominant optimization, why batching matters so much, why H200 beats H100 for inference, why GQA was adopted universally, and why speculative decoding works.
5. The regime changes with configuration#
The same kernel changes regime as you change parameters:
Decode GEMM, H100 (ridge 296):
batch 1 I = 1 memory bound (300x below ridge)
batch 32 I = 32 memory bound (9x below)
batch 128 I = 128 memory bound (2.3x below)
batch 512 I = 512 COMPUTE bound
Prefill attention:
S = 128 I ≈ 32 memory bound
S = 2048 I ≈ 512 compute bound
S = 32768 I ≈ 8192 deeply compute bound
Decode attention, context length:
ctx 1k, B=8 KV bytes small, weights dominate → weight-bound
ctx 32k, B=64 KV bytes dominate → KV-bandwidth boundTherefore: classify at your actual operating point, not in general. A system that’s memory-bound at batch 8 and compute-bound at batch 512 needs different optimizations at different times of day.
6. The latency-bound case (the one people miss)#
Both sm% and dram% low. The GPU is doing neither compute nor memory work
at capacity. It's waiting.
CAUSES:
1. Not enough parallelism (too few warps → nothing to switch to)
2. Long dependency chains (each instruction waits for the previous)
3. Launch overhead (the kernel is too small to matter)
4. Synchronization (barriers, or waiting on another stream)
5. Instruction cache misses (rare; huge unrolled loops)
DIAGNOSIS: Nsight Compute's Warp State Statistics
"Long Scoreboard" high + low occupancy → cause 1
"Not Selected" low, "Stall Wait" high → cause 2
kernel duration < 10 µs → cause 3
"Barrier" high → cause 4Small decode kernels at batch 1 are frequently latency-bound rather than memory-bound, which is why CUDA graphs and fusion help there more than quantization does.
7. Worked example — misdiagnosis and correction#
SITUATION
Team reports: "decode is slow, 45 ms per step for a 70B model at TP=8."
Expected from the roofline: 17.6 GB / 3.35 TB/s = 5.3 ms.
8.5x off.
FIRST HYPOTHESIS (wrong)
"It's memory bound and we need to quantize."
→ they spend 3 weeks on FP8. Result: 42 ms. 7% better.
CORRECT DIAGNOSIS
ncu shows: sm% = 12, dram% = 18. BOTH LOW → latency bound.
nsys shows: 4,800 kernel launches per step, 34 ms of gaps.
py-spy shows: a Python logits processor called per step, with .item().
ACTUAL FIXES
1. Remove the per-step .item() → 45 → 31 ms
2. Enable CUDA graphs → 31 → 9 ms
3. THEN quantize to FP8 → 9 → 6 ms
Total: 45 → 6 ms. The quantization was the LAST 33%, not the first.The three weeks on FP8 weren’t wasted — but they were done in the wrong order, and the
misdiagnosis (memory-bound when it was latency-bound) is what caused it. Ten minutes with ncu
would have revealed sm% = 12, dram% = 18.
8. Production implications#
- Classify before optimizing. Two profiler metrics, ten minutes.
- Reclassify at each operating point. Batch 8 and batch 256 are different regimes.
- The “both low” case means look at the CPU and the timeline, not the kernels.
- Track dram% for decode as a health metric. For a well-tuned system it should be 70-90%. A drop means something regressed.
- Reject optimizations that don’t match the regime. “We reduced FLOPs 40%” for a memory-bound kernel is worth zero; say so.
9. Common mistakes#
Assuming memory-bound because “LLM decode is memory-bound.” Usually true, but verify — latency-bound looks similar in wall-clock terms and needs completely different fixes.
Optimizing FLOPs for a memory-bound kernel.
Optimizing bandwidth for a compute-bound kernel.
Not recognizing the latency-bound case.
Classifying once and never rechecking as the workload changes.
Using the whole-model average. Classify per kernel.
10. Hands-on exercise#
A. Classify everything. For a decode step, get sm% and dram% for every kernel. Classify
each. Which are at their roof? Which are latency bound?
B. Watch the regime change. For the main decode GEMM, measure intensity and classification at batch 1, 8, 32, 128, 512. At what batch does it cross the ridge point? Does the measured crossover match the calculated one?
C. The clock test. Lock SM clocks to a low value and re-measure a kernel you believe is memory-bound. Does its time change? Repeat for one you believe is compute-bound.
D. Construct a latency-bound kernel. Write a kernel with a long dependency chain and low occupancy. Verify both sm% and dram% are low. Then fix it (more independent work per thread) and re-measure.
E. The misdiagnosis exercise. Take the worked example in section 7. On a real system,
deliberately introduce a per-step .item() and measure the effect. Then verify with ncu that
it presents as latency-bound rather than memory-bound.
11. Interview questions#
- How do you classify a kernel as memory-bound, compute-bound, or latency-bound?
- Name three optimizations that help memory-bound kernels and three that don’t.
- What does it mean when both sm% and dram% are low?
- At what batch size does LLM decode become compute-bound on an H100? Show the reasoning.
- Why does weight-only quantization not help a compute-bound kernel?
- How would you classify a kernel without a profiler?
- Walk through a case where someone misdiagnosed the regime. What did it cost?
12. Further reading#
- [FUNDAMENTAL] Williams et al., “Roofline” (CACM 2009)
- [REFERENCE] Nsight Compute “Speed of Light” section documentation
- [ESTABLISHED] Volkov, “Better Performance at Lower Occupancy” — the latency-bound case
- Next: 04 — GPU utilization is a lie