Below the API

Bandwidth and the Roofline on GPU

Intermediate Intermediate 1h Difficulty 3/5

Prerequisites I.07, I.08, 03


1. What is it?#

The roofline model applied to GPUs: a single plot that tells you, for any kernel, whether it’s limited by memory bandwidth or by arithmetic, and how far from the limit you are.

 attainable
 FLOP/s
   ▲
   │                    ╭──────────────────── peak compute (990 TF/s)
   │                   ╱
   │                  ╱
   │                 ╱  ← ridge point (I = 296 FLOP/byte on H100)
   │                ╱
   │               ╱
   │              ╱  slope = memory bandwidth (3.35 TB/s)
   │             ╱
   │    ●decode ╱          ●prefill
   │   ●norm   ╱
   └──────────────────────────────────────────► arithmetic intensity
       1      10     100    1000   10000       (FLOP/byte)

Plot your kernels on it. Points below the sloped line are memory-bound and could go faster with better access patterns. Points on the line are at the bandwidth limit. Points under the flat line are compute-bound.


2. Why does it exist?#

Because “is this kernel fast?” is unanswerable without a reference. The roofline provides the reference: the fastest this kernel could possibly run given its intensity.

It converts optimization from guesswork into a decision procedure:

Am I near the roof?          → done; change the algorithm, not the kernel
Am I far below and left?     → memory-bound; improve access patterns or reduce bytes
Am I far below and right?    → compute-bound; improve ILP, use tensor cores

3. Simple analogy#

A road with a speed limit and a fuel-delivery limit.

The flat roof is the speed limit (peak FLOP/s) — you can’t go faster no matter how much fuel you have.

The sloped line is the fuel supply (bandwidth) — if your engine burns fuel faster than the tanker can deliver it, you’re limited by delivery, not by the engine.

The ridge point is where they balance: the fuel efficiency (FLOP per byte) at which the tanker exactly keeps up with the engine at full speed.


4. Tiny example#

Build your own roofline:

import torch, time, math

def bench(fn, iters=50):
    for _ in range(5): fn()
    torch.cuda.synchronize(); t0 = time.perf_counter()
    for _ in range(iters): fn()
    torch.cuda.synchronize()
    return (time.perf_counter()-t0)/iters

N = 100_000_000
x = torch.randn(N, device='cuda', dtype=torch.float32)

points = []

# Vary arithmetic intensity by doing K fused ops per element
for K in [1, 2, 4, 8, 16, 32, 64, 128, 256]:
    def f(K=K):
        y = x
        for _ in range(K): y = y * 1.0001 + 0.5
        return y
    # NOTE: run this under torch.compile so the ops fuse into one kernel,
    # otherwise each op is its own memory round trip.
    fc = torch.compile(f)
    t = bench(fc)
    flops = 2 * K * N          # one FMA per op
    byts  = 2 * 4 * N          # read x, write y, once (fused)
    points.append((flops/byts, flops/t/1e12))
    print(f"K={K:4d}  I={flops/byts:7.1f}  {flops/t/1e12:7.2f} TFLOP/s")

Plot TFLOP/s vs I on log-log axes. You’ll see the characteristic roofline: a rising line that flattens. The elbow is your machine’s ridge point, measured.


5. Technical explanation#

The model#

attainable_FLOPs = min(peak_FLOPs, bandwidth × arithmetic_intensity)
Ridge point = peak_FLOPs / bandwidth

A100:   312e12 / 2.04e12  = 153 FLOP/byte
H100:   990e12 / 3.35e12  = 296
H200:   990e12 / 4.80e12  = 206
L40S:   362e12 / 0.864e12 = 419
B200:  2250e12 / 8.0e12   = 281

Where LLM kernels sit#

Kernel                              Intensity      Regime
Residual add                        0.08           memory ██████████
RMSNorm                             1.25           memory ██████████
Softmax                             ~1             memory ██████████
Decode GEMV (batch 1)               1              memory ██████████
Decode GEMM (batch 32)              32             memory ████████
Decode GEMM (batch 256)             250            borderline
Decode attention, MHA               1              memory ██████████
Decode attention, GQA-8             8              memory ████████
Prefill attention (S=2048)          ~500           compute ██████
Prefill GEMM (S=2048)               ~1400          compute ████████

Almost everything is memory-bound. The list of compute-bound LLM kernels is short: large-S prefill GEMMs and prefill attention.

The hierarchical roofline#

The simple model uses HBM bandwidth. But a kernel that hits L2 or shared memory has a higher effective bandwidth:

       ╭──────────────── peak compute
      ╱ ╱ ╱
     ╱ ╱ ╱  ← shared memory roofline (19 TB/s)
    ╱ ╱ ╱   ← L2 roofline (7 TB/s)
   ╱ ╱ ╱    ← HBM roofline (3.35 TB/s)
  ╱ ╱ ╱

A kernel that appears to exceed the HBM roofline isn’t violating physics — it’s getting cache hits. This is precisely what tiling achieves: move up to a higher roofline.

FlashAttention’s speedup is exactly “moved from the HBM roofline to the shared-memory roofline.”

Reading a roofline diagnostically#

Point far below the sloped line, low I:
  → memory-bound AND not achieving peak bandwidth
  → uncoalesced access, bad layout, or insufficient parallelism
  → FIX: access patterns, vectorized loads, more warps

Point ON the sloped line:
  → at the bandwidth limit
  → FIX: reduce bytes (quantization, fusion, better algorithm). Kernel tuning won't help.

Point far below the flat roof, high I:
  → compute-bound but not achieving peak FLOPs
  → not using tensor cores, warp divergence, low ILP, or bad tiling
  → FIX: tensor cores, better tiling

Point ON the flat roof:
  → done. Change the algorithm or the hardware.

6. Under the hood#

Nsight Compute generates rooflines automatically:

ncu --set roofline --kernel-name regex:".*gemm.*" -o profile ./your_app
ncu-ui profile.ncu-rep     # open the GUI, view the roofline chart

Or compute it from metrics:

ncu --metrics \
  sm__sass_thread_inst_executed_op_ffma_pred_on.sum,\
  dram__bytes.sum,\
  gpu__time_duration.sum ./your_app
intensity = (2 × ffma_count) / dram_bytes
achieved  = (2 × ffma_count) / duration

For tensor-core kernels, use sm__inst_executed_pipe_tensor.sum and the appropriate FLOP multiplier for the MMA shape.


7. Performance implications#

The roofline tells you which optimizations can possibly help:

SituationUselessUseful
Memory-bound, at the rooffaster math, more FLOPs/elementquantization, fusion, fewer bytes
Memory-bound, below the roofquantization (won’t fix access)coalescing, layout, vectorized loads, more parallelism
Compute-bound, below the roofmore bandwidth, better layouttensor cores, tiling, ILP, reduce divergence
Compute-bound, at the roofeverythingsmaller model, sparsity, lower precision

Half of all wasted optimization effort is applying a fix from the wrong row.


8. Production implications#

  • Establish the roofline position of your top 5 kernels once. It tells you where engineering effort will pay and where it won’t.
  • Track achieved-vs-roofline as a regression metric. If your decode step historically hits 85% of the HBM roofline and drops to 55%, something changed — a kernel fell back, a layout broke, a shape shifted.
  • Use it to argue for hardware. “We’re at 92% of the HBM roofline; the only remaining lever is more bandwidth” is a defensible procurement case.
  • Use it to reject bad ideas. “This optimization reduces FLOPs by 40%” — if the kernel is at the memory roof, that’s worth zero.

9. Common mistakes#

Using peak FLOPs that include sparsity. Halves your apparent efficiency.

Ignoring the hierarchical roofline. A kernel “above” the HBM roofline is cache-resident, not broken.

Computing intensity from the algorithm rather than the implementation. An unfused chain has much lower effective intensity than the math implies.

Forgetting that reads and writes both count.

Assuming the roofline applies to the whole model. It applies per kernel. A model contains kernels in both regimes.


10. Hands-on exercise#

A. Measure your ridge point. Run the sweep in section 4. Find the elbow. Compare to peak_FLOPs / bandwidth from the spec sheet. Explain any difference.

B. Plot your model’s kernels. Profile a real LLM decode step with ncu. For the top 10 kernels by time, compute intensity and achieved FLOP/s. Plot them on a roofline. Which are at the roof? Which have headroom?

C. Move a kernel up. Take a memory-bound unfused chain, fuse it with torch.compile, and plot both points on the roofline. Confirm the intensity increased and the point moved right.

D. The hierarchical roofline. Measure a tiled matmul’s DRAM bytes and L1/shared bytes. Compute both rooflines. Which one is it actually near?

E. Predict then measure. For a kernel you haven’t profiled, predict its position from the algorithm. Then measure. How close were you?


11. Interview questions#

  1. Explain the roofline model and how to read it.
  2. Compute the ridge point for an H100 and an L40S. What does the difference imply?
  3. A kernel is at 92% of the memory roofline. What optimizations remain?
  4. A kernel is at 30% of the memory roofline. What do you investigate?
  5. What is a hierarchical roofline and how does FlashAttention relate to it?
  6. Why does a 40% FLOP reduction sometimes produce zero speedup?
  7. Where do LLM decode, prefill, and normalization sit on the roofline?

12. Further reading#

  • [FUNDAMENTAL] Williams, Waterman, Patterson, “Roofline: An Insightful Visual Performance Model” (CACM 2009)
  • [REFERENCE] NVIDIA Nsight Compute roofline documentation
  • [ESTABLISHED] “Hierarchical Roofline Analysis for GPUs” (Yang et al., NERSC)
  • Next: 05 — Kernel launch and host-device interaction

↑↓ navigate ↵ open