Below the API

The Roofline Model

Intermediate Intermediate 1h 10m Difficulty 4/5

Prerequisites II.03, II.04

The idea in one minute#

Any operation needs two things from the GPU: arithmetic (FLOPs) and data (bytes moved to and from memory). The GPU supplies each at a maximum rate. Whichever runs out first sets the speed.

The roofline model turns that into one picture and one formula. It tells you, before you write any code, the best speed an operation can reach on a given GPU and — more usefully — which resource is the limit, so you know what kind of optimization can help.

An analogy#

A bakery can be limited by its ovens or by its flour deliveries. If flour arrives slowly, buying more ovens changes nothing. If flour is piled to the ceiling and the ovens are full, a faster truck changes nothing.

The first question for any slow bakery is “ovens or flour?”. The roofline answers that for a GPU.

A picture#

flowchart LR
  OP["Operation"] --> I["Arithmetic intensity<br/>FLOPs per byte moved"]
  I --> Q{"Intensity below<br/>the GPU's ridge?"}
  Q -->|"yes"| MB["Memory-bound<br/>speed = bandwidth x intensity"]
  Q -->|"no"| CB["Compute-bound<br/>speed = peak FLOP/s"]
  MB --> FIX1["Helps: move fewer bytes<br/>batching, quantization, fusion"]
  CB --> FIX2["Helps: cheaper arithmetic<br/>tensor cores, lower precision"]
  class OP,I neutral
  class Q queue
  class MB,FIX1 memory
  class CB,FIX2 compute

How it really works#

Arithmetic intensity#

intensity = FLOPs performed ÷ bytes moved through memory

It is a property of the operation, not the hardware.

OperationFLOPsBytes (FP16)Intensity
y = a*x + y, n elements2n6n (read x, read y, write y)0.33
Matrix × vector, d×d2d²~2d²~1
Matrix × B vectors (batched)2Bd²~2d²~B
Matrix × matrix, n×n2n³6n²n/3

The ridge point#

A GPU has a peak arithmetic rate P (FLOP/s) and a bandwidth W (bytes/s). Their ratio is the GPU’s ridge point:

ridge = P ÷ W      (FLOPs per byte)

H100, FP16:  990e12 ÷ 3.35e12 ≈ 295 FLOPs per byte

The GPU can do 295 FLOPs in the time it takes to fetch one byte. An operation that needs fewer FLOPs per byte than that leaves arithmetic idle: it is memory-bound. One that needs more is compute-bound.

The formula#

attainable FLOP/s = min(P, W × intensity)

Plotted against intensity, it rises as a slope (the bandwidth limit) and then goes flat (the compute limit) — the shape of a roof.

What it tells you about LLMs#

One decode step for one sequence is essentially matrix × vector: intensity ≈ 1. Against a ridge of ~295, that operation can use at most 1/295 of the GPU’s arithmetic. The other 99.7% is idle, waiting for bytes.

Batch 64 sequences and the same pass over the weights does 64x the arithmetic: intensity ≈ 64. Throughput rises almost 64x, and each sequence is barely slower. This is the entire reason batching dominates LLM serving, derived from two data-sheet numbers.

What it tells you about fixes#

RegimeWorksDoes nothing
Memory-boundBigger batches; fewer bytes per weight (quantization); fusing kernels so data is read once; a GPU with more bandwidthFaster arithmetic, more tensor cores
Compute-boundLower-precision arithmetic on tensor cores; a GPU with more FLOP/s; a smaller modelMore bandwidth

Applying a fix from the wrong row is the most common waste of optimization effort.

Limits of the model#

The roofline is a ceiling. Real kernels land below it because of launch overhead, imperfect memory access (lesson 02), idle SMs (lesson 03) and practical arithmetic efficiency of 50–80% of peak. If you measure 60–80% of the roofline, the kernel is good. If you measure 10%, there is something to find.

Code#

// roofline.go — predict the regime and speed of an operation on a GPU.
package main

import "fmt"

type GPU struct {
	Name      string
	PeakFLOPs float64 // FLOP/s in the format you will use
	Bandwidth float64 // bytes/s
}

func (g GPU) Ridge() float64 { return g.PeakFLOPs / g.Bandwidth }

// Attainable returns the best possible FLOP/s for an operation of the given intensity.
func (g GPU) Attainable(intensity float64) (flops float64, regime string) {
	if mem := g.Bandwidth * intensity; mem < g.PeakFLOPs {
		return mem, "memory-bound"
	}
	return g.PeakFLOPs, "compute-bound"
}

func main() {
	h100 := GPU{"H100 (FP16)", 990e12, 3.35e12}
	fmt.Printf("%s ridge: %.0f FLOPs/byte\n\n", h100.Name, h100.Ridge())

	// One decode step of an 8B-parameter model in FP16, at different batch sizes.
	const params, bytesPerParam = 8e9, 2.0
	fmt.Println("batch  intensity  regime          step time  tokens/s total  tokens/s per user")
	for _, B := range []float64{1, 4, 16, 64, 256, 1024} {
		flops := 2 * params * B         // 2 FLOPs per weight per sequence
		bytes := params * bytesPerParam // weights are read once, whatever B is
		rate, regime := h100.Attainable(flops / bytes)
		step := flops / rate
		fmt.Printf("%5.0f  %9.0f  %-14s  %6.1f ms  %14.0f  %17.0f\n",
			B, flops/bytes, regime, step*1e3, B/step, 1/step)
	}
}

Read the output: tokens/s total climbs with the batch while tokens/s per user stays flat — until the regime flips to compute-bound, after which each user starts to slow down.

Remember this#

  • Intensity = FLOPs per byte moved. It belongs to the operation.
  • Ridge = peak FLOP/s ÷ bandwidth. It belongs to the GPU.
  • Below the ridge: memory-bound. Above: compute-bound. speed = min(P, W × intensity).
  • Know the regime before optimizing; fixes for one regime do nothing in the other.

Try it#

  1. Run roofline.go. At what batch size does the H100 become compute-bound for this model?
  2. Add an A100 (312e12 FLOP/s, 2.0e12 B/s) and a B200 (2250e12, 8e12). Which has the highest ridge? What does that mean for the batch size needed to use it fully?
  3. Halve bytesPerParam (8-bit weights). What happens to batch-1 speed? To the crossover?

Check yourself#

  1. Define arithmetic intensity and the ridge point.
  2. An operation is memory-bound. Name two changes that speed it up and one that will not.
  3. Why does batching raise throughput without (at first) slowing individual users?

↑↓ navigate ↵ open