The idea in one minute#
Any operation needs two things from the GPU: arithmetic (FLOPs) and data (bytes moved to and from memory). The GPU supplies each at a maximum rate. Whichever runs out first sets the speed.
The roofline model turns that into one picture and one formula. It tells you, before you write any code, the best speed an operation can reach on a given GPU and — more usefully — which resource is the limit, so you know what kind of optimization can help.
An analogy#
A bakery can be limited by its ovens or by its flour deliveries. If flour arrives slowly, buying more ovens changes nothing. If flour is piled to the ceiling and the ovens are full, a faster truck changes nothing.
The first question for any slow bakery is “ovens or flour?”. The roofline answers that for a GPU.
A picture#
flowchart LR
OP["Operation"] --> I["Arithmetic intensity<br/>FLOPs per byte moved"]
I --> Q{"Intensity below<br/>the GPU's ridge?"}
Q -->|"yes"| MB["Memory-bound<br/>speed = bandwidth x intensity"]
Q -->|"no"| CB["Compute-bound<br/>speed = peak FLOP/s"]
MB --> FIX1["Helps: move fewer bytes<br/>batching, quantization, fusion"]
CB --> FIX2["Helps: cheaper arithmetic<br/>tensor cores, lower precision"]
class OP,I neutral
class Q queue
class MB,FIX1 memory
class CB,FIX2 computeHow it really works#
Arithmetic intensity#
intensity = FLOPs performed ÷ bytes moved through memoryIt is a property of the operation, not the hardware.
| Operation | FLOPs | Bytes (FP16) | Intensity |
|---|---|---|---|
y = a*x + y, n elements | 2n | 6n (read x, read y, write y) | 0.33 |
| Matrix × vector, d×d | 2d² | ~2d² | ~1 |
| Matrix × B vectors (batched) | 2Bd² | ~2d² | ~B |
| Matrix × matrix, n×n | 2n³ | 6n² | n/3 |
The ridge point#
A GPU has a peak arithmetic rate P (FLOP/s) and a bandwidth W (bytes/s). Their ratio is the
GPU’s ridge point:
ridge = P ÷ W (FLOPs per byte)
H100, FP16: 990e12 ÷ 3.35e12 ≈ 295 FLOPs per byteThe GPU can do 295 FLOPs in the time it takes to fetch one byte. An operation that needs fewer FLOPs per byte than that leaves arithmetic idle: it is memory-bound. One that needs more is compute-bound.
The formula#
attainable FLOP/s = min(P, W × intensity)Plotted against intensity, it rises as a slope (the bandwidth limit) and then goes flat (the compute limit) — the shape of a roof.
What it tells you about LLMs#
One decode step for one sequence is essentially matrix × vector: intensity ≈ 1. Against a ridge of ~295, that operation can use at most 1/295 of the GPU’s arithmetic. The other 99.7% is idle, waiting for bytes.
Batch 64 sequences and the same pass over the weights does 64x the arithmetic: intensity ≈ 64. Throughput rises almost 64x, and each sequence is barely slower. This is the entire reason batching dominates LLM serving, derived from two data-sheet numbers.
What it tells you about fixes#
| Regime | Works | Does nothing |
|---|---|---|
| Memory-bound | Bigger batches; fewer bytes per weight (quantization); fusing kernels so data is read once; a GPU with more bandwidth | Faster arithmetic, more tensor cores |
| Compute-bound | Lower-precision arithmetic on tensor cores; a GPU with more FLOP/s; a smaller model | More bandwidth |
Applying a fix from the wrong row is the most common waste of optimization effort.
Limits of the model#
The roofline is a ceiling. Real kernels land below it because of launch overhead, imperfect memory access (lesson 02), idle SMs (lesson 03) and practical arithmetic efficiency of 50–80% of peak. If you measure 60–80% of the roofline, the kernel is good. If you measure 10%, there is something to find.
Code#
// roofline.go — predict the regime and speed of an operation on a GPU.
package main
import "fmt"
type GPU struct {
Name string
PeakFLOPs float64 // FLOP/s in the format you will use
Bandwidth float64 // bytes/s
}
func (g GPU) Ridge() float64 { return g.PeakFLOPs / g.Bandwidth }
// Attainable returns the best possible FLOP/s for an operation of the given intensity.
func (g GPU) Attainable(intensity float64) (flops float64, regime string) {
if mem := g.Bandwidth * intensity; mem < g.PeakFLOPs {
return mem, "memory-bound"
}
return g.PeakFLOPs, "compute-bound"
}
func main() {
h100 := GPU{"H100 (FP16)", 990e12, 3.35e12}
fmt.Printf("%s ridge: %.0f FLOPs/byte\n\n", h100.Name, h100.Ridge())
// One decode step of an 8B-parameter model in FP16, at different batch sizes.
const params, bytesPerParam = 8e9, 2.0
fmt.Println("batch intensity regime step time tokens/s total tokens/s per user")
for _, B := range []float64{1, 4, 16, 64, 256, 1024} {
flops := 2 * params * B // 2 FLOPs per weight per sequence
bytes := params * bytesPerParam // weights are read once, whatever B is
rate, regime := h100.Attainable(flops / bytes)
step := flops / rate
fmt.Printf("%5.0f %9.0f %-14s %6.1f ms %14.0f %17.0f\n",
B, flops/bytes, regime, step*1e3, B/step, 1/step)
}
}
Read the output: tokens/s total climbs with the batch while tokens/s per user stays flat — until the regime flips to compute-bound, after which each user starts to slow down.
Remember this#
- Intensity = FLOPs per byte moved. It belongs to the operation.
- Ridge = peak FLOP/s ÷ bandwidth. It belongs to the GPU.
- Below the ridge: memory-bound. Above: compute-bound.
speed = min(P, W × intensity). - Know the regime before optimizing; fixes for one regime do nothing in the other.
Try it#
- Run
roofline.go. At what batch size does the H100 become compute-bound for this model? - Add an A100 (312e12 FLOP/s, 2.0e12 B/s) and a B200 (2250e12, 8e12). Which has the highest ridge? What does that mean for the batch size needed to use it fully?
- Halve
bytesPerParam(8-bit weights). What happens to batch-1 speed? To the crossover?
Check yourself#
- Define arithmetic intensity and the ridge point.
- An operation is memory-bound. Name two changes that speed it up and one that will not.
- Why does batching raise throughput without (at first) slowing individual users?