The idea in one minute#
Data on a GPU can live in several places. The closer to the arithmetic units, the faster and the smaller. From nearest to furthest: registers, shared memory / L1, L2 cache, global memory (HBM), and finally host RAM across the PCIe link.
Fast GPU code keeps what it is working on in the near levels and touches the far levels as few times as possible.
An analogy#
A carpenter’s workshop:
- Hands — the tool in use right now. Instant. (Registers)
- Workbench — a few tools shared by the crew at this bench. One step away. (Shared memory)
- Tool wall — shared by the whole workshop. A short walk. (L2 cache)
- Warehouse out back — holds everything, but each trip takes a while. (HBM)
- Supplier across town — a truck delivery. (Host RAM over PCIe)
Nobody walks to the warehouse for every screw. You fetch a box, put it on the bench, and work from there.
A picture#
flowchart TB
REG[("Registers<br/>per thread, about 1 cycle")]
SHM[("Shared memory and L1<br/>per SM, up to 256 KB, a few cycles")]
L2[("L2 cache<br/>whole chip, 50 MB, tens of cycles")]
HBM[("Global memory: HBM<br/>80 GB, hundreds of cycles, 3.35 TB/s")]
HOST[("Host RAM<br/>across PCIe, about 64 GB/s")]
REG --- SHM --- L2 --- HBM --- HOST
class REG compute
class SHM,L2 memory
class HBM queue
class HOST warn(Sizes are for an H100. Other GPUs differ in the numbers, not the shape.)
How it really works#
| Level | Scope | Size (H100) | Who manages it |
|---|---|---|---|
| Registers | One thread | 65,536 32-bit registers per SM, divided among resident threads | Compiler |
| Shared memory | Threads of one block, on one SM | Configurable, up to ~228 KB per SM | You |
| L1 cache | One SM | Shares a 256 KB pool with shared memory | Hardware |
| L2 cache | Whole chip | 50 MB | Hardware |
| Global memory (HBM) | Whole chip | 80 GB | You allocate, hardware serves |
| Host RAM | The CPU | Hundreds of GB | You copy across |
Registers#
Each thread’s local variables live in registers. The pool is fixed per SM, so a kernel that needs many registers per thread lets fewer threads be resident. This is one of the main inputs to occupancy (IV.03).
Shared memory: the one you control#
Shared memory is a small scratchpad that threads cooperating on the same piece of work can all read and write. It is as fast as a cache, but you decide what goes in it. The classic pattern:
- Load a tile of data from global memory into shared memory, once.
- Have every thread in the block use it many times.
- Write results back to global memory, once.
That turns many slow reads into one. Almost every high-performance kernel (matrix multiply, convolution, attention) is built on this pattern.
Caches#
L1 and L2 work like CPU caches but are much smaller relative to the data. An 80 GB memory behind a 50 MB L2 means that for large arrays the cache holds well under 0.1% of the data. Do not expect caches to rescue you; plan your memory traffic.
Global memory#
This is “GPU memory” in the everyday sense — where the model weights, inputs and outputs live. It is big and, per byte, fast in bulk (terabytes per second), but each individual access is slow (hundreds of cycles). GPUs cope by streaming large contiguous amounts and by keeping other warps busy during the wait.
Host RAM#
The far end. Moving 16 GB of weights from host to device at ~25 GB/s (a realistic PCIe rate) takes more than half a second; reading them back inside the GPU takes 5 milliseconds. The ratio — roughly 100x — is why data should cross PCIe once.
Code#
How long does it take to read 16 GB (an 8-billion-parameter model in FP16) at each level’s speed? The numbers make the hierarchy concrete.
// hierarchy.go — the same 16 GB, moved at the speed of each level.
package main
import "fmt"
func main() {
const gb = 16.0
levels := []struct {
name string
gbps float64
}{
{"NVMe SSD (cold start)", 5},
{"Host RAM over PCIe, pageable", 12},
{"Host RAM over PCIe, pinned", 25},
{"GPU global memory (HBM, H100)", 3350},
}
for _, l := range levels {
sec := gb / l.gbps
fmt.Printf("%-32s %8.1f GB/s %9.1f ms\n", l.name, l.gbps, sec*1000)
}
fmt.Println("\nA decode step reads every weight once, so only the last row is fast")
fmt.Println("enough to do per token. Everything above it happens once, at load time.")
}Remember this#
- Nearer the arithmetic = faster and smaller: registers → shared/L1 → L2 → HBM → host.
- Shared memory is the level you program directly. Load once, reuse many times.
- Caches are tiny compared with GPU data. Design for streaming, not for cache hits.
- Crossing PCIe is ~100x slower than reading HBM. Do it once.
Try it#
- Run
hierarchy.go. Add a row for a 10 Gbit/s network (1.25 GB/s). How long to pull the model from another machine? - A kernel uses 64 registers per thread. With 65,536 registers per SM, how many threads can be resident? What if it used 128?
- Explain in two sentences why the “load a tile into shared memory” pattern speeds up matrix multiplication. (You will build it in V.01.)
Check yourself#
- List the five levels in order, with a rough size for each.
- Which level is managed by the programmer rather than the hardware?
- Why do GPU caches help less than CPU caches for AI workloads?