PidokuInfra

HBM and Memory Bandwidth

Basic Intermediate 45 min Difficulty 3/5 Topic 03 of 05

Prerequisites 02

The idea in one minute#

Data-center GPUs use a special memory called HBM (high-bandwidth memory): memory chips stacked vertically and placed right next to the GPU, connected by thousands of microscopic wires. The goal is not capacity. It is bandwidth — how many bytes per second can flow between memory and the arithmetic units.

For large AI models, bandwidth is frequently the number that sets the speed. A GPU that can do a trillion multiplications per second is still stuck if the numbers to multiply arrive slowly.

An analogy#

A kitchen with fifty chefs and one narrow door to the pantry. Hiring more chefs does nothing; they queue at the door. Widening the door is the only thing that helps.

HBM is a very wide door.

A picture#

flowchart LR
  subgraph PKG["One GPU package"]
    direction LR
    subgraph STACK["HBM stack (several around the chip)"]
      direction TB
      D1["DRAM layer"] --- D2["DRAM layer"] --- D3["DRAM layer"] --- D4["... 8 to 16 layers"]
    end
    INT["Silicon interposer<br/>thousands of short wires"]
    DIE["GPU die"]
    STACK --- INT --- DIE
  end
  class D1,D2,D3,D4 memory
  class INT neutral
  class DIE compute

How it really works#

What bandwidth means#

bandwidth = (number of data wires) × (bits per second on each wire)

Ordinary memory modules sit centimetres from the CPU on a circuit board; only a few hundred wires fit, and long wires cannot be driven very fast. HBM attacks the first term: by stacking memory dies and connecting them with through-silicon vias (vertical wires through the chips), then mounting the stack on a silicon interposer beside the GPU, it gets over a thousand data wires per stack, each only millimetres long.

The generations#

MemoryUsed inBandwidth per GPU
GDDR6 / GDDR6X (not stacked)T4, L4, RTX 40900.3–1.0 TB/s
HBM2eA100~2.0 TB/s
HBM3H100~3.35 TB/s
HBM3eH200, B200, B300~4.8–8 TB/s
HBM4Rubin (shipping since August 2026); AMD MI455X~22 TB/s (vendor figures)

Consumer cards use GDDR: fast, cheap, conventional chips around the GPU. It is why a gaming card with strong arithmetic still has a fraction of a data-center card’s bandwidth.

HBM is also the scarce part. It is difficult to manufacture, made by three companies, and for several years its supply — not the GPU chips themselves — has limited how many AI GPUs exist.

Why AI cares so much#

Generating one token with a language model means reading every weight once. For a 16 GB model:

time per token ≥ 16 GB ÷ bandwidth
H100: 16 ÷ 3350 GB/s = 4.8 ms   →  at most ~210 tokens/s for one sequence
L4:   16 ÷ 300  GB/s = 53 ms    →  at most ~19 tokens/s

The arithmetic for that token would take the H100 well under a millisecond. The GPU spends most of the step waiting for bytes. Work like this is called memory-bound.

The cure is to make each byte do more work — process many sequences per pass over the weights (batching), or store each weight in fewer bytes (quantization). Module V covers both.

Capacity vs bandwidth#

They are different numbers and they fail differently:

  • Too little capacity → the program does not run (out of memory).
  • Too little bandwidth → the program runs, slowly, with the GPU reporting “100% utilization”.

Code#

Measure your own machine’s memory bandwidth. The result will be far below a GPU’s, and the experiment shows what “memory-bound” feels like: adding goroutines stops helping long before you run out of cores.

Go
// bandwidth.go — measure how fast this machine can stream bytes through memory.
package main

import (
	"fmt"
	"runtime"
	"sync"
	"time"
)

func main() {
	const n = 1 << 28 // 256M float32 = 1 GiB per array
	src, dst := make([]float32, n), make([]float32, n)
	for i := range src {
		src[i] = 1
	}

	for workers := 1; workers <= runtime.NumCPU(); workers *= 2 {
		chunk := n / workers
		t0 := time.Now()
		var wg sync.WaitGroup
		for w := 0; w < workers; w++ {
			a, b := src[w*chunk:(w+1)*chunk], dst[w*chunk:(w+1)*chunk]
			wg.Add(1)
			go func() {
				defer wg.Done()
				copy(b, a) // read 4 bytes, write 4 bytes, no arithmetic at all
			}()
		}
		wg.Wait()
		dt := time.Since(t0).Seconds()
		fmt.Printf("%2d workers: %6.1f GB/s\n", workers, 2*4*float64(n)/dt/1e9)
	}
}

Typical laptop: 15–30 GB/s with one worker, flattening at 30–100 GB/s no matter how many you add. That ceiling is your memory bandwidth. An H100’s is 3,350.

Remember this#

  • HBM = stacked memory beside the GPU, built for bytes per second, not capacity.
  • Bandwidth = wires × speed per wire. HBM wins by having far more, far shorter wires.
  • Reading every weight once per token makes LLM generation memory-bound at small batch sizes.
  • Capacity failures crash; bandwidth shortages just make things slow.

Try it#

  1. Run bandwidth.go. At what worker count does the number stop improving? Why does it stop?
  2. With your measured bandwidth, what is the upper limit on tokens/s for a 4 GB model on your CPU?
  3. A 70 GB model on a GPU with 8 TB/s: what is the maximum tokens/s for one sequence?

Check yourself#

  1. What physical trick gives HBM its bandwidth?
  2. Why does a gaming GPU with similar FLOP/s serve LLMs more slowly than a data-center GPU?
  3. What does “memory-bound” mean?

↑↓ navigate↵ openesc close