1. What is it?#
A GPU is a chip organized around one idea: do the same operation to enormous amounts of data, and spend every available transistor on arithmetic and memory bandwidth rather than on making a single instruction stream fast.
H100 SXM:
┌─────────────────────────────────────────────────────────────┐
│ GPC 0 GPC 1 ... GPC 7 │
│ ┌────────┐ ┌────────┐ ┌────────┐ │
│ │ TPC×9 │ │ TPC×9 │ │ TPC×9 │ │
│ │ (SMs) │ │ │ │ │ │
│ └────────┘ └────────┘ └────────┘ │
│ │
│ 132 SMs total, 16,896 FP32 CUDA cores │
│ 528 Tensor Cores │
├─────────────────────────────────────────────────────────────┤
│ L2 Cache 50 MB │
├─────────────────────────────────────────────────────────────┤
│ HBM3 80 GB @ 3.35 TB/s (5 stacks) │
└─────────────────────────────────────────────────────────────┘
↕ NVLink 900 GB/s ↕ PCIe Gen5 x16 64 GB/sFor inference, the two numbers that matter are on the bottom two rows: HBM capacity (what fits) and HBM bandwidth (how fast decode runs).
Diagram — A GPU from the outside in#
flowchart TB
CPU["Host CPU + system RAM"] -->|"PCIe"| GPU
subgraph GPU["GPU"]
direction TB
subgraph SMS["Streaming Multiprocessors - a hundred or more"]
direction LR
SM1["SM<br/>warp schedulers<br/>CUDA cores<br/>tensor cores<br/>registers<br/>shared memory / L1"]
SM2["SM"]
SM3["SM ..."]
end
L2["L2 cache - tens of MB, shared by all SMs"]
HBM[("HBM global memory<br/>tens to hundreds of GB at TB/s")]
SMS --> L2 --> HBM
end
GPU <-->|"NVLink"| PEER["Other GPUs"]
class SM1,SM2,SM3 compute
class L2,HBM memory
class CPU,PEER neutral2. Why does it exist?#
Graphics: color a million pixels, each independently, with the same shader. Perfect data parallelism, no branching, huge memory traffic. Building a chip for that shape of problem gives you something very different from a CPU.
Neural networks turned out to have the same shape. GPUs were already there.
The design bet, explicitly: give up single-thread performance, latency-avoidance machinery, and complex control flow; buy arithmetic units and memory bandwidth. For inference, that bet pays roughly 10x on bandwidth and 100x on matrix throughput.
3. Simple analogy#
A chip is a factory floor, and you’re choosing between craftsmen and an assembly line.
A CPU is eight master craftsmen, each with a full workshop, capable of any job, each finishing one item quickly.
A GPU is an assembly line with 16,000 stations, each doing one simple operation. Any single item takes longer to traverse the line. But throughput is thousands of times higher — provided every station has work, which is why occupancy and batch size matter so much.
4. Tiny example#
Read your own hardware:
import torch
p = torch.cuda.get_device_properties(0)
print(f"name {p.name}")
print(f"compute capability {p.major}.{p.minor}")
print(f"SMs {p.multi_processor_count}")
print(f"total memory {p.total_memory/1e9:.1f} GB")
print(f"max threads/SM {p.max_threads_per_multi_processor}")
print(f"shared mem/block {p.shared_memory_per_block/1024:.0f} KB")
print(f"warp size {p.warp_size}")
print(f"clock {p.clock_rate/1e6:.2f} GHz")
Then compute the peak yourself:
FP32 peak = SMs × 128 CUDA cores/SM × 2 (FMA) × clock
H100: 132 × 128 × 2 × 1.755e9 = 59.3 TFLOP/s (spec says 67 with boost)And measure the bandwidth:
import torch, time
n = 2_000_000_000 // 4
a = torch.empty(n, device='cuda', dtype=torch.float32)
b = torch.empty_like(a)
for _ in range(3): b.copy_(a)
torch.cuda.synchronize(); t0 = time.perf_counter()
for _ in range(20): b.copy_(a)
torch.cuda.synchronize()
dt = (time.perf_counter()-t0)/20
print(f"{2*n*4/dt/1e9:.0f} GB/s") # read + write
Record both in numbers.md. These are the two constants you’ll use in every estimate for
the rest of the curriculum.
5. Technical explanation#
Inside an SM (Hopper)#
┌────────────────────────── SM ──────────────────────────┐
│ L1 Instruction Cache │
│ ┌──────────────┬──────────────┬────────────┬────────┐ │
│ │ Processing │ Processing │ Processing │ Proc. │ │
│ │ Block 0 │ Block 1 │ Block 2 │ Blk 3 │ │
│ │ │ │ │ │ │
│ │ Warp Sched │ Warp Sched │ ... │ ... │ │
│ │ 32 FP32 cores│ 32 FP32 │ │ │ │
│ │ 16 INT32 │ │ │ │ │
│ │ 1 Tensor Core│ 1 TC │ 1 TC │ 1 TC │ │
│ │ 16K registers│ 16K │ 16K │ 16K │ │
│ │ 8 LD/ST │ │ │ │ │
│ │ 4 SFU │ │ │ │ │
│ └──────────────┴──────────────┴────────────┴────────┘ │
│ 256 KB combined L1 / Shared Memory (configurable split)│
│ Tensor Memory Accelerator (TMA) ← Hopper: async copy │
└─────────────────────────────────────────────────────────┘Key facts:
- 4 warp schedulers per SM, each issuing 1 instruction per cycle to its 32-lane group.
- Up to 64 warps resident per SM (2,048 threads) — this is how latency is hidden.
- 256 KB of register file per SM, statically partitioned among resident threads. More registers per thread = fewer resident warps.
- 256 KB combined L1/shared memory, split configurably (e.g. 228 KB shared / 28 KB L1).
- 4 tensor cores per SM, one per processing block — these do the matrix math.
The generations#
| V100 | A100 | H100 | B200* | |
|---|---|---|---|---|
| SMs | 80 | 108 | 132 | 148×2 dies |
| Memory | 32 GB HBM2 | 80 GB HBM2e | 80 GB HBM3 | 192 GB HBM3e |
| Bandwidth | 900 GB/s | 2.0 TB/s | 3.35 TB/s | ~8 TB/s |
| FP16 tensor | 125 TF | 312 TF | 990 TF | ~2,250 TF |
| FP8 | — | — | 1,979 TF | ~4,500 TF |
| L2 | 6 MB | 40 MB | 50 MB | larger |
| NVLink | 300 GB/s | 600 GB/s | 900 GB/s | 1.8 TB/s |
| New for inference | tensor cores | sparsity, MIG | FP8, TMA, DSMEM | FP4, faster |
*Blackwell figures are approximate; check current specs.
Notice the trend: FLOPs grew 18x from V100 to H100; bandwidth grew 3.7x. The ridge point keeps rising, so workloads keep becoming more memory-bound. Every generation makes batching and quantization more important, not less.
MIG (Multi-Instance GPU)#
A100 and later can be partitioned into up to 7 isolated instances, each with its own SMs, L2 slice, and memory.
nvidia-smi mig -cgi 1g.10gb,1g.10gb,2g.20gb -C
Use for inference: hosting several small models with hard isolation on one GPU. Downsides: fixed partition sizes, reconfiguration requires draining, and you lose the ability to burst. Generally useful for multi-tenant platforms serving small models (Section XII.05); rarely useful for large models.
What a GPU is bad at#
✗ Branchy code (warp divergence, Section 10)
✗ Pointer chasing / random access
✗ Serial dependencies
✗ Small amounts of work (launch overhead dominates)
✗ Frequent host communication
✗ Anything needing low single-operation latencyAll of these are relevant: decode has serial dependencies, tokenization is branchy, sampling has small work. The parts of an LLM stack that stay on the CPU are exactly the parts on this list.
6. Under the hood#
Where the transistors go, roughly:
CPU (server): ~5% arithmetic units, ~40% cache, ~30% control/prediction, ~25% I/O
GPU (H100): ~35% arithmetic (incl. tensor cores), ~25% memory system,
~15% register file, ~10% scheduling, ~15% I/OThe GPU has more register file than the CPU has L1 cache — because it must hold state for 2,048 concurrent threads per SM to enable latency hiding.
7. Performance implications#
For inference, rank the specs by how much they matter:
1. HBM bandwidth decode speed, directly
2. HBM capacity what fits; how much KV cache; concurrency
3. Tensor FLOPs prefill speed
4. NVLink bandwidth multi-GPU scaling (Section IX)
5. L2 size modest; helps small working sets
6. Clock speed mostly irrelevant; it's a bandwidth gameChoose GPUs by bandwidth and capacity for LLM serving. An H200 (same FLOPs as H100, 43% more bandwidth, 76% more memory) is meaningfully better for inference and identical for training.
8. Production implications#
- Know your fleet’s specs.
nvidia-smi -qandtorch.cuda.get_device_properties. - Check the actual power and clock limits. A GPU in a poorly-cooled chassis runs at reduced
clocks. Monitor
clocks_throttle_reasons. - Enable persistence mode (
nvidia-smi -pm 1) to avoid driver re-initialization latency. - Match GPU to workload: bandwidth-heavy decode → H200/H100; compute-heavy prefill or vision → any high-FLOP part; small models at low QPS → L4/A10G or MIG slices.
- Watch for Xid errors in
dmesg— hardware faults that manifest as mysterious CUDA errors.
9. Common mistakes#
Choosing GPUs by TFLOPs for a decode-bound workload. Bandwidth is what you need.
Assuming “CUDA cores” are comparable to CPU cores. They’re SIMD lanes.
Ignoring memory capacity. A model that doesn’t fit is infinitely slow.
Not checking for throttling. A power-capped GPU runs 20% slower silently.
Using MIG for large models. The partitions are too small.
10. Hands-on exercise#
A. Inventory. Run the property script and record everything in numbers.md. Compute
theoretical FP32 and FP16 peaks from first principles and compare to the spec sheet.
B. Measure the two constants. Measure achieved HBM bandwidth (copy benchmark) and achieved FP16 tensor TFLOPs (large matmul). What fraction of spec do you get for each?
C. Compare generations. Build a table of the GPUs you have access to: bandwidth, capacity, FLOPs, ridge point, and $/hour if known. Compute $/TB/s and $/TFLOP for each. Which is best value for decode? For prefill?
D. Throttling. Run a sustained heavy load and watch nvidia-smi -q -d PERFORMANCE. Can you
induce a power cap? What happens to throughput?
E. If you have MIG-capable hardware: partition a GPU and measure the bandwidth available to one 1g slice vs the full GPU. Is it proportional?
11. Interview questions#
- Describe an SM and what it contains.
- Why has the ridge point increased with each GPU generation, and what does that mean for inference?
- Which GPU spec matters most for decode? For prefill? Why?
- What is MIG and when would you use it for inference?
- Name five things GPUs are bad at, and where they appear in an LLM stack.
- How would you compute a GPU’s theoretical FP32 peak from its specifications?
12. Further reading#
- [FUNDAMENTAL] NVIDIA architecture whitepapers (Volta, Ampere, Hopper, Blackwell)
- [FUNDAMENTAL] Kirk & Hwu, Programming Massively Parallel Processors, ch. 1-4
- [REFERENCE] CUDA C++ Programming Guide, “Hardware Implementation”
- [ESTABLISHED] Jia et al., “Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking”
- Next: 02 — Execution model