Goal: understand the machine well enough to predict its behavior, read a profiler, and write a kernel when you need to.
You do not need to become a CUDA expert to be an excellent inference engineer. You do need to understand the execution model, the memory hierarchy, and how to profile — because without them you cannot tell a hardware limit from a software bug.
Hardware numbers for every card you are likely to meet — VRAM, memory type, bandwidth, interconnects — live in GPU-HARDWARE.md. Keep it open while reading this section.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | GPU architecture | Beginner | 75 min |
| 02 | Execution model: warps, blocks, grids | Intermediate | 75 min |
| 03 | Memory hierarchy | Intermediate | 75 min |
| 04 | Bandwidth and the roofline on GPU | Intermediate | 60 min |
| 05 | Kernel launch and host-device interaction | Intermediate | 60 min |
| 06 | Streams and synchronization | Intermediate | 60 min |
| 07 | CUDA graphs | Advanced | 60 min |
| 08 | CUDA programming fundamentals | Advanced | 120 min |
| 09 | Coalescing and access patterns | Advanced | 75 min |
| 10 | Occupancy and warp divergence | Advanced | 75 min |
| 11 | Tensor cores | Advanced | 75 min |
| 12 | Profiling CUDA | Advanced | 90 min |
The thread#
flowchart TD N0["A GPU is many simple cores plus enormous bandwidth<br/><b>01</b>"] N1["Organized as warps of 32 lanes executing in lockstep<br/><b>02</b>"] N2["Fed by a memory hierarchy: registers → shared → L2 → HBM<br/><b>03</b>"] N3["Whose bandwidth, not FLOPs, usually binds you<br/><b>04</b>"] N4["Driven by the CPU via launches (05) on streams<br/><b>06</b>"] N5["Which can be pre-recorded to eliminate overhead<br/><b>07</b>"] N6["And you can write your own<br/><b>08</b>"] N7["If you respect coalescing (09), occupancy, and divergence<br/><b>10</b>"] N8["And use tensor cores for anything matmul-shaped<br/><b>11</b>"] N9["And measure everything<br/><b>12</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 class N0,N1 neutral class N2,N3 io class N4,N5 queue class N6,N7 compute class N8,N9 memory
No GPU?#
Sections 01-07 and 09-11 are readable and the exercises have CPU-analogue or Colab alternatives. Google Colab’s free T4 is sufficient for every exercise in this section — it has tensor cores, Nsight works, and the concepts all transfer. Do not skip this section for lack of an H100.