Below the API

GPU Computing

IntermediateModule12 topics~15h
completed

Topics, in order

About this module

Goal: understand the machine well enough to predict its behavior, read a profiler, and write a kernel when you need to.

You do not need to become a CUDA expert to be an excellent inference engineer. You do need to understand the execution model, the memory hierarchy, and how to profile — because without them you cannot tell a hardware limit from a software bug.

Hardware numbers for every card you are likely to meet — VRAM, memory type, bandwidth, interconnects — live in GPU-HARDWARE.md. Keep it open while reading this section.

Files#

#FileLevelTime
01GPU architectureBeginner75 min
02Execution model: warps, blocks, gridsIntermediate75 min
03Memory hierarchyIntermediate75 min
04Bandwidth and the roofline on GPUIntermediate60 min
05Kernel launch and host-device interactionIntermediate60 min
06Streams and synchronizationIntermediate60 min
07CUDA graphsAdvanced60 min
08CUDA programming fundamentalsAdvanced120 min
09Coalescing and access patternsAdvanced75 min
10Occupancy and warp divergenceAdvanced75 min
11Tensor coresAdvanced75 min
12Profiling CUDAAdvanced90 min

The thread#

flowchart TD
  N0["A GPU is many simple cores plus enormous bandwidth<br/><b>01</b>"]
  N1["Organized as warps of 32 lanes executing in lockstep<br/><b>02</b>"]
  N2["Fed by a memory hierarchy: registers → shared → L2 → HBM<br/><b>03</b>"]
  N3["Whose bandwidth, not FLOPs, usually binds you<br/><b>04</b>"]
  N4["Driven by the CPU via launches (05) on streams<br/><b>06</b>"]
  N5["Which can be pre-recorded to eliminate overhead<br/><b>07</b>"]
  N6["And you can write your own<br/><b>08</b>"]
  N7["If you respect coalescing (09), occupancy, and divergence<br/><b>10</b>"]
  N8["And use tensor cores for anything matmul-shaped<br/><b>11</b>"]
  N9["And measure everything<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8,N9 memory

No GPU?#

Sections 01-07 and 09-11 are readable and the exercises have CPU-analogue or Colab alternatives. Google Colab’s free T4 is sufficient for every exercise in this section — it has tensor cores, Nsight works, and the concepts all transfer. Do not skip this section for lack of an H100.

↑↓ navigate ↵ open