Below the API

Learning path

GPU Engineering

The GPU itself — what is inside the chip, how it is programmed, why it is fast, and how thousands are run together.

5 levels · 8 modules · 32 topics · ~27h 15m completed ·
Start learning →

The path

  1. Foundations

    Build the mental model.

    4 topics · ~2h 30m

    Why GPUs Exist

    A GPU is not a faster CPU. It is a different bargain: give up speed on any single task, and in exchange do an enormous number of simple tasks at the same time. Module overview →

    1. What Is a GPU?30 min
    2. Latency vs Throughput40 min
    3. From Pixels to Tensors35 min
    4. Reading a Spec Sheet45 min
  2. Basic

    Understand the core mechanisms.

    5 topics · ~4h 10m

    Inside the Chip

    Module I said a GPU has "thousands of lanes" and "fast memory". This module opens the lid and names every part, using NVIDIA's H100 as the worked example. Module overview →

    1. SMs, Warps and SIMT1h
    2. The Memory Hierarchy50 min
    3. HBM and Memory Bandwidth45 min
    4. Tensor Cores and Number Formats55 min
    5. Power, Heat and the Host Link40 min
  3. Intermediate

    Learn the optimization techniques.

    10 topics · ~9h 45m

    The Programming Model

    You know what the hardware is. Now: how do you tell it what to do? Module overview →

    1. Kernels, Threads, Blocks, Grids1h
    2. Your First Kernel, Called from Go1h 30m
    3. Host and Device Memory50 min
    4. Asynchronous Execution and Streams50 min
    5. The Software Stack40 min
    Performance

    A GPU can be a thousand times faster than a CPU or barely faster at all. The difference is almost never "better arithmetic". It is whether the program respects four constraints: memory … Module overview →

    1. The Roofline Model1h 10m
    2. Memory Access Patterns1h
    3. Occupancy and Divergence55 min
    4. Launch Overhead and Fusion50 min
    5. Measuring a GPU1h
  4. Advanced

    Study systems at production scale.

    8 topics · ~7h 50m

    GPUs for AI

    Modules II–IV were about GPUs in general. This module applies them to the workload that now buys most GPUs: neural networks, and large language models in particular. Module overview →

    1. Matrix Multiplication on a GPU1h 15m
    2. Precision and Quantization on Hardware1h
    3. Memory Planning for LLMs1h 10m
    4. Training vs Inference on a GPU40 min
    Multi-GPU and Sharing

    One GPU is a component. Real systems have two opposite problems: a job too big for one GPU, and a GPU too big for one job. This module covers both, and how a cluster scheduler hands GPUs … Module overview →

    1. Interconnects and Topology55 min
    2. Splitting Work Across GPUs1h
    3. Sharing One GPU50 min
    4. GPUs in Kubernetes1h
  5. Expert

    Design platforms and read the frontier.

    4 topics · ~3h

    Data Center and Frontier

    The last module zooms out: from one chip to product generations, to competing designs, to the buildings full of them, and finally to where the technology is heading. Module overview →

    1. GPU Generations45 min
    2. Other Accelerators50 min
    3. Racks, Power and Cost50 min
    4. Where This Is Going35 min

Build

Reference

About this path

Understand the GPU completely: what it is, how it is built, how you program it, why it is fast, why it is sometimes slow, and how thousands of them are run together.

This course starts at “what is a GPU?” and ends at “plan the GPUs, network and power for a cluster”. It assumes you can program in Go. It does not assume any hardware, graphics, C or machine-learning background.


How this course works#

Every lesson follows the same shape, so you always know where you are:

  1. The idea in one minute — the whole lesson in a few sentences.
  2. An analogy — something from everyday life with the same shape.
  3. A picture — a hand-drawn diagram of the mechanism.
  4. How it really works — the precise version, with real numbers.
  5. Code — a small Go program you can run, usually without owning a GPU.
  6. Remember this — the three or four facts worth keeping.
  7. Try it and Check yourself — exercises and questions.

Why Go, when GPUs are programmed in C?#

A GPU runs kernels — small functions written in CUDA C (or generated by a compiler). That will not change, and this course shows you those kernels in CUDA C where they matter.

Everything around the kernel is ordinary systems programming, and Go is very good at it:

  • Models and simulators. Most GPU behaviour — warps, coalescing, the roofline, memory planning — can be reproduced in fifty lines of Go and run on a laptop. You learn the mechanism by building it.
  • Talking to the GPU. NVIDIA’s own Kubernetes device plugin, GPU operator and container toolkit are written in Go, on top of the go-nvml bindings. Reading GPU state from Go is a first-class path.
  • Calling kernels. cgo lets a Go program call CUDA C directly. You will do this in module III.
  • Serving. The layers above the GPU — schedulers, gateways, autoscalers — are where most Go engineers meet GPUs, and they are the subject of the sister course, Inference Engineering.
flowchart LR
  subgraph GO["Written in Go"]
    APP["Your service<br/>scheduler, gateway"]
    SIM["Simulators and<br/>capacity models"]
    MON["Monitoring<br/>go-nvml"]
  end
  subgraph C["Written in CUDA C"]
    K["Kernels<br/>the code that runs on the GPU"]
  end
  DRV["NVIDIA driver"]
  GPU[("GPU")]
  APP -->|"cgo"| K
  MON --> DRV
  K --> DRV --> GPU
  class APP,SIM,MON compute
  class K io
  class DRV neutral
  class GPU memory

The path#

flowchart LR
  I["I Why GPUs exist"] --> II["II Inside the chip"]
  II --> III["III Programming model"]
  III --> IV["IV Performance"]
  IV --> V["V GPUs for AI"]
  V --> VI["VI Multi-GPU and sharing"]
  VI --> VII["VII Data center and frontier"]
  class I,II neutral
  class III,IV compute
  class V memory
  class VI queue
  class VII io
ModuleYou will be able toLevel
I — Why GPUs ExistExplain what a GPU is, why it beats a CPU at some work and loses at other work, and read a spec sheetBeginner
II — Inside the ChipDescribe SMs, warps, the memory hierarchy, HBM and tensor cores, with numbersBeginner → Intermediate
III — The Programming ModelWrite a CUDA kernel, call it from Go, and manage host/device memory and streamsIntermediate
IV — PerformancePredict speed with the roofline, fix slow memory access, and profile a GPUIntermediate
V — GPUs for AIExplain how matmul, precision and memory planning decide AI performanceIntermediate → Advanced
VI — Multi-GPU and SharingReason about NVLink, collectives, MPS/MIG and GPUs in KubernetesAdvanced
VII — Data Center and FrontierCompare GPU generations and other accelerators; plan power, cooling and costAdvanced
ProjectsBuild five Go programs that make the ideas stickAll

What you need#

  • Go 1.22 or newer.
  • A terminal.
  • No GPU required for about 90% of the course. Lessons that need real hardware say so and give a no-GPU alternative. A free Colab T4 or any rented NVIDIA card covers the rest.

A promise about numbers#

GPU specifications change every year. This course teaches you the mechanisms, which do not change, and uses real parts (mostly NVIDIA’s H100, because it is the best documented) as worked examples. When a number is a vendor announcement rather than something measured, the lesson says so. Always check the current data sheet before spending money.

Product facts were last checked on 3 October 2026: Rubin (HBM4) had been shipping since August, Blackwell and Blackwell Ultra made up most new capacity, and Hopper remained the largest installed base. VII.01 holds the current table.

Where to go next#

↑↓ navigate ↵ open