One GPU is a component. Real systems have two opposite problems: a job too big for one GPU, and a GPU too big for one job. This module covers both, and how a cluster scheduler hands GPUs out.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Interconnects and Topology | How are GPUs wired to each other, and why does the wiring matter? |
| 02 | Splitting Work Across GPUs | What are the ways to use several GPUs, and what does each cost? |
| 03 | Sharing One GPU | How do several jobs use one GPU safely? |
| 04 | GPUs in Kubernetes | How does a cluster discover, allocate and monitor GPUs? |