Modules II–IV were about GPUs in general. This module applies them to the workload that now buys most GPUs: neural networks, and large language models in particular.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Matrix Multiplication on a GPU | How is the single most important operation made fast? |
| 02 | Precision and Quantization on Hardware | What do fewer bits buy, in bytes and FLOPs? |
| 03 | Memory Planning for LLMs | What fills an AI GPU’s memory, and how many users fit? |
| 04 | Training vs Inference on a GPU | Why do the two stress the same hardware so differently? |
The sister course, Inference Engineering, takes these ideas up into full serving systems. This module gives you the hardware-side half.