Below the API

Neural Network Inference

BasicModule12 topics~13h 30m
completed

Topics, in order

About this module

Goal: turn “a model is math” into “a model is a program that a specific machine executes.”

Section III taught the mathematics. This section is about what actually runs: operators, kernels, layouts, graphs, and the transformations a compiler applies. It is the bridge between the model and the hardware.

Files#

#FileLevelTime
01Tracing one request through a modelBeginner75 min
02Computational graphs and operatorsIntermediate60 min
03GEMM and GEMVIntermediate90 min
04ConvolutionsIntermediate45 min
05Normalization layersIntermediate45 min
06Attention computation in practiceAdvanced90 min
07Tensor layouts and memoryAdvanced75 min
08Kernels and kernel launchesIntermediate60 min
09Operator fusionAdvanced75 min
10Graph optimization and compilersAdvanced75 min
11Static vs dynamic shapesAdvanced60 min
12Numerical precision and stability in practiceAdvanced60 min

The thread#

flowchart TD
  N0["A model is a graph of operators<br/><b>02</b>"]
  N1["Each operator becomes one or more kernels<br/><b>08</b>"]
  N2["Kernels read tensors whose LAYOUT determines their speed<br/><b>07</b>"]
  N3["Most kernels are GEMM (03) or attention (06) or memory-bound glue<br/><b>05</b>"]
  N4["The glue should be fused away<br/><b>09</b>"]
  N5["A compiler can do that automatically — if shapes cooperate<br/><b>10, 11</b>"]
  N6["And all of it must stay numerically sane<br/><b>12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6

  class N0,N1 neutral
  class N2 io
  class N3,N4 queue
  class N5 compute
  class N6 memory

↑↓ navigate ↵ open