Below the API

Machine Learning Fundamentals

FoundationsModule12 topics~15h 30m
completed

Topics, in order

About this module

Goal: teach exactly the ML you need to reason about inference — no more, no less.

This is not an ML course. There is no training theory, no optimization, no generalization bounds. What is here: the linear algebra you will do by hand, the architecture you will trace through, and the numerics you will trade away for speed.

If you already know ML, do not skip files 07, 11, and 12 — softmax stability, number formats, and quantization are where ML people most often have gaps that matter for inference.

Files#

#FileLevelTime
01Vectors and matricesBeginner60 min
02Matrix multiplication by handBeginner75 min
03Tensors, shapes, and broadcastingBeginner60 min
04Neural networks from scratchBeginner90 min
05Activation functionsBeginner45 min
06EmbeddingsBeginner60 min
07Softmax and numerical stabilityIntermediate75 min
08Attention from first principlesIntermediate120 min
09The transformerIntermediate120 min
10Training vs inference: the mathIntermediate45 min
11Number formats: FP32, FP16, BF16, FP8Intermediate90 min
12Quantization fundamentalsIntermediate90 min

The thread#

flowchart TD
  N0["Numbers in arrays<br/><b>01, 03</b>"]
  N1["Multiplied by learned arrays<br/><b>02</b>"]
  N2["Stacked with nonlinearity between them<br/><b>04, 05</b>"]
  N3["Text becomes vectors<br/><b>06</b>"]
  N4["Scores become probabilities<br/><b>07</b>"]
  N5["Tokens look at each other<br/><b>08</b>"]
  N6["Assembled into a transformer<br/><b>09</b>"]
  N7["Which we run forward only<br/><b>10</b>"]
  N8["In as few bits as we can get away with<br/><b>11, 12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint B#

  1. Multiply a 2×3 by a 3×2 matrix by hand, showing every dot product.
  2. Explain attention in three sentences with no equations.
  3. Why is softmax computed with a max subtraction? What breaks without it?
  4. What does BF16 trade away relative to FP16, and why does that trade matter?
  5. Write the parameter count formula for a decoder-only transformer and apply it to a real model.
  6. What is the difference between per-tensor and per-channel quantization, and why does the latter usually win?

↑↓ navigate ↵ open