Below the API

Projects

Module15 topics~148h
completed

Topics, in order

About this module

Goal: turn every section of the curriculum into something that runs, and that you measured.

Fifteen builds. Each one reuses the code from the ones before it, so keep them in one repo (inference-lab/ is a fine name) with one folder per project and a shared numbers.md.

Index#

#ProjectAfter sectionLevelTimeGPU needed?
01From-scratch inference engineIIIBeginner6-8 hNo
02CPU matmul benchmarkII, IIIBeginner6-8 hNo
03First GPU kernelVIIntermediate8-10 hYes (T4 is enough)
04Tiny transformer engineIV, VIntermediate8-12 hNo
05KV cache ★VIntermediate6-8 hNo
06LLM inference serverVIIIIntermediate8-12 hNo
07Dynamic batchingVIIIIntermediate8-10 hHelps
08Continuous batching ★V, VIIIAdvanced12-18 hHelps
09KV cache manager ★V, XAdvanced10-14 hNo
10Quantized inferenceVIIAdvanced10-14 hHelps
11GPU benchmark suiteVI, XAdvanced8-12 hYes
12Inference gatewayVIII, XIAdvanced12-16 hNo
13Multi-GPU inferenceIXExpert12-16 h2 GPUs, or simulate
14Distributed inferenceIXExpert14-20 hOptional
15Mini inference platformXIIExpert20-30 hOptional

★ = do not skip. These three are the ones interviewers ask you to whiteboard.

The thread#

flowchart TD
  N0["A forward pass is just array math<br/><b>01</b>"]
  N1["Whose speed is set by FLOPs and bytes moved, which you can measure<br/><b>02</b>"]
  N2["On a GPU, where you launch kernels yourself<br/><b>03</b>"]
  N3["A transformer is the same thing with attention<br/><b>04</b>"]
  N4["Made affordable by caching K and V<br/><b>05</b>"]
  N5["Wrapped in an API that streams<br/><b>06</b>"]
  N6["That batches requests to amortize weight reads<br/><b>07</b>"]
  N7["At every step rather than once per batch<br/><b>08</b>"]
  N8["With KV memory managed in pages<br/><b>09</b>"]
  N9["And weights shrunk to move fewer bytes<br/><b>10</b>"]
  N10["On hardware you have characterized yourself<br/><b>11</b>"]
  N11["Behind a front door that protects it<br/><b>12</b>"]
  N12["Split across GPUs when one is not enough<br/><b>13</b>"]
  N13["And across machines when one box is not enough<br/><b>14</b>"]
  N14["All of it run as a platform<br/><b>15</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8 --> N9 --> N10 --> N11 --> N12 --> N13 --> N14

  class N0,N1,N2 neutral
  class N3,N4,N5 io
  class N6,N7,N8 queue
  class N9,N10,N11 compute
  class N12,N13,N14 memory

Rules that apply to every project#

  1. Correctness before speed. Every project has a reference to match (PyTorch, Hugging Face, or your own earlier project). Do not benchmark code that produces different tokens.
  2. Write the prediction first. Before each measurement, write down the number you expect and why. The gap between prediction and measurement is the lesson.
  3. Record everything in numbers.md with the machine, the date, and the command.
  4. Warm up, then measure. Discard the first runs; report median and p95, never a single run.
  5. Synchronize before timing GPU code, or you are timing the launch, not the work (Section VI.06).

Doing the projects in Go#

The projects are designed to be built in Go, with one deliberate exception: the model weights you compare against come from the Python ecosystem, because that is where trained models are published.

ProjectsIn GoWhat still touches Python / CUDA
01, 02, 04, 05 — engine, matmul, transformer, KV cacheEverything: tensors, operators, the forward pass, the cache. Standard library only.A ten-line script, run once, to export reference weights and expected outputs from PyTorch to safetensors (I.03 shows how to read that format in Go).
06, 07, 08, 09, 12 — server, batching, KV manager, gatewayEverything. These are the projects closest to professional Go work: net/http, channels, contexts, schedulers.Nothing. The engine behind your server is your own Project 04/05 model, or any OpenAI-compatible server you put it in front of.
10 — quantizationThe quantizers and the accuracy measurements.Reference perplexity numbers, if you want to compare.
03 — first GPU kernelThe host program, through cgo (see the GPU course, III.02).The kernel itself is CUDA C. That is true in every language.
11, 13, 14 — GPU benchmarks, multi-GPU, distributedLoad generators, harnesses, the KV transfer protocol.The on-GPU model runs in an existing engine (vLLM or PyTorch). The skeletons in these three projects are shown in Python for that reason.
15 — platformThe gateway, router and autoscaler.The operator skeleton is shown with a Python framework; in Go you would use controller-runtime.

Every project asks for a Model you can swap. Define it once and reuse it everywhere:

// Model is the seam between your serving code and whatever does the arithmetic.
type Model interface {
	Prefill(ids []int) (logits []float32, cache KVCache)
	Decode(token int, cache KVCache) (logits []float32, next KVCache)
}

Start with your own pure-Go implementation (slow, fully understood). Later, add a second implementation that forwards to a real engine over HTTP. Your server, batcher and gateway code does not change — which is the point of the interface.

Suggested repo layout#

inference-lab/
  numbers.md
  go.mod
  common/            # timers, load generator, plotting — grows as you go
  p01_engine/
  p02_matmul_bench/
  ...
  p15_platform/

Model to use throughout#

GPT-2 small (124M) is the default from Project 04 onward: it runs on a laptop CPU, its weights are public, and every formula in Section V applies to it unchanged. Where a project benefits from a modern architecture (RoPE, GQA, RMSNorm), swap in a ≤1B model such as Qwen2.5-0.5B or SmolLM2-360M.

↑↓ navigate ↵ open