Your checklist for the whole curriculum: every concept file, every project, every checkpoint.
How to use it
- Tick a file only after you have done its hands-on exercise, not after reading it.
- Tick a project only when every box in its “Done when” section is ticked.
- Tick a checkpoint only if you answered its questions without notes (see ROADMAP.md).
- Put the date next to anything you finish. Momentum is easier to keep when you can see it.
Started: ____-__-__
Target finish: ____-__-__
Hours per week: __I — Fundamentals#
01 — What Is Inference?
07 — Compute vs Memory
09 — CPU vs GPU
Checkpoint A passed without notes
II — Computer Systems#
- 01 — CPU Architecture
- 02 — Caches and the Memory Hierarchy
- 03 — RAM, DRAM, and NUMA
- 04 — SIMD and Vectorization
- 05 — Threads, Processes, and Context Switching
- 06 — Memory Allocation and Virtual Memory
- 07 — Storage and Model Loading
- 08 — Networking Fundamentals
- 09 — PCIe, DMA, and Interconnects
- 10 — OS Scheduling
- 11 — Linux for Inference Engineers
- 12 — Containers, Namespaces, and cgroups
- 13 — Syscalls and Profiling Basics
III — ML Fundamentals#
01 — Vectors and Matrices
05 — Activation Functions
06 — Embeddings
09 — The Transformer
Checkpoint B passed without notes
IV — Neural Network Inference#
- 01 — Tracing One Request Through a Model
- 02 — Computational Graphs and Operators
- 03 — GEMM and GEMV
- 04 — Convolutions
- 05 — Normalization Layers
- 06 — Attention Computation in Practice
- 07 — Tensor Layouts and Memory
- 08 — Kernels and Kernel Launches
- 09 — Operator Fusion
- 10 — Graph Optimization and Compilers
- 11 — Static vs Dynamic Shapes
- 12 — Numerical Precision and Stability in Practice
V — LLM Inference#
03 — Prefill vs Decode
05 — The KV Cache
06 — KV Cache Math
09 — Continuous Batching
10 — PagedAttention
12 — Speculative Decoding
14 — Streaming Inference
Checkpoint C — the big one passed without notes
VI — GPU Computing#
- 01 — GPU Architecture
- 02 — Execution Model: Threads, Warps, Blocks, Grids
- 03 — GPU Memory Hierarchy
- 04 — Bandwidth and the Roofline on GPU
- 05 — Kernel Launch and Host-Device Interaction
- 06 — Streams and Synchronization
- 07 — CUDA Graphs
- 08 — CUDA Programming Fundamentals
- 09 — Coalescing and Access Patterns
- 10 — Occupancy and Warp Divergence
- 11 — Tensor Cores
- 12 — Profiling CUDA
VII — Inference Optimization#
10 — FlashAttention
Checkpoint D passed without notes
VIII — Serving Systems#
- 01 — Anatomy of an Inference Server
- 02 — APIs: HTTP, gRPC, Streaming
- 03 — Queues, Scheduling, and Admission Control
- 04 — Batching in Servers
- 05 — Backpressure, Timeouts, and Retries
- 06 — Load Balancing and Routing
- 07 — Autoscaling and Cold Starts
- 08 — Model Lifecycle: Loading, Warmup, Unloading
- 09 — Multi-Model Serving
- 10 — Versioning, Canary Deployments, and A/B Testing
- 11 — vLLM Architecture
- 12 — SGLang Architecture
- 13 — TGI, Triton, and ONNX Runtime
- 14 — Choosing a Serving Stack
IX — Distributed Inference#
- 01 — Why Distribute (and When Not To)
- 02 — Data Parallelism
- 03 — Tensor Parallelism
- 04 — Pipeline Parallelism
- 05 — Expert Parallelism
- 06 — Sequence and Context Parallelism
- 07 — Collectives and NCCL
- 08 — Interconnects: NVLink, PCIe, InfiniBand, RDMA
- 09 — Communication Cost Math
- 10 — Multi-Node Inference
- 11 — Distributed KV Cache
- 12 — When More GPUs Make Things Worse
X — Memory & Performance#
11 — Case Studies
Checkpoint E passed without notes
XI — Production Inference#
- 01 — SLOs, SLIs, and Latency Budgets
- 02 — Capacity Planning
- 03 — Cost Per Token
- 04 — Autoscaling in Production
- 05 — Fault Tolerance and Failure Modes
- 06 — Multi-Region and Disaster Recovery
- 07 — Model Rollouts
- 08 — Observability for LLM Services
- 09 — Security and Isolation
- 10 — Abuse Prevention, Rate Limiting, and Quotas
- 11 — Scenario: 70B Model, 10,000 Concurrent Users, Fixed Budget
XII — Inference Platform Engineering#
02 — Model Registry
03 — Kubernetes for GPUs
06 — Routing Strategies
07 — Prefix-Aware Routing
08 — Inference Gateways
Checkpoint F passed without notes
XIII — Advanced LLM Inference#
- 01 — MQA, GQA, and MLA
- 02 — Mixture of Experts
- 03 — MoE Serving and Expert Parallelism
- 04 — Long-Context Inference
- 05 — Chunked Prefill
- 06 — Disaggregated Prefill/Decode
- 07 — KV Offloading and Transfer
- 08 — Request Scheduling and Priorities
- 09 — Speculative Decoding Variants
- 10 — Low-Bit Inference: FP4 and Beyond
- 11 — Triton Kernels
- 12 — Compilers: torch.compile and Beyond
XIV — Research & Frontier#
- 01 — How to Read an Inference Paper
- 02 — Attention Research
- 03 — KV Cache Research
- 04 — Quantization Research
- 05 — Decoding Research
- 06 — MoE Research
- 07 — Long-Context Research
- 08 — Scheduling and Systems Research
- 09 — NVIDIA Architecture Trends
- 10 — Alternative Accelerators
- 11 — Memory-Centric and Disaggregated Futures
Projects#
- 01 — NumPy Inference Engine — README
- 02 — CPU Matmul Benchmark — README
- 03 — First GPU Kernel — README
- 04 — Tiny Transformer Engine — README
- 05 — KV Cache — README
- 06 — LLM Inference Server — README
- 07 — Dynamic Batching — README
- 08 — Continuous Batching — README
- 09 — KV Cache Manager — README
- 10 — Quantized Inference — README
- 11 — GPU Benchmark Suite — README
- 12 — Inference Gateway — README
- 13 — Multi-GPU Inference — README
- 14 — Distributed Inference — README
- 15 — Mini Inference Platform — README
Reference pages#
- GPU Hardware & Memory Reference — read once, exercises A-C done
Measurement journal#
-
numbers.mdcreated - CPU: cache sizes, DRAM bandwidth, peak GFLOP/s (Project 02)
- Storage and page-cache read bandwidth (II.07)
- PCIe host↔device bandwidth, pageable and pinned (II.09, Project 03)
- GPU: memory bandwidth, peak TFLOP/s per dtype, ridge point (Project 11)
- Model: tokens/s at batch 1 and batch 64, KV bytes per token (Projects 05, 08)
Tally#
| Block | Done | Total |
|---|---|---|
| Concept files | 171 | |
| Checkpoints | 6 | |
| Projects | 15 |