From “what is a matrix multiply” to “design the serving stack for a 70B model with 10,000 concurrent users on a fixed GPU budget.”
This repository is a textbook + lab + roadmap for inference engineering: the discipline of running trained machine-learning models — especially large language models — as fast, cheap, and reliable production systems.
What this repository teaches#
Inference engineering sits at the intersection of four fields that are usually taught separately:
Machine learning Computer architecture
(what the model is) (what the hardware does)
\ /
\ /
+--------------------+
| INFERENCE |
| ENGINEERING |
+--------------------+
/ \
/ \
Distributed systems Production operations
(many machines) (SLOs, cost, reliability)Most people learn one corner and stall. A model researcher can explain attention but cannot explain why their server falls over at 40 concurrent users. A backend engineer can build a bulletproof API but cannot explain why a 7B model at batch size 1 leaves 97% of the GPU idle. This curriculum builds all four corners in the order in which they actually depend on each other.
By the end you should be able to:
- Read a model config and predict its memory footprint, KV cache growth, and peak throughput before deploying it.
- Look at a slow inference service and systematically determine whether it is memory-bandwidth bound, compute bound, queueing bound, or bound by something silly like tokenizer overhead.
- Explain — and implement — continuous batching, paged attention, prefix caching, speculative decoding, tensor parallelism, and quantization.
- Write a CUDA kernel and a Triton kernel, and profile both.
- Design a multi-tenant inference platform with routing, autoscaling, quotas, and cost attribution.
- Read a frontier inference paper and judge whether it is production-relevant or a lab curiosity.
Who this is for#
| You are | This will |
|---|---|
| A backend / platform / DevOps engineer moving into AI infra | Give you the ML and GPU foundations you’re missing, using systems vocabulary you already have |
| An ML engineer who trains models | Give you the systems, GPU, and production layers that training rarely teaches |
| A student targeting inference / performance roles | Give you a complete, sequential path with projects and interview questions |
| An SRE inheriting a GPU fleet | Give you the mental models to reason about capacity, cost, and failure |
Assumed: you can program in Go, you understand basic software engineering, you can use a terminal, and you know what an HTTP request is.
Not assumed: any ML, linear algebra beyond high school, GPU knowledge, CUDA, or distributed systems.
Prerequisites#
Practical, not theoretical:
- Go 1.22+. Every runnable example is a single-file Go program using only the standard
library: save it,
go runit. - Python 3.10+ is useful later, but only to run existing engines (vLLM, PyTorch) that your Go code talks to. You will not need to write it.
- Comfort with Linux/macOS command line.
- Optional but strongly recommended: access to any NVIDIA GPU. A free Colab T4, a rented A10G/L4, or a consumer RTX card is enough for 90% of the exercises. Sections VI, IX, and several projects have CPU-only fallbacks marked [No GPU? Do this instead].
git,dockerfor the later sections.
You do not need an H100. You need curiosity and the willingness to measure things.
Why Go for inference engineering#
The arithmetic of a model runs in CUDA kernels driven by C++ and Python engines. That is not going to change, and this course teaches you how those engines work from the inside.
But most inference engineering is everything around the kernel: schedulers, batchers, KV-cache managers, routers, gateways, rate limiters, autoscalers, load generators. That is concurrent systems programming, and it is what Go is for. So in this course:
- Mechanisms are built in Go. A transformer forward pass, a KV cache, paged attention, continuous batching, speculative decoding, quantization, a fair scheduler, a consistent-hash router — each as a small program you can read in one sitting and run on a laptop.
- Real engines are measured from Go. Clients that drive any OpenAI-compatible server and measure time-to-first-token, inter-token latency and goodput correctly.
- Framework internals stay in their own language. Where a lesson is specifically about a PyTorch, CUDA or Triton API (mostly module VI and parts of X and XIII), the snippet is shown as that API really looks. Translating it would teach you something that does not exist.
Every complete Go program in the lessons is compiled and vetted by tools/gocheck, so what you
copy is known to build.
For the hardware underneath all of this, see the sister course GPU Engineering.
Curriculum map#
flowchart LR
subgraph F["Foundations"]
direction TB
I["I Fundamentals"]
II["II Computer Systems"]
III["III ML Fundamentals"]
IV["IV NN Inference"]
end
subgraph C["Core"]
direction TB
V["V LLM Inference"]
VI["VI GPU Computing"]
end
subgraph B["Build"]
direction TB
VII["VII Optimization"]
VIII["VIII Serving Systems"]
end
subgraph S["Scale"]
direction TB
IX["IX Distributed"]
X["X Memory & Performance"]
end
subgraph O["Operate"]
direction TB
XI["XI Production"]
XII["XII Platform"]
end
subgraph R["Frontier"]
direction TB
XIII["XIII Advanced LLM"]
XIV["XIV Research"]
end
F --> C --> B --> S --> O --> R
class I,II,III,IV neutral
class V,VI memory
class VII,VIII compute
class IX,X queue
class XI,XII io
class XIII,XIV warnI Fundamentals ............... the vocabulary and the physics
II Computer Systems .......... CPU, memory, OS, network, PCIe
III ML Fundamentals ........... linear algebra -> transformers -> number formats
IV NN Inference .............. the forward pass as an engineering artifact
V LLM Inference ............. prefill/decode, KV cache, batching, PagedAttention [CORE]
VI GPU Computing ............. architecture, CUDA, memory hierarchy, profiling [CORE]
VII Inference Optimization .... quantization, fusion, FlashAttention, spec decoding
VIII Serving Systems ........... vLLM/SGLang/TGI/Triton architectures, APIs, queues
IX Distributed Inference ..... TP/PP/EP, NCCL, NVLink, when more GPUs hurt
X Memory & Performance ...... roofline, profiling, "why is my system slow?"
XI Production Inference ...... SLOs, capacity, cost/token, observability, security
XII Platform Engineering ...... registry, K8s GPU scheduling, routing, gateways
XIII Advanced LLM Inference .... MoE, MLA, disaggregation, chunked prefill, kernels
XIV Research & Frontier ....... how to read the field and judge what mattersEach section has its own README.md with a per-file index, difficulty labels, and estimated
time.
Recommended learning order#
The main path (sequential)#
Read I → II → III → IV → V → VI → VII → VIII → IX → X → XI → XII → XIII → XIV.
This is the intended order and every file assumes only what came before it. See ROADMAP.md for the full dependency graph, time estimates, and checkpoints.
If you are in a hurry (“I need to be useful in 3 weeks”)#
I (all) -> III.11-12 -> V (all) -> VI.01-05 -> VII.01-02, 10
-> VIII.01-04, 11 -> X.01-03 -> XI.01-03This is the “operate an LLM serving stack competently” path. It skips the deep hardware and distributed material, which you can return to.
If you already know ML#
Skim III, read IV.03/IV.07/IV.09 closely, then go straight to V and VI. Do not skip V — it is where most ML people have the largest gaps.
If you already know systems/GPUs#
Skim II and VI, read III fully (do not skip: the attention math matters), then V, VII, IX.
How to use this repository#
Read sequentially, but do the exercises. Every concept file ends with a hands-on exercise. The exercises are where the learning actually happens. Reading alone produces the illusion of understanding — inference engineering punishes that illusion very quickly, because the hardware does not care what you believe.
Keep a measurement journal. Every time a file asks you to measure something, record the number and your machine. Six months later, “an H100 does ~3.3 TB/s of HBM bandwidth” will be a fact you own rather than one you looked up.
Do the projects in order.
projects/contains fifteen builds that mirror the sections. Project N is designed to be doable after section N-ish. Each project README lists its exact prerequisites.Track progress in PROGRESS.md. Check items off. It is a real checklist of every file and project.
When you meet an unfamiliar term, check GLOSSARY.md first. It is alphabetical and deliberately blunt.
Read the “Common mistakes” section of every file twice. Those sections encode the mistakes that cost real teams real money.
File format#
Every major concept file follows the same twelve-part structure:
1. What is it? 7. Performance implications
2. Why does it exist? 8. Production implications
3. Simple analogy 9. Common mistakes
4. Tiny example 10. Hands-on exercise
5. Technical explanation 11. Interview questions
6. Under the hood 12. Further readingShorter connective files (indexes, methodology notes) use a lighter structure. Every file carries a header:
Project progression#
| # | Project | After section | What it proves |
|---|---|---|---|
| 01 | NumPy inference engine | III | You understand a forward pass end to end |
| 02 | CPU matmul benchmark | II, III | You can measure FLOPs and bandwidth |
| 03 | First GPU kernel | VI | You can write and launch CUDA |
| 04 | Tiny transformer engine | IV, V | You understand transformer inference mechanically |
| 05 | KV cache | V | You understand the single most important optimization |
| 06 | LLM inference server | VIII | You can wrap a model in a real API |
| 07 | Dynamic batching | VIII | You understand throughput/latency tradeoffs |
| 08 | Continuous batching | V, VIII | You understand modern LLM serving |
| 09 | KV cache manager | V, X | You understand paged memory management |
| 10 | Quantized inference | VII | You can trade accuracy for speed deliberately |
| 11 | GPU benchmark suite | VI, X | You can characterize hardware |
| 12 | Inference gateway | VIII, XI | You can build the production front door |
| 13 | Multi-GPU inference | IX | You understand tensor parallelism concretely |
| 14 | Distributed inference | IX | You understand multi-node and collectives |
| 15 | Mini inference platform | XII | You can build the whole thing |
A note on honesty#
This curriculum distinguishes three tiers of knowledge and labels them:
- [FUNDAMENTAL] — physics and math. True in 1995, true in 2035. Memory bandwidth, arithmetic intensity, Amdahl’s law.
- [ESTABLISHED] — production-proven engineering. Continuous batching, PagedAttention, FP8/INT8 quantization, tensor parallelism. Deployed at scale by many organizations.
- [EMERGING] — promising but not universally production-ready. Some speculative decoding variants, aggressive low-bit formats, disaggregated serving in its newer forms.
Numbers in this curriculum (bandwidths, prices, model sizes) are illustrative and drift. The methods for computing them do not. Always re-derive with current hardware specs.
Repository files#
- ROADMAP.md — the full I→XIV path with dependencies and checkpoints
- GLOSSARY.md — every term, defined bluntly
- GPU-HARDWARE.md — GPU memory types, spec sheet, memory budgets,
nvidia-smicheat sheet - RESOURCES.md — papers, docs, blogs, talks, organized by section
- PROGRESS.md — your checklist
Currency#
The mechanisms in this curriculum are stable; product facts are not. Hardware tables, engine recommendations and project statuses were last reviewed on 3 October 2026. What that review changed:
- Hardware. Rubin (HBM4, ~22 TB/s) has shipped since August 2026 and Blackwell Ultra (B300) since 2025; both are in GPU-HARDWARE.md and XIV.09.
- Engines. Hugging Face archived TGI in March 2026; VIII.13 and VIII.14 now treat it as a migration source. vLLM and SGLang remain the reference engines.
- Platform. Kubernetes DRA and the Gateway API Inference Extension moved from “emerging” to stable APIs; llm-d and NVIDIA Dynamo are the open-source layers above the engine.
- Newer papers and feeds are under “2025-2026 additions” in RESOURCES.md.
Sister paths#
- Go Engineering — the language this curriculum’s code is written in: memory, the runtime, concurrency, and AI building blocks in Go.
- GPU Engineering — the device underneath all of this.
- Observability Engineering — measuring it in production: SLOs, GPU and engine telemetry, tracing, cost per token.
The one idea to carry through everything#
If you remember nothing else from this repository, remember this:
Modern LLM inference is not limited by how fast your GPU can multiply. It is limited by how fast your GPU can move bytes.
Almost every technique in sections V, VII, IX, X, and XIII is, at bottom, a scheme to move fewer bytes or to do more arithmetic per byte moved. Once you see that, the field stops being a list of tricks and becomes a single coherent story.
Start with I-fundamentals/.