Below the API

Learning path

Inference Engineering

How modern LLM inference works — from a single matrix multiply to a multi-tenant serving platform.

5 levels · 15 modules · 186 topics · ~354h 45m completed ·
Start learning →

The path

  1. Foundations

    Build the mental model.

    36 topics · ~40h 45m

    Fundamentals

    Goal of this section: give you the vocabulary and the physics of inference. By the end you will be able to read any inference discussion without getting lost, and — more importantly — you … Module overview →

    1. What Is Inference?45 min
    2. Training vs Inference45 min
    3. Model Anatomy: Parameters, Weights, Activations1h
    4. Tensors and the Forward Pass1h
    5. Latency, Throughput, and the Metrics That Matter1h 15m
    6. Batching: The Central Tradeoff1h
    7. Compute vs Memory1h 15m
    8. FLOPs, Bandwidth, and Arithmetic Intensity1h 30m
    9. CPU vs GPU1h
    10. Why Inference Is Not Normal Backend Engineering1h
    11. Why Inference Gets Expensive at Scale1h
    Computer Systems Foundations

    Goal: give you the systems knowledge that inference engineering assumes and rarely teaches. Every file answers "why does an inference engineer care about this?" — if a systems topic doesn't … Module overview →

    1. CPU Architecture1h
    2. Caches and the Memory Hierarchy1h 15m
    3. RAM, DRAM, and NUMA1h
    4. SIMD and Vectorization1h
    5. Threads, Processes, and Context Switching1h
    6. Memory Allocation and Virtual Memory1h 15m
    7. Storage and Model Loading1h
    8. Networking Fundamentals1h
    9. PCIe, DMA, and Interconnects1h
    10. OS Scheduling45 min
    11. Linux for Inference Engineers1h 15m
    12. Containers, Namespaces, and cgroups1h
    13. Syscalls and Profiling Basics1h 15m
    Machine Learning Fundamentals

    Goal: teach exactly the ML you need to reason about inference — no more, no less. Module overview →

    1. Vectors and Matrices1h
    2. Matrix Multiplication by Hand1h 15m
    3. Tensors, Shapes, and Broadcasting1h
    4. Neural Networks from Scratch1h 30m
    5. Activation Functions45 min
    6. Embeddings1h
    7. Softmax and Numerical Stability1h 15m
    8. Attention from First Principles2h
    9. The Transformer2h
    10. Training vs Inference: The Math45 min
    11. Number Formats: FP32, FP16, BF16, FP8, INT8, INT41h 30m
    12. Quantization Fundamentals1h 30m
  2. Intermediate

    Learn the optimization techniques.

    40 topics · ~49h

    GPU Computing

    Goal: understand the machine well enough to predict its behavior, read a profiler, and write a kernel when you need to. Module overview →

    1. GPU Architecture1h 15m
    2. Execution Model: Threads, Warps, Blocks, Grids1h 15m
    3. GPU Memory Hierarchy1h 15m
    4. Bandwidth and the Roofline on GPU1h
    5. Kernel Launch and Host-Device Interaction1h
    6. Streams and Synchronization1h
    7. CUDA Graphs1h
    8. CUDA Programming Fundamentals2h
    9. Coalescing and Access Patterns1h 15m
    10. Occupancy and Warp Divergence1h 15m
    11. Tensor Cores1h 15m
    12. Profiling CUDA1h 30m
    Inference Optimization

    Goal: a catalogue of the techniques that make inference faster, each analyzed the same way: Module overview →

    1. Optimization Methodology1h
    2. Quantization Overview and Decision Guide1h 30m
    3. Weight-Only Quantization: GPTQ and AWQ1h 30m
    4. Activation Quantization and SmoothQuant1h 15m
    5. FP8 and INT8 Inference1h 15m
    6. INT4 and Low-Bit Inference1h 15m
    7. Kernel and Operator Fusion1h
    8. CUDA Graphs in Serving45 min
    9. TensorRT and TensorRT-LLM1h 15m
    10. FlashAttention2h
    11. KV Cache Optimization1h 30m
    12. KV Cache Quantization1h
    13. Speculative Decoding in Practice1h 15m
    14. Pruning, Distillation, and Architecture Optimization1h
    Inference Serving Systems

    Goal: understand how models become services — and, more importantly, understand why real inference servers are built the way they are. Module overview →

    1. Anatomy of an Inference Server1h 15m
    2. APIs: HTTP, gRPC, Streaming1h
    3. Queues, Scheduling, and Admission Control1h 30m
    4. Batching in Servers1h
    5. Backpressure, Timeouts, and Retries1h 15m
    6. Load Balancing and Routing1h 15m
    7. Autoscaling and Cold Starts1h 15m
    8. Model Lifecycle: Loading, Warmup, Unloading1h
    9. Multi-Model Serving1h 15m
    10. Versioning, Canary Deployments, and A/B Testing1h
    11. vLLM Architecture1h 30m
    12. SGLang Architecture1h
    13. TGI, Triton, and ONNX Runtime1h 15m
    14. Choosing a Serving Stack1h
  3. Advanced

    Study systems at production scale.

    34 topics · ~41h 45m

    Distributed Inference

    Goal: understand what happens when one GPU isn't enough — and, equally important, when adding GPUs makes things worse. Module overview →

    1. Why Distribute (and When Not To)1h
    2. Data Parallelism45 min
    3. Tensor Parallelism2h
    4. Pipeline Parallelism1h 15m
    5. Expert Parallelism1h 15m
    6. Sequence and Context Parallelism1h
    7. Collectives and NCCL1h 15m
    8. Interconnects: NVLink, PCIe, InfiniBand, RDMA1h 15m
    9. Communication Cost Math1h 30m
    10. Multi-Node Inference1h 15m
    11. Distributed KV Cache1h
    12. When More GPUs Make Things Worse1h 15m
    Memory and Performance Engineering

    Goal: answer "why is my inference system slow?" systematically, every time, without guessing. Module overview →

    1. The Performance Debugging Methodology1h 30m
    2. The Roofline Model in Practice1h
    3. Memory-Bound vs Compute-Bound1h
    4. GPU Utilization Is a Lie1h
    5. Memory Fragmentation and Allocators1h 15m
    6. OOM: Causes and Cures1h
    7. Benchmarking Inference Correctly1h 30m
    8. GPU Profiling with Nsight1h 30m
    9. CPU and Python Profiling1h
    10. Metrics, Prometheus, and Dashboards1h 15m
    11. Case Studies1h 30m
    Production Inference Engineering

    Goal: run an inference service that meets commitments, costs what you planned, and doesn't wake you at 3 a.m. Module overview →

    1. SLOs, SLIs, and Latency Budgets1h 15m
    2. Capacity Planning1h 30m
    3. Cost Per Token1h 15m
    4. Autoscaling in Production1h
    5. Fault Tolerance and Failure Modes1h 15m
    6. Multi-Region and Disaster Recovery1h
    7. Model Rollouts1h
    8. Observability for LLM Services1h
    9. Security and Isolation1h 15m
    10. Abuse Prevention, Rate Limiting, and Quotas1h
    11. Scenario: 70B Model, 10,000 Concurrent Users, Fixed Budget2h

Build

Reference

About this path

From “what is a matrix multiply” to “design the serving stack for a 70B model with 10,000 concurrent users on a fixed GPU budget.”

This repository is a textbook + lab + roadmap for inference engineering: the discipline of running trained machine-learning models — especially large language models — as fast, cheap, and reliable production systems.


What this repository teaches#

Inference engineering sits at the intersection of four fields that are usually taught separately:

   Machine learning            Computer architecture
   (what the model is)         (what the hardware does)
            \                        /
             \                      /
              +--------------------+
              |  INFERENCE         |
              |  ENGINEERING       |
              +--------------------+
             /                      \
            /                        \
   Distributed systems          Production operations
   (many machines)              (SLOs, cost, reliability)

Most people learn one corner and stall. A model researcher can explain attention but cannot explain why their server falls over at 40 concurrent users. A backend engineer can build a bulletproof API but cannot explain why a 7B model at batch size 1 leaves 97% of the GPU idle. This curriculum builds all four corners in the order in which they actually depend on each other.

By the end you should be able to:

  • Read a model config and predict its memory footprint, KV cache growth, and peak throughput before deploying it.
  • Look at a slow inference service and systematically determine whether it is memory-bandwidth bound, compute bound, queueing bound, or bound by something silly like tokenizer overhead.
  • Explain — and implement — continuous batching, paged attention, prefix caching, speculative decoding, tensor parallelism, and quantization.
  • Write a CUDA kernel and a Triton kernel, and profile both.
  • Design a multi-tenant inference platform with routing, autoscaling, quotas, and cost attribution.
  • Read a frontier inference paper and judge whether it is production-relevant or a lab curiosity.

Who this is for#

You areThis will
A backend / platform / DevOps engineer moving into AI infraGive you the ML and GPU foundations you’re missing, using systems vocabulary you already have
An ML engineer who trains modelsGive you the systems, GPU, and production layers that training rarely teaches
A student targeting inference / performance rolesGive you a complete, sequential path with projects and interview questions
An SRE inheriting a GPU fleetGive you the mental models to reason about capacity, cost, and failure

Assumed: you can program in Go, you understand basic software engineering, you can use a terminal, and you know what an HTTP request is.

Not assumed: any ML, linear algebra beyond high school, GPU knowledge, CUDA, or distributed systems.


Prerequisites#

Practical, not theoretical:

  • Go 1.22+. Every runnable example is a single-file Go program using only the standard library: save it, go run it.
  • Python 3.10+ is useful later, but only to run existing engines (vLLM, PyTorch) that your Go code talks to. You will not need to write it.
  • Comfort with Linux/macOS command line.
  • Optional but strongly recommended: access to any NVIDIA GPU. A free Colab T4, a rented A10G/L4, or a consumer RTX card is enough for 90% of the exercises. Sections VI, IX, and several projects have CPU-only fallbacks marked [No GPU? Do this instead].
  • git, docker for the later sections.

You do not need an H100. You need curiosity and the willingness to measure things.


Why Go for inference engineering#

The arithmetic of a model runs in CUDA kernels driven by C++ and Python engines. That is not going to change, and this course teaches you how those engines work from the inside.

But most inference engineering is everything around the kernel: schedulers, batchers, KV-cache managers, routers, gateways, rate limiters, autoscalers, load generators. That is concurrent systems programming, and it is what Go is for. So in this course:

  • Mechanisms are built in Go. A transformer forward pass, a KV cache, paged attention, continuous batching, speculative decoding, quantization, a fair scheduler, a consistent-hash router — each as a small program you can read in one sitting and run on a laptop.
  • Real engines are measured from Go. Clients that drive any OpenAI-compatible server and measure time-to-first-token, inter-token latency and goodput correctly.
  • Framework internals stay in their own language. Where a lesson is specifically about a PyTorch, CUDA or Triton API (mostly module VI and parts of X and XIII), the snippet is shown as that API really looks. Translating it would teach you something that does not exist.

Every complete Go program in the lessons is compiled and vetted by tools/gocheck, so what you copy is known to build.

For the hardware underneath all of this, see the sister course GPU Engineering.


Curriculum map#

flowchart LR
  subgraph F["Foundations"]
    direction TB
    I["I Fundamentals"]
    II["II Computer Systems"]
    III["III ML Fundamentals"]
    IV["IV NN Inference"]
  end
  subgraph C["Core"]
    direction TB
    V["V LLM Inference"]
    VI["VI GPU Computing"]
  end
  subgraph B["Build"]
    direction TB
    VII["VII Optimization"]
    VIII["VIII Serving Systems"]
  end
  subgraph S["Scale"]
    direction TB
    IX["IX Distributed"]
    X["X Memory & Performance"]
  end
  subgraph O["Operate"]
    direction TB
    XI["XI Production"]
    XII["XII Platform"]
  end
  subgraph R["Frontier"]
    direction TB
    XIII["XIII Advanced LLM"]
    XIV["XIV Research"]
  end
  F --> C --> B --> S --> O --> R

  class I,II,III,IV neutral
  class V,VI memory
  class VII,VIII compute
  class IX,X queue
  class XI,XII io
  class XIII,XIV warn
I    Fundamentals ............... the vocabulary and the physics
II   Computer Systems .......... CPU, memory, OS, network, PCIe
III  ML Fundamentals ........... linear algebra -> transformers -> number formats
IV   NN Inference .............. the forward pass as an engineering artifact
V    LLM Inference ............. prefill/decode, KV cache, batching, PagedAttention  [CORE]
VI   GPU Computing ............. architecture, CUDA, memory hierarchy, profiling      [CORE]
VII  Inference Optimization .... quantization, fusion, FlashAttention, spec decoding
VIII Serving Systems ........... vLLM/SGLang/TGI/Triton architectures, APIs, queues
IX   Distributed Inference ..... TP/PP/EP, NCCL, NVLink, when more GPUs hurt
X    Memory & Performance ...... roofline, profiling, "why is my system slow?"
XI   Production Inference ...... SLOs, capacity, cost/token, observability, security
XII  Platform Engineering ...... registry, K8s GPU scheduling, routing, gateways
XIII Advanced LLM Inference .... MoE, MLA, disaggregation, chunked prefill, kernels
XIV  Research & Frontier ....... how to read the field and judge what matters

Each section has its own README.md with a per-file index, difficulty labels, and estimated time.


The main path (sequential)#

Read I → II → III → IV → V → VI → VII → VIII → IX → X → XI → XII → XIII → XIV.

This is the intended order and every file assumes only what came before it. See ROADMAP.md for the full dependency graph, time estimates, and checkpoints.

If you are in a hurry (“I need to be useful in 3 weeks”)#

I (all)  ->  III.11-12  ->  V (all)  ->  VI.01-05  ->  VII.01-02, 10
         ->  VIII.01-04, 11  ->  X.01-03  ->  XI.01-03

This is the “operate an LLM serving stack competently” path. It skips the deep hardware and distributed material, which you can return to.

If you already know ML#

Skim III, read IV.03/IV.07/IV.09 closely, then go straight to V and VI. Do not skip V — it is where most ML people have the largest gaps.

If you already know systems/GPUs#

Skim II and VI, read III fully (do not skip: the attention math matters), then V, VII, IX.


How to use this repository#

  1. Read sequentially, but do the exercises. Every concept file ends with a hands-on exercise. The exercises are where the learning actually happens. Reading alone produces the illusion of understanding — inference engineering punishes that illusion very quickly, because the hardware does not care what you believe.

  2. Keep a measurement journal. Every time a file asks you to measure something, record the number and your machine. Six months later, “an H100 does ~3.3 TB/s of HBM bandwidth” will be a fact you own rather than one you looked up.

  3. Do the projects in order. projects/ contains fifteen builds that mirror the sections. Project N is designed to be doable after section N-ish. Each project README lists its exact prerequisites.

  4. Track progress in PROGRESS.md. Check items off. It is a real checklist of every file and project.

  5. When you meet an unfamiliar term, check GLOSSARY.md first. It is alphabetical and deliberately blunt.

  6. Read the “Common mistakes” section of every file twice. Those sections encode the mistakes that cost real teams real money.

File format#

Every major concept file follows the same twelve-part structure:

1.  What is it?              7.  Performance implications
2.  Why does it exist?       8.  Production implications
3.  Simple analogy           9.  Common mistakes
4.  Tiny example             10. Hands-on exercise
5.  Technical explanation    11. Interview questions
6.  Under the hood           12. Further reading

Shorter connective files (indexes, methodology notes) use a lighter structure. Every file carries a header:


Project progression#

#ProjectAfter sectionWhat it proves
01NumPy inference engineIIIYou understand a forward pass end to end
02CPU matmul benchmarkII, IIIYou can measure FLOPs and bandwidth
03First GPU kernelVIYou can write and launch CUDA
04Tiny transformer engineIV, VYou understand transformer inference mechanically
05KV cacheVYou understand the single most important optimization
06LLM inference serverVIIIYou can wrap a model in a real API
07Dynamic batchingVIIIYou understand throughput/latency tradeoffs
08Continuous batchingV, VIIIYou understand modern LLM serving
09KV cache managerV, XYou understand paged memory management
10Quantized inferenceVIIYou can trade accuracy for speed deliberately
11GPU benchmark suiteVI, XYou can characterize hardware
12Inference gatewayVIII, XIYou can build the production front door
13Multi-GPU inferenceIXYou understand tensor parallelism concretely
14Distributed inferenceIXYou understand multi-node and collectives
15Mini inference platformXIIYou can build the whole thing

A note on honesty#

This curriculum distinguishes three tiers of knowledge and labels them:

  • [FUNDAMENTAL] — physics and math. True in 1995, true in 2035. Memory bandwidth, arithmetic intensity, Amdahl’s law.
  • [ESTABLISHED] — production-proven engineering. Continuous batching, PagedAttention, FP8/INT8 quantization, tensor parallelism. Deployed at scale by many organizations.
  • [EMERGING] — promising but not universally production-ready. Some speculative decoding variants, aggressive low-bit formats, disaggregated serving in its newer forms.

Numbers in this curriculum (bandwidths, prices, model sizes) are illustrative and drift. The methods for computing them do not. Always re-derive with current hardware specs.


Repository files#

  • ROADMAP.md — the full I→XIV path with dependencies and checkpoints
  • GLOSSARY.md — every term, defined bluntly
  • GPU-HARDWARE.md — GPU memory types, spec sheet, memory budgets, nvidia-smi cheat sheet
  • RESOURCES.md — papers, docs, blogs, talks, organized by section
  • PROGRESS.md — your checklist

Currency#

The mechanisms in this curriculum are stable; product facts are not. Hardware tables, engine recommendations and project statuses were last reviewed on 3 October 2026. What that review changed:

  • Hardware. Rubin (HBM4, ~22 TB/s) has shipped since August 2026 and Blackwell Ultra (B300) since 2025; both are in GPU-HARDWARE.md and XIV.09.
  • Engines. Hugging Face archived TGI in March 2026; VIII.13 and VIII.14 now treat it as a migration source. vLLM and SGLang remain the reference engines.
  • Platform. Kubernetes DRA and the Gateway API Inference Extension moved from “emerging” to stable APIs; llm-d and NVIDIA Dynamo are the open-source layers above the engine.
  • Newer papers and feeds are under “2025-2026 additions” in RESOURCES.md.

Sister paths#

  • Go Engineering — the language this curriculum’s code is written in: memory, the runtime, concurrency, and AI building blocks in Go.
  • GPU Engineering — the device underneath all of this.
  • Observability Engineering — measuring it in production: SLOs, GPU and engine telemetry, tracing, cost per token.

The one idea to carry through everything#

If you remember nothing else from this repository, remember this:

Modern LLM inference is not limited by how fast your GPU can multiply. It is limited by how fast your GPU can move bytes.

Almost every technique in sections V, VII, IX, X, and XIII is, at bottom, a scheme to move fewer bytes or to do more arithmetic per byte moved. Once you see that, the field stops being a list of tricks and becomes a single coherent story.

Start with I-fundamentals/.

↑↓ navigate ↵ open