Inference Engineering
How modern LLM inference works — from a single matrix multiply to a multi-tenant serving platform.
5 levels · 186 topics · ~354h 45m
Edition 01 — Engineering curriculum
Deep technical courses on how inference systems, GPUs, the telemetry that watches them and the language they are built in actually work — from first principles to production.
04 paths 285 topics 5 levels
How modern LLM inference works — from a single matrix multiply to a multi-tenant serving platform.
5 levels · 186 topics · ~354h 45m
The GPU itself — what is inside the chip, how it is programmed, why it is fast, and how thousands are run together.
5 levels · 32 topics · ~27h 15m
Seeing inside running systems — metrics, logs, traces and profiles, then the GPU, token and cost telemetry that AI infrastructure adds.
5 levels · 30 topics · ~24h 25m
Go from the first program to the runtime — how values sit in memory, how the allocator, garbage collector and scheduler work, and how to build AI systems in it.
5 levels · 37 topics · ~33h 5m
Building systems where models plan, call tools and act — loops, context, evaluation and safety.
Coming soon
New paths are added as plain Markdown folders. Distributed systems for AI and networking for GPU clusters are on the list.
Every path runs through the same five levels. Each level assumes the ones before it and nothing else.