Goal: give you the systems knowledge that inference engineering assumes and rarely teaches. Every file answers “why does an inference engineer care about this?” — if a systems topic doesn’t affect inference, it isn’t here.
You can serve a model without this section. You cannot debug one.
Files#
| # | File | Level | Time |
|---|---|---|---|
| 01 | CPU architecture | Beginner | 60 min |
| 02 | Caches and the memory hierarchy | Intermediate | 75 min |
| 03 | RAM, DRAM, and NUMA | Intermediate | 60 min |
| 04 | SIMD and vectorization | Intermediate | 60 min |
| 05 | Threads, processes, context switching | Beginner | 60 min |
| 06 | Memory allocation and virtual memory | Intermediate | 75 min |
| 07 | Storage and model loading | Intermediate | 60 min |
| 08 | Networking fundamentals | Intermediate | 60 min |
| 09 | PCIe, DMA, and interconnects | Intermediate | 60 min |
| 10 | OS scheduling | Intermediate | 45 min |
| 11 | Linux for inference engineers | Beginner | 75 min |
| 12 | Containers, namespaces, cgroups | Intermediate | 60 min |
| 13 | Syscalls and profiling basics | Intermediate | 75 min |
The thread#
flowchart TD N0["Compute is fast; memory is slow<br/><b>01, 02</b>"] N1["So hardware hides latency with caches, prefetch, SIMD, threads<br/><b>02, 04, 05</b>"] N2["And the OS multiplexes all of it<br/><b>05, 06, 10</b>"] N3["Data must cross buses to reach the GPU (09) and networks to reach other nodes<br/><b>08</b>"] N4["And models must be read from storage before any of it starts<br/><b>07</b>"] N5["All of it inside containers with limits you must understand<br/><b>12</b>"] N6["And when it's slow, you need tools to see it<br/><b>11, 13</b>"] N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 class N0,N1 neutral class N2 io class N3,N4 queue class N5 compute class N6 memory
Why this matters (concrete payoffs)#
- Section 02 explains why tokenization is slow and how to fix it.
- Section 03 explains a class of mysterious 2x slowdowns on dual-socket boxes.
- Section 06 is the direct conceptual ancestor of PagedAttention (Section V.10).
- Section 07 explains your 8-minute cold start and how to cut it to 40 seconds.
- Section 09 explains why you never stream weights over PCIe per request.
- Section 12 explains why your container sees 128 CPUs and gets throttled at 4.
- Section 13 gives you the tools you will use in Section X.