Below the API

Computer Systems Foundations

FoundationsModule13 topics~13h 45m
completed

Topics, in order

About this module

Goal: give you the systems knowledge that inference engineering assumes and rarely teaches. Every file answers “why does an inference engineer care about this?” — if a systems topic doesn’t affect inference, it isn’t here.

You can serve a model without this section. You cannot debug one.

Files#

#FileLevelTime
01CPU architectureBeginner60 min
02Caches and the memory hierarchyIntermediate75 min
03RAM, DRAM, and NUMAIntermediate60 min
04SIMD and vectorizationIntermediate60 min
05Threads, processes, context switchingBeginner60 min
06Memory allocation and virtual memoryIntermediate75 min
07Storage and model loadingIntermediate60 min
08Networking fundamentalsIntermediate60 min
09PCIe, DMA, and interconnectsIntermediate60 min
10OS schedulingIntermediate45 min
11Linux for inference engineersBeginner75 min
12Containers, namespaces, cgroupsIntermediate60 min
13Syscalls and profiling basicsIntermediate75 min

The thread#

flowchart TD
  N0["Compute is fast; memory is slow<br/><b>01, 02</b>"]
  N1["So hardware hides latency with caches, prefetch, SIMD, threads<br/><b>02, 04, 05</b>"]
  N2["And the OS multiplexes all of it<br/><b>05, 06, 10</b>"]
  N3["Data must cross buses to reach the GPU (09) and networks to reach other nodes<br/><b>08</b>"]
  N4["And models must be read from storage before any of it starts<br/><b>07</b>"]
  N5["All of it inside containers with limits you must understand<br/><b>12</b>"]
  N6["And when it's slow, you need tools to see it<br/><b>11, 13</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6

  class N0,N1 neutral
  class N2 io
  class N3,N4 queue
  class N5 compute
  class N6 memory

Why this matters (concrete payoffs)#

  • Section 02 explains why tokenization is slow and how to fix it.
  • Section 03 explains a class of mysterious 2x slowdowns on dual-socket boxes.
  • Section 06 is the direct conceptual ancestor of PagedAttention (Section V.10).
  • Section 07 explains your 8-minute cold start and how to cut it to 40 seconds.
  • Section 09 explains why you never stream weights over PCIe per request.
  • Section 12 explains why your container sees 128 CPUs and gets throttled at 4.
  • Section 13 gives you the tools you will use in Section X.

↑↓ navigate ↵ open