Below the API

Memory and Performance Engineering

AdvancedModule11 topics~13h 30m
completed

Topics, in order

About this module

Goal: answer “why is my inference system slow?” systematically, every time, without guessing.

This section is the toolbox. Sections I-IX taught you what the system does; this one teaches you how to find out what it’s actually doing.

Files#

#FileLevelTime
01The performance debugging methodology ★Advanced90 min
02The roofline model in practiceAdvanced60 min
03Memory-bound vs compute-boundAdvanced60 min
04GPU utilization is a lie ★Intermediate60 min
05Memory fragmentation and allocatorsAdvanced75 min
06OOM: causes and curesAdvanced60 min
07Benchmarking inference correctly ★Advanced90 min
08GPU profiling with NsightAdvanced90 min
09CPU and Python profilingIntermediate60 min
10Metrics, Prometheus, GrafanaIntermediate75 min
11Case studies ★Advanced90 min

The method, in one diagram#

flowchart TB
  S["It's slow"] --> A["1. WHICH METRIC?<br/>TTFT, ITL, throughput, cost"]
  A --> B["2. WHICH PHASE?<br/>queue, tokenize, prefill, decode, detokenize, network<br/><b>file 10</b>"]
  B --> C["3. GPU, CPU, OR WAITING?<br/>timeline analysis<br/><b>files 08, 09</b>"]
  C --> D["4. WHICH REGIME?<br/>memory, compute, latency, launch<br/><b>files 02, 03</b>"]
  D --> E["5. WHICH KERNEL?<br/>per-kernel analysis<br/><b>file 08</b>"]
  E --> F["6. FIX ONE THING, RE-MEASURE"]
  F -.->|"still slow"| A

  class S warn
  class A,B neutral
  class C io
  class D queue
  class E compute
  class F memory

Never skip a level. Starting at step 5 optimizes kernels that don’t matter.

Checkpoint E#

  1. Given a slow inference service, list your first five measurements in order.
  2. Draw a roofline and place decode, prefill, and RMSNorm on it.
  3. nvidia-smi says 100% GPU utilization. Why might the GPU still be mostly idle?
  4. What makes an inference benchmark valid? Name five requirements.
  5. Your p99 ITL is 5x your p50. Give four hypotheses and how you’d distinguish them.

↑↓ navigate ↵ open