The final file. It looks at where the field is heading, and closes the loop on the idea this curriculum opened with.
1. The thesis#
INFERENCE IS A MEMORY PROBLEM.
decode is memory-bandwidth-bound (Section I.07)
capacity limits concurrency (Section V.06)
the KV cache is the binding constraint (Section V.05)
quantization is a bandwidth optimization (Section VII.02)
batching is a bandwidth amortization (Section I.06)
MoE trades memory capacity for quality (Section XIII.02)
disaggregation is about where KV lives (Section XIII.06)
→ EVERY major technique in this curriculum is, at bottom, about
moving fewer bytes or moving them from somewhere closer.The architectural implication: systems designed around memory rather than around compute should win. That’s the thesis of everything in this file.
2. The memory hierarchy is extending#
TODAY EMERGING
registers registers
SRAM (shared memory) SRAM
L2 L2
HBM (80-192 GB) HBM (192+ GB)
─────── PCIe wall ─────── ─────── coherent link ───────
CPU DRAM (rarely used) CPU DRAM (a real tier, 500 GB-2 TB)
NVMe (never used) CXL-attached memory (TBs)
NVMe (for cold KV)
remote memory (over RDMA)WHAT MAKES A TIER USABLE
bandwidth relative to what you'd otherwise do (Section XIII.07)
CPU DRAM over PCIe Gen5: 55 GB/s → marginal
CPU DRAM over NVLink-C2C: 900 GB/s → GENUINELY USABLE
CXL memory: ~64 GB/s per link, aggregatable
NVMe: 7 GB/s → capacity only, not speed
→ the coherent-link change (Grace-Hopper and successors) is what
turns CPU DRAM from a swap tier into a real level of the hierarchyThis is the most consequential hardware trend for inference software architecture. It changes: KV tiering (Section XIII.07), expert offloading (Section XIV.06), and what “fits” means.
3. Disaggregation, generalized#
THE PATTERN, APPLIED REPEATEDLY
DISAGGREGATE PREFILL FROM DECODE (Section XIII.06)
because they have different bottlenecks
DISAGGREGATE KV FROM COMPUTE (Mooncake, and CXL-based designs)
because KV is data with its own lifecycle
DISAGGREGATE MEMORY FROM COMPUTE (CXL memory pooling)
because the ratio of memory to compute needed varies by workload
DISAGGREGATE EXPERTS FROM ROUTING (MoE with expert offload)
because most experts are idle for any given token
THE COMMON IDEA
a monolithic GPU couples compute, memory capacity, and memory
bandwidth in a fixed ratio.
Different workloads need different ratios.
→ disaggregation lets you provision them independently.The counter-argument, which is currently winning: every disaggregation adds a link, and links are slower than what they replace. Coupling is fast; decoupling is flexible. The trade only pays when the link is fast enough — which is why coherent interconnect matters so much.
4. CXL and memory pooling#
CXL (Compute Express Link)
a cache-coherent interconnect over PCIe physical layers
CXL.mem: attach memory to a host, or POOL memory across hosts
→ a shared memory pool that any host can allocate from
FOR INFERENCE
✓ a large, shared tier for KV cache and cold model weights
✓ provision memory independently of GPUs
✓ a model too large for HBM could keep cold parts in CXL memory
✗ bandwidth: ~64 GB/s per x16 link. Aggregatable, but far below HBM.
✗ latency: 150-300 ns, vs HBM's ~500 cycles — comparable, actually
✗ maturity: deployment is early
STATUS: [RESEARCH → EMERGING]. Watch for it in inference-specific
products rather than general server designs.5. Processing in / near memory#
THE IDEA: if moving data is the bottleneck, compute where the data is.
PIM (Processing In Memory)
arithmetic units inside the DRAM die
→ HBM-PIM (Samsung), and others
→ for GEMV specifically — exactly what decode does
NEAR-MEMORY COMPUTE
compute units on the memory controller or the HBM base die
WHY IT'S APPEALING FOR LLM DECODE
decode is GEMV: read a weight, multiply once, discard.
arithmetic intensity 1 (Section I.07).
→ the ideal PIM workload: the computation is trivial and the
data movement is everything
→ a PIM device could do the multiply where the weight lives
WHY IT HASN'T HAPPENED
✗ programming model: how do you express a transformer for PIM?
✗ the arithmetic must be simple (DRAM dies have little area for logic)
✗ integration with the rest of the model (attention isn't GEMV)
✗ ecosystem: no software stack
✗ economics: HBM is already expensive; adding logic makes it more so
STATUS: [RESEARCH]. Genuinely well-matched to the problem, and
persistently 5 years away.Worth understanding because the match to LLM decode is unusually good. If anything makes PIM practical, LLM inference is the workload that would justify it.
6. What would change the field#
Ranked by potential impact:
1. COHERENT CPU-GPU MEMORY AS STANDARD
→ the hierarchy extends; offloading and tiering become architecture
rather than compromise
→ PROBABILITY: high. Already shipping.
→ IMPACT: large. Changes what fits and what you can cache.
2. HYBRID SSM/ATTENTION AT FRONTIER QUALITY (Section XIV.02)
→ O(1) state instead of O(S) KV cache
→ PROBABILITY: moderate. Shipping in some models; not yet at
the frontier.
→ IMPACT: very large for long context. 10-100x on the binding
constraint.
3. NATIVE LOW-PRECISION MODELS (FP4-trained)
→ removes the quantization step and its quality question
→ PROBABILITY: high. FP8 already happened; FP4 is the same path.
→ IMPACT: 2x on memory and compute, with no quality debate.
4. INFERENCE-TIME SCALING BECOMING UNIVERSAL (Section XIV.05)
→ workload shape shifts to overwhelmingly decode-dominated
→ PROBABILITY: already happening
→ IMPACT: large. Changes which optimizations matter.
5. BANDWIDTH GROWING FASTER THAN COMPUTE
→ the ridge point falls; less batching needed
→ PROBABILITY: moderate. H200 suggests the industry is trying.
→ IMPACT: moderate. Eases the constraint without removing it.
6. PIM / NEAR-MEMORY COMPUTE
→ PROBABILITY: low near-term
→ IMPACT: would be transformative for decode
7. OPTICAL INTERCONNECT AT SCALE
→ changes multi-node economics
→ PROBABILITY: low near-term
→ IMPACT: large for very large modelsItems 1-4 are the ones to plan around. They’re happening or likely, and each changes what you should invest in.
7. What stays true#
Whatever changes, these don’t:
1. THE ROOFLINE
FLOPs, bytes, and their ratio determine performance. The numbers
change; the model doesn't.
2. BATCHING AMORTIZES FIXED COSTS
Whatever the memory hierarchy, reading a weight once for many
tokens beats reading it once per token.
3. MEASURE BEFORE OPTIMIZING
The methodology in Sections VII.01 and X.01 outlives any technique.
4. THE BOTTLENECK MOVES
Fix one, another appears. The skill is finding the current one,
not knowing a fixed list of optimizations.
5. QUALITY IS PART OF PERFORMANCE
A faster system that answers worse is not faster.
6. SCHEDULING BEATS KERNELS
The 3-10x from batching and memory management exceeds the 20-60%
from compilation and kernel work. Order accordingly.
7. THE SIMPLE THING FIRST
FP8 before FP4. Chunked prefill before disaggregation. Routing
before a KV store. n-gram before EAGLE.If the curriculum has a single thesis beyond “inference is a memory problem,” it’s number 7.
8. Where to go from here#
YOU HAVE FINISHED THE CURRICULUM. NEXT:
1. BUILD SOMETHING
The projects (projects/) if you haven't. Or contribute to vLLM
or SGLang — the schedulers and attention backends are where the
education is.
2. MEASURE SOMETHING REAL
Take a production system. Apply Section X.01's methodology.
Find the bottleneck. Fix it. Write it up.
3. HALVE A COST
Take one system's cost per million tokens and cut it in half.
Document what you did. That document is your career.
4. TEACH IT
The fastest way to find your gaps. Section 3 of every file
(the analogy) exists partly for this.
5. STAY CALIBRATED
Re-read Section XIV.01 when a new technique appears.
Re-measure your `numbers.md` on new hardware.
Re-run your capacity plan quarterly.9. Hands-on exercise#
A. Project the hierarchy. For a hypothetical system with coherent 900 GB/s CPU-GPU memory and 1 TB of CPU DRAM, recompute: max concurrency for a 70B model at 128k context, and whether KV offloading becomes the default. What changes?
B. The hybrid projection. For a hybrid SSM/attention model with 1 attention layer per 8, compute the KV cache at 1M context. Compare to pure attention. What does it do to concurrency?
C. Revisit your priorities. Take your current optimization backlog. For each of the four likely changes in section 6, note which items become more or less valuable.
D. The unchanging test. For each of the seven things in section 7, find an example in this curriculum where it applied. Then find a case in your own work.
E. Write your own section 6. Based on your reading and your hardware roadmap, write your own ranked list of what would change the field. Revisit it in a year and see how you did.
10. Interview questions#
- Why is inference fundamentally a memory problem?
- What does coherent CPU-GPU memory change about inference architecture?
- What is disaggregation, generalized, and what’s the counter-argument?
- Why is LLM decode a good match for processing-in-memory?
- What would hybrid SSM/attention models change about long-context serving?
- What stays true regardless of hardware changes?
- If you had to bet on one change transforming inference in the next three years, what would it be and why?
11. Further reading#
- [REFERENCE] CXL specification and consortium materials
- [RESEARCH] HBM-PIM and near-memory computing literature (Samsung, SK Hynix, UPMEM)
- [EMERGING] Grace-Hopper and coherent-interconnect architecture documentation
- [EMERGING] Qin et al., “Mooncake” — the KV-centric framing
- [FUNDAMENTAL] Wulf & McKee, “Hitting the Memory Wall” (1995) — where this all started
- Back to: README.md | Projects: projects/