1. Why hardware trends determine software priorities#
Every optimization in this curriculum exists because of a hardware
ratio. Change the ratio, change what matters.
compute ≫ bandwidth → batching, quantization, fusion matter most
bandwidth catches up → those matter less
memory grows → KV cache pressure eases
interconnect improves → parallelism strategies change
→ read the roadmap to know what to invest in.2. The generational trend#
V100 A100 H100 H200 B200* B300* Rubin*
Year 2017 2020 2022 2023 2024 2025 2026
Memory (GB) 32 80 80 141 192 288 288
Bandwidth (TB/s) 0.9 2.0 3.35 4.8 ~8 ~8 ~22
FP16 dense (TF) 125 312 990 990 ~2250 n/p n/p
FP8 (TF) — — 1979 1979 ~4500 n/p n/p
FP4 (TF) — — — — ~9000 ~15000 ~50000
NVLink (GB/s) 300 600 900 900 1800 1800 3600
RIDGE (FP16) 139 156 296 206 281 — —
RIDGE (FP4) — — — — ~1125 ~1875 ~2270
* Blackwell, Blackwell Ultra (B300) and Rubin figures are vendor numbers, rounded;
n/p = not published. Rubin's FP4 figure is NVIDIA's inference number. Rubin began
production shipments in August 2026. Checked 3 October 2026 — verify against
current specs.THE TWO TRENDS THAT MATTER
1. RIDGE POINT ROSE THEN PARTIALLY RECOVERED
V100 → H100: 139 → 296 (compute grew faster than bandwidth)
H100 → H200: 296 → 206 (bandwidth caught up: same FLOPs, +43% BW)
→ H200 was a BANDWIDTH generation, which is exactly what
inference needed
→ memory-bound workloads got 43% faster with zero software change
2. MEMORY CAPACITY GREW SLOWLY THEN JUMPED
32 → 80 → 80 → 141 → 192 GB
→ the 80 GB plateau (A100 to H100) is why so much of this
curriculum is about fitting things in memory
→ 141-192 GB changes what fits: a 70B model at FP16 fits on
ONE H200 (141 GB > 140 GB), barely
→ 288 GB (B300, Rubin) fits it with ~150 GB left for KV cache
3. AT FP4 THE RIDGE POINT IS CLIMBING AGAIN
B200 → B300 → Rubin: ~1125 → ~1875 → ~2270
→ Rubin's bandwidth grew 2.75x over Blackwell; its FP4 arithmetic
grew ~5.5x. Bandwidth "catching up" was true at fixed precision
and stopped being true once FP4 became the headline format.
→ decode stays memory-bound up to even larger batches, so
batching, quantization and speculative decoding matter MORE on
the newest hardware, not lessThe H100→H200 transition is instructive: no new compute, 43% more bandwidth and 76% more memory, and it was a large inference improvement. That’s evidence that the industry recognized inference is bandwidth- and capacity-bound.
3. What each generation added, and what it enabled#
VOLTA (V100) tensor cores (FP16)
→ made transformers economically trainable and servable at all
AMPERE (A100) TF32, BF16, INT8/INT4 tensor cores,
2:4 sparsity, MIG, async copy (cp.async)
→ BF16 became the training default
→ MIG enabled hard multi-tenant isolation
→ cp.async enabled software-pipelined GEMM (better FlashAttention)
HOPPER (H100) FP8 tensor cores, TMA, thread block clusters,
distributed shared memory, warp-group MMA
→ FP8 became the inference standard (2x, minimal quality cost)
→ TMA + warp specialization enabled FlashAttention-3
→ this is the generation the current software stack targets
HOPPER-refresh (H200) HBM3e: 141 GB, 4.8 TB/s
→ pure inference improvement; no software change needed
BLACKWELL (B200) FP4/FP6 tensor cores, microscaling formats,
2-die design, 192 GB, faster NVLink,
decompression engine
→ FP4 for inference (compute AND memory — Section XIII.10)
→ larger memory eases KV pressure
→ this is what the next generation of software will target
BLACKWELL ULTRA (B300) 288 GB HBM3e, ~1.5x the FP4 arithmetic,
SAME ~8 TB/s bandwidth
→ a capacity generation: bigger models and more KV per GPU,
no faster at memory-bound decode
RUBIN HBM4: 288 GB at ~22 TB/s; NVLink 6 (3.6 TB/s per
GPU); sold only as liquid-cooled racks with
NVIDIA's own Vera CPU, joined by NVLink-C2C
→ the largest single bandwidth step so far: batch-1 decode
ceilings rise ~2.75x with no software change
→ the rack adds a shared, RDMA-attached storage tier intended for
KV cache ("inference context memory") — hardware built around
the tiering in Section XIII.07
→ shipping since August 2026; independent numbers are still few4. The four ratios to track#
1. FLOPs : BANDWIDTH (the ridge point)
→ determines how much batching you need to be compute-bound
→ rising means batching and quantization matter more
→ H200 and Blackwell partially reversed the rise
2. MEMORY CAPACITY : MODEL SIZE
→ determines how much parallelism you need to fit
→ 192 GB means a 70B model at FP8 fits with 120 GB of KV room
→ capacity growth reduces the need for TP
3. NVLINK : HBM BANDWIDTH
→ determines TP efficiency (Section IX.03)
H100: 900/3350 = 0.27
B200: 1800/8000 = 0.23
Rubin: 3600/22000 = 0.16
→ roughly constant through Blackwell, lower on Rubin: memory got
faster than the links between GPUs, so the relative cost of
tensor parallelism rises — and 288 GB per GPU means you need it
less often
4. GPU : CPU BANDWIDTH
PCIe Gen5: 64 GB/s; NVLink-C2C (Grace): 900 GB/s;
NVLink-C2C (Vera, in Rubin racks): ~1,800 GB/s (vendor)
→ determines whether CPU offload is viable (Section XIII.07)
→ coherent designs change this by 14-28xRatio 4 is the one with the largest potential impact on software architecture. Coherent CPU-GPU memory makes tiered KV caching, expert offloading, and large-model-on-small-GPU deployment all viable in a way they currently aren’t.
5. What the trends imply for software#
IF THE RIDGE POINT KEEPS RISING
→ batching, quantization, and fusion matter more
→ memory-bound techniques (speculative decoding, MoE) become
more valuable
→ but H200/Blackwell suggest it may not
IF MEMORY CAPACITY KEEPS GROWING FAST
→ less tensor parallelism needed (fit on fewer GPUs)
→ KV cache pressure eases → less pressure on eviction/compression
→ larger models become single-node
IF LOW-PRECISION SUPPORT KEEPS DEEPENING (FP4, FP6)
→ quantization becomes a hardware feature rather than a
software technique
→ the question shifts to whether models can be TRAINED there
(Section XIV.04)
IF COHERENT CPU-GPU MEMORY BECOMES STANDARD
→ the memory hierarchy extends: HBM → CPU DRAM is a real tier
→ offloading, tiering, and huge-context serving change character
→ this is the biggest potential architectural shift6. Reading a whitepaper#
WHAT TO LOOK FOR, IN ORDER
1. HBM capacity and bandwidth ← the two numbers that matter most
2. Supported precisions and their throughputs
3. Interconnect bandwidth (NVLink, and CPU-GPU if coherent)
4. SM count and per-SM resources (registers, shared memory)
5. New instructions or units (TMA, WGMMA, decompression engines)
6. L2 size
THEN COMPUTE
ridge point = peak_FLOPs / bandwidth
→ tells you the batch size at which decode becomes compute-bound
memory / your_model_size
→ tells you the parallelism you'll need
NVLink / HBM bandwidth
→ tells you TP efficiency
BE SKEPTICAL OF
✗ sparse FLOPs (2:4 sparsity is rarely used — Section VI.11)
✗ "AI TOPS" aggregate numbers that mix precisions
✗ comparisons against a much older generationAlways compute the ridge point yourself from the whitepaper’s numbers. It’s the single most informative derived quantity and it’s never stated directly.
7. Beyond NVIDIA’s roadmap#
Worth tracking because they change the ratios:
HBM ROADMAP
HBM3e → HBM4 (shipping in 2026) → HBM4e (announced for 2027 parts)
→ the primary lever on inference performance
INFERENCE-SPECIFIC SILICON FROM THE GPU VENDOR
NVIDIA licensed Groq's technology (December 2025) and announced
a non-GPU rack, Groq 3 LPX, at GTC 2026 to sit beside Rubin racks
for low-latency decode. The "Rubin CPX" long-context GPU announced
in 2025 left the public roadmap.
→ hardware-level prefill/decode disaggregation (Section XIII.06)
→ watch for independent measurements and for engine support
RACK POWER
~120 kW (Blackwell) → roughly 200 kW (Rubin) → ~600 kW announced
→ tokens per watt becomes the planning metric (Section XI.03)
CHIPLET / MULTI-DIE
Blackwell is two dies presented as one GPU
→ more compute per package; interconnect between dies matters
COHERENT CPU-GPU (Grace-Hopper, and successors)
→ the tier-extension change (section 4, ratio 4)
OPTICAL INTERCONNECT
→ would change multi-node economics substantially
→ co-packaged optics appear in the Rubin-generation Ethernet
switch line; watch for deployment, not announcements
IN-NETWORK COMPUTE (SHARP and similar)
→ collectives partially executed in the switch
→ would reduce All-to-All and AllReduce cost (Sections IX.05, IX.07)8. Hands-on exercise#
A. Build the ratio table. For every GPU you have access to, compute: ridge point, memory/model-size for a model you serve, and NVLink/HBM ratio. Which is best for your workload?
B. Predict a generation’s impact. For a hardware generation you don’t have, compute what its ratios imply for your workload. Which of your current optimizations would matter more or less?
C. Verify the specs. For a GPU you have, measure achieved bandwidth and achieved FLOPs (Section VI.01). What fraction of spec do you get? Recompute the ridge point with measured values.
D. The H200 experiment. If you have both H100 and H200, measure decode throughput for the same model and configuration. Is the improvement close to the 43% bandwidth ratio?
E. Read a whitepaper. Take the most recent NVIDIA architecture whitepaper. Extract the six items from section 6 and compute the three derived quantities. What does it imply for your stack?
9. Interview questions#
- Why does the FLOPs:bandwidth ratio determine which optimizations matter?
- Why was H200 a significant inference improvement despite the same FLOPs as H100?
- How would you compute a GPU’s ridge point from a whitepaper?
- What did Hopper add that FlashAttention-3 depends on?
- What would coherent CPU-GPU memory change about inference architecture?
- Why should you be skeptical of sparse FLOP numbers?
- Which hardware trend would most change your current priorities?
10. Further reading#
- [FUNDAMENTAL] NVIDIA architecture whitepapers: Volta, Ampere, Hopper, Blackwell
- [REFERENCE] “Inside the NVIDIA Vera Rubin Platform” (NVIDIA developer blog) — the Rubin numbers used above; vendor figures
- [REFERENCE] NVIDIA GTC keynotes and technical sessions
- [REFERENCE] HBM specifications and roadmaps (JEDEC)
- [ESTABLISHED] Jia et al., microbenchmarking papers — for what the whitepapers don’t say
- Next: 10 — Alternative accelerators