Below the API

NVIDIA Architecture Trends

Expert Advanced 1h Difficulty 3/5

Prerequisites VI.01, I.08


Every optimization in this curriculum exists because of a hardware
ratio. Change the ratio, change what matters.

  compute ≫ bandwidth  →  batching, quantization, fusion matter most
  bandwidth catches up  →  those matter less
  memory grows          →  KV cache pressure eases
  interconnect improves →  parallelism strategies change

→ read the roadmap to know what to invest in.

2. The generational trend#

                V100    A100    H100    H200    B200*   B300*   Rubin*
Year            2017    2020    2022    2023    2024    2025    2026
Memory (GB)      32      80      80     141     192     288     288
Bandwidth (TB/s) 0.9     2.0     3.35    4.8     ~8      ~8     ~22
FP16 dense (TF)  125     312     990     990    ~2250    n/p     n/p
FP8 (TF)          —       —     1979    1979   ~4500     n/p     n/p
FP4 (TF)          —       —       —       —    ~9000  ~15000  ~50000
NVLink (GB/s)    300     600     900     900    1800    1800    3600
RIDGE (FP16)     139     156     296     206     281      —       —
RIDGE (FP4)       —       —       —       —    ~1125   ~1875   ~2270

* Blackwell, Blackwell Ultra (B300) and Rubin figures are vendor numbers, rounded;
  n/p = not published. Rubin's FP4 figure is NVIDIA's inference number. Rubin began
  production shipments in August 2026. Checked 3 October 2026 — verify against
  current specs.
THE TWO TRENDS THAT MATTER

1. RIDGE POINT ROSE THEN PARTIALLY RECOVERED
   V100 → H100: 139 → 296 (compute grew faster than bandwidth)
   H100 → H200: 296 → 206 (bandwidth caught up: same FLOPs, +43% BW)
   
   → H200 was a BANDWIDTH generation, which is exactly what
     inference needed
   → memory-bound workloads got 43% faster with zero software change

2. MEMORY CAPACITY GREW SLOWLY THEN JUMPED
   32 → 80 → 80 → 141 → 192 GB
   
   → the 80 GB plateau (A100 to H100) is why so much of this
     curriculum is about fitting things in memory
   → 141-192 GB changes what fits: a 70B model at FP16 fits on
     ONE H200 (141 GB > 140 GB), barely
   → 288 GB (B300, Rubin) fits it with ~150 GB left for KV cache

3. AT FP4 THE RIDGE POINT IS CLIMBING AGAIN
   B200 → B300 → Rubin: ~1125 → ~1875 → ~2270
   → Rubin's bandwidth grew 2.75x over Blackwell; its FP4 arithmetic
     grew ~5.5x. Bandwidth "catching up" was true at fixed precision
     and stopped being true once FP4 became the headline format.
   → decode stays memory-bound up to even larger batches, so
     batching, quantization and speculative decoding matter MORE on
     the newest hardware, not less

The H100→H200 transition is instructive: no new compute, 43% more bandwidth and 76% more memory, and it was a large inference improvement. That’s evidence that the industry recognized inference is bandwidth- and capacity-bound.


3. What each generation added, and what it enabled#

VOLTA (V100)          tensor cores (FP16)
  → made transformers economically trainable and servable at all

AMPERE (A100)         TF32, BF16, INT8/INT4 tensor cores,
                      2:4 sparsity, MIG, async copy (cp.async)
  → BF16 became the training default
  → MIG enabled hard multi-tenant isolation
  → cp.async enabled software-pipelined GEMM (better FlashAttention)

HOPPER (H100)         FP8 tensor cores, TMA, thread block clusters,
                      distributed shared memory, warp-group MMA
  → FP8 became the inference standard (2x, minimal quality cost)
  → TMA + warp specialization enabled FlashAttention-3
  → this is the generation the current software stack targets

HOPPER-refresh (H200) HBM3e: 141 GB, 4.8 TB/s
  → pure inference improvement; no software change needed

BLACKWELL (B200)      FP4/FP6 tensor cores, microscaling formats,
                      2-die design, 192 GB, faster NVLink,
                      decompression engine
  → FP4 for inference (compute AND memory — Section XIII.10)
  → larger memory eases KV pressure
  → this is what the next generation of software will target

BLACKWELL ULTRA (B300) 288 GB HBM3e, ~1.5x the FP4 arithmetic,
                      SAME ~8 TB/s bandwidth
  → a capacity generation: bigger models and more KV per GPU,
    no faster at memory-bound decode

RUBIN                 HBM4: 288 GB at ~22 TB/s; NVLink 6 (3.6 TB/s per
                      GPU); sold only as liquid-cooled racks with
                      NVIDIA's own Vera CPU, joined by NVLink-C2C
  → the largest single bandwidth step so far: batch-1 decode
    ceilings rise ~2.75x with no software change
  → the rack adds a shared, RDMA-attached storage tier intended for
    KV cache ("inference context memory") — hardware built around
    the tiering in Section XIII.07
  → shipping since August 2026; independent numbers are still few

4. The four ratios to track#

1. FLOPs : BANDWIDTH  (the ridge point)
   → determines how much batching you need to be compute-bound
   → rising means batching and quantization matter more
   → H200 and Blackwell partially reversed the rise

2. MEMORY CAPACITY : MODEL SIZE
   → determines how much parallelism you need to fit
   → 192 GB means a 70B model at FP8 fits with 120 GB of KV room
   → capacity growth reduces the need for TP

3. NVLINK : HBM BANDWIDTH
   → determines TP efficiency (Section IX.03)
   H100:  900/3350  = 0.27
   B200:  1800/8000 = 0.23
   Rubin: 3600/22000 = 0.16
   → roughly constant through Blackwell, lower on Rubin: memory got
     faster than the links between GPUs, so the relative cost of
     tensor parallelism rises — and 288 GB per GPU means you need it
     less often

4. GPU : CPU BANDWIDTH
   PCIe Gen5: 64 GB/s;  NVLink-C2C (Grace): 900 GB/s;
   NVLink-C2C (Vera, in Rubin racks): ~1,800 GB/s (vendor)
   → determines whether CPU offload is viable (Section XIII.07)
   → coherent designs change this by 14-28x

Ratio 4 is the one with the largest potential impact on software architecture. Coherent CPU-GPU memory makes tiered KV caching, expert offloading, and large-model-on-small-GPU deployment all viable in a way they currently aren’t.


IF THE RIDGE POINT KEEPS RISING
  → batching, quantization, and fusion matter more
  → memory-bound techniques (speculative decoding, MoE) become
    more valuable
  → but H200/Blackwell suggest it may not

IF MEMORY CAPACITY KEEPS GROWING FAST
  → less tensor parallelism needed (fit on fewer GPUs)
  → KV cache pressure eases → less pressure on eviction/compression
  → larger models become single-node

IF LOW-PRECISION SUPPORT KEEPS DEEPENING (FP4, FP6)
  → quantization becomes a hardware feature rather than a
    software technique
  → the question shifts to whether models can be TRAINED there
    (Section XIV.04)

IF COHERENT CPU-GPU MEMORY BECOMES STANDARD
  → the memory hierarchy extends: HBM → CPU DRAM is a real tier
  → offloading, tiering, and huge-context serving change character
  → this is the biggest potential architectural shift

6. Reading a whitepaper#

WHAT TO LOOK FOR, IN ORDER

1. HBM capacity and bandwidth        ← the two numbers that matter most
2. Supported precisions and their throughputs
3. Interconnect bandwidth (NVLink, and CPU-GPU if coherent)
4. SM count and per-SM resources (registers, shared memory)
5. New instructions or units (TMA, WGMMA, decompression engines)
6. L2 size

THEN COMPUTE
  ridge point = peak_FLOPs / bandwidth
  → tells you the batch size at which decode becomes compute-bound
  
  memory / your_model_size
  → tells you the parallelism you'll need
  
  NVLink / HBM bandwidth
  → tells you TP efficiency

BE SKEPTICAL OF
  ✗ sparse FLOPs (2:4 sparsity is rarely used — Section VI.11)
  ✗ "AI TOPS" aggregate numbers that mix precisions
  ✗ comparisons against a much older generation

Always compute the ridge point yourself from the whitepaper’s numbers. It’s the single most informative derived quantity and it’s never stated directly.


7. Beyond NVIDIA’s roadmap#

Worth tracking because they change the ratios:

HBM ROADMAP
  HBM3e → HBM4 (shipping in 2026) → HBM4e (announced for 2027 parts)
  → the primary lever on inference performance

INFERENCE-SPECIFIC SILICON FROM THE GPU VENDOR
  NVIDIA licensed Groq's technology (December 2025) and announced
  a non-GPU rack, Groq 3 LPX, at GTC 2026 to sit beside Rubin racks
  for low-latency decode. The "Rubin CPX" long-context GPU announced
  in 2025 left the public roadmap.
  → hardware-level prefill/decode disaggregation (Section XIII.06)
  → watch for independent measurements and for engine support

RACK POWER
  ~120 kW (Blackwell) → roughly 200 kW (Rubin) → ~600 kW announced
  → tokens per watt becomes the planning metric (Section XI.03)

CHIPLET / MULTI-DIE
  Blackwell is two dies presented as one GPU
  → more compute per package; interconnect between dies matters

COHERENT CPU-GPU (Grace-Hopper, and successors)
  → the tier-extension change (section 4, ratio 4)

OPTICAL INTERCONNECT
  → would change multi-node economics substantially
  → co-packaged optics appear in the Rubin-generation Ethernet
    switch line; watch for deployment, not announcements

IN-NETWORK COMPUTE (SHARP and similar)
  → collectives partially executed in the switch
  → would reduce All-to-All and AllReduce cost (Sections IX.05, IX.07)

8. Hands-on exercise#

A. Build the ratio table. For every GPU you have access to, compute: ridge point, memory/model-size for a model you serve, and NVLink/HBM ratio. Which is best for your workload?

B. Predict a generation’s impact. For a hardware generation you don’t have, compute what its ratios imply for your workload. Which of your current optimizations would matter more or less?

C. Verify the specs. For a GPU you have, measure achieved bandwidth and achieved FLOPs (Section VI.01). What fraction of spec do you get? Recompute the ridge point with measured values.

D. The H200 experiment. If you have both H100 and H200, measure decode throughput for the same model and configuration. Is the improvement close to the 43% bandwidth ratio?

E. Read a whitepaper. Take the most recent NVIDIA architecture whitepaper. Extract the six items from section 6 and compute the three derived quantities. What does it imply for your stack?


9. Interview questions#

  1. Why does the FLOPs:bandwidth ratio determine which optimizations matter?
  2. Why was H200 a significant inference improvement despite the same FLOPs as H100?
  3. How would you compute a GPU’s ridge point from a whitepaper?
  4. What did Hopper add that FlashAttention-3 depends on?
  5. What would coherent CPU-GPU memory change about inference architecture?
  6. Why should you be skeptical of sparse FLOP numbers?
  7. Which hardware trend would most change your current priorities?

10. Further reading#

  • [FUNDAMENTAL] NVIDIA architecture whitepapers: Volta, Ampere, Hopper, Blackwell
  • [REFERENCE] “Inside the NVIDIA Vera Rubin Platform” (NVIDIA developer blog) — the Rubin numbers used above; vendor figures
  • [REFERENCE] NVIDIA GTC keynotes and technical sessions
  • [REFERENCE] HBM specifications and roadmaps (JEDEC)
  • [ESTABLISHED] Jia et al., microbenchmarking papers — for what the whitepapers don’t say
  • Next: 10 — Alternative accelerators

↑↓ navigate ↵ open