Below the API

GPU Hardware & Memory Reference

Beginner → Advanced (reference; dip in as needed)

Prerequisites I.07, I.08 ; pairs with VI.01, VI.03, V.06, V.15

One page for the hardware facts the rest of the curriculum keeps reaching for: what GPU memory is, how much each card has, how fast it moves bytes, where the bytes go, and how to inspect all of it.

Numbers are vendor-published figures, rounded. They drift with SKUs and revisions — treat them as starting points and replace them with what you measure in Project 11. The tables were last checked on 3 October 2026; figures for parts that began shipping in 2026 (Rubin, MI455X) are vendor claims with little independent measurement yet.


1. The three numbers that decide everything#

flowchart LR
  CAP["CAPACITY<br/>GB of VRAM"] --> FIT["What fits<br/>weights + KV cache<br/>= model size x concurrency"]
  BW["BANDWIDTH<br/>GB/s of memory"] --> DEC["Decode speed<br/>memory-bound: every token<br/>re-reads every weight"]
  FL["COMPUTE<br/>TFLOP/s"] --> PRE["Prefill speed and<br/>large-batch throughput<br/>compute-bound"]
  FIT --> COST["Cost per token"]
  DEC --> COST
  PRE --> COST

  class CAP,BW,FIT,DEC memory
  class FL,PRE compute
  class COST neutral

For LLM serving, read a spec sheet in this order: capacity, bandwidth, then FLOPs. Most buying mistakes come from reading it in the opposite order.


2. What “GPU memory” physically is#

Two families of memory sit next to GPU dies. Which one a card uses explains most of its price and most of its inference performance.

flowchart TB
  subgraph GD["GDDR card (T4, L4, A10G, L40S, RTX)"]
    direction LR
    D1["GPU die"] --- PCB["Memory chips soldered on the PCB<br/>around the package<br/>narrow, fast-clocked links"]
  end
  subgraph HB["HBM card (A100, H100, H200, B200, MI300X)"]
    direction LR
    D2["GPU die"] --- INT["HBM stacks on the SAME silicon interposer<br/>DRAM dies stacked vertically, wired with TSVs<br/>very wide, short links"]
  end
  GD --> R1["Cheaper, lower bandwidth<br/>~0.3 - 1.8 TB/s"]
  HB --> R2["Expensive, supply-constrained, high bandwidth<br/>~2 - 22 TB/s"]

  class PCB,INT memory
  class D1,D2 compute
  class R1,R2 neutral
TechnologyWhere it sitsTypical card bandwidthFound in
GDDR6Discrete chips on the board0.3 - 0.9 TB/sT4, L4, A10G, L40S
GDDR6X / GDDR7Discrete chips, faster signalling0.9 - 1.8 TB/sRTX 3090/4090, RTX 5090
HBM2 / HBM2eStacked, on interposer1.5 - 2.0 TB/sA100
HBM3Stacked, on interposer3.35 TB/sH100 SXM, MI300X (5.3)
HBM3eStacked, on interposer4.8 - 8 TB/sH200, B200, B300, MI325X/MI355X
HBM4Stacked, interface twice as wide~22 TB/sRubin (shipping since August 2026), MI455X
Unified LPDDRShared with the CPU0.1 - 0.8 TB/sApple M-series, DGX Spark, Jetson

Why it matters to you:

  • HBM is the scarce component. GPU supply and price track HBM and advanced-packaging capacity more than they track logic wafers.
  • Bandwidth comes from width × speed. HBM gets it from thousands of short wires through the interposer; GDDR has to do it with far fewer, longer traces. That is a physical ceiling, not a tuning problem.
  • ECC. Datacenter cards have error-correcting memory; consumer cards do not. A flipped bit in a weight tensor on a consumer card is silent.

Inside the GPU the hierarchy above HBM — registers, shared memory, L2 — is covered in VI.03.


3. Spec sheet: the cards you will meet#

Datacenter — NVIDIA#

GPUYearArchVRAMMemory typeBandwidthFP16/BF16 densePowerGPU↔GPU link
T42018Turing16 GBGDDR60.32 TB/s65 TF70 WPCIe only
A10G2021Ampere24 GBGDDR60.60 TB/s~125 TF150 WPCIe only
L42023Ada24 GBGDDR60.30 TB/s~120 TF72 WPCIe only
L40S2023Ada48 GBGDDR60.86 TB/s~362 TF350 WPCIe only
A100 40GB2020Ampere40 GBHBM21.56 TB/s312 TF400 WNVLink 600 GB/s
A100 80GB2020Ampere80 GBHBM2e2.04 TB/s312 TF400 WNVLink 600 GB/s
H100 PCIe2022Hopper80 GBHBM2e~2.0 TB/s~756 TF350 WNVLink bridge
H100 SXM2022Hopper80 GBHBM33.35 TB/s~990 TF700 WNVLink 900 GB/s
H2002023Hopper141 GBHBM3e4.8 TB/s~990 TF700 WNVLink 900 GB/s
B2002024Blackwell~180-192 GBHBM3e~8 TB/s~2,250 TF~1,000 WNVLink 1.8 TB/s
B3002025Blackwell Ultra288 GBHBM3e~8 TB/snot published (~15 PF at FP4)~1,400 WNVLink 1.8 TB/s
Rubin2026Rubin288 GBHBM4~22 TB/snot published (~50 PF at FP4, NVIDIA’s inference figure)~1,800-2,300 WNVLink 3.6 TB/s

Rubin is sold only inside liquid-cooled Vera Rubin racks; you rent it, you do not slot it into a server. B300 has the B200’s bandwidth: for decode it is a capacity upgrade, not a speed one.

Datacenter — AMD#

GPUVRAMMemory typeBandwidthPower
Instinct MI300X192 GBHBM35.3 TB/s750 W
Instinct MI325X256 GBHBM3e~6 TB/s~1,000 W
Instinct MI355X288 GBHBM3e~8 TB/s~1,400 W
Instinct MI455X432 GBHBM4~23 TB/s (vendor)sold in Helios racks, shipping from H2 2026

Workstation / consumer (learning, local serving)#

GPUVRAMMemory typeBandwidthPower
RTX 309024 GBGDDR6X0.94 TB/s350 W
RTX 409024 GBGDDR6X1.01 TB/s450 W
RTX 509032 GBGDDR71.79 TB/s575 W
Apple M-series (Max/Ultra)up to 128-512 GB unifiedLPDDR5(X)~0.4 - 0.8 TB/s60-200 W system

Things to notice

  • H100 → H200 added no compute — only 43% more bandwidth and 76% more capacity — and was a large inference upgrade. That tells you what inference is bound by.
  • L4 and T4 are capacity-per-watt cards, not speed cards. An L4 has the VRAM of an A10G and half its bandwidth.
  • A 4090 out-reads an L40S on bandwidth but has half the memory and no ECC.
  • “H100” is two different cards. The PCIe variant has HBM2e and roughly 60% of the SXM bandwidth. Check which one your cloud instance has.
  • MIG (hardware partitioning into up to 7 isolated instances) exists on A100 / H100 / H200 / Blackwell, not on T4 / L4 / A10G / L40S.

For the generation-over-generation ratios and where they are heading, see XIV.09.


4. Where the gigabytes go#

pie showData
  title 80 GB H100 serving an 8B model in FP16 (illustrative)
  "Model weights" : 16
  "KV cache pool" : 52
  "Activations and workspace" : 2
  "CUDA context, framework, graphs" : 2
  "Unreserved headroom (10%)" : 8
ConsumerSizeScales withNotes
Weightsparams × bytes/paramModel, precisionFixed once loaded
KV cache2 × layers × kv_heads × head_dim × bytes × tokensConcurrency × contextThe part you actually manage (V.06)
Activationsbatch_tokens × d_model × small factorTokens per stepPeaks during prefill of long prompts
CUDA context~0.3 - 0.5 GB per processNumber of processesPaid per process, not per GPU
Framework overhead0.5 - 2 GBEnginecuDNN/cuBLAS workspaces, CUDA graphs, compiled kernels
Allocator slackvariesAllocation patternreserved − allocated; fragmentation (X.05)

Engines such as vLLM claim a fixed fraction of VRAM at startup (--gpu-memory-utilization, default 0.9), load the weights, and turn everything left over into the KV block pool. So nvidia-smi shows ~90% used from the first second, whether you have 1 user or 100.

Weight sizes to memorize#

ParametersFP32FP16 / BF16FP8 / INT8INT4 (≈4.5 bit effective)
1 B4 GB2 GB1 GB~0.6 GB
8 B32 GB16 GB8 GB~4.5 GB
32 B128 GB64 GB32 GB~18 GB
70 B280 GB140 GB70 GB~39 GB
120 B480 GB240 GB120 GB~68 GB

What fits where (weights only — leave room for KV)#

Model @ precision16 GB24 GB48 GB80 GB141 GB~180 GB
8B FP16 (16 GB)no room for KVyesyesyesyesyes
8B INT4 (4.5 GB)yesyesyesyesyesyes
32B FP16 (64 GB)———tightyesyes
32B INT4 (18 GB)—tightyesyesyesyes
70B FP16 (140 GB)———2 GPUsbarely, no KVyes
70B FP8 (70 GB)———tightyesyes
70B INT4 (39 GB)——tightyesyesyes

“Fits” without KV room is not serving — it is loading. A model that leaves 1 GB for KV serves almost nobody.


5. From bandwidth to tokens per second#

At batch 1, a decode step reads every weight once. So:

tokens/s ceiling (batch 1)  ≈  memory bandwidth / weight bytes
realistic                   ≈  60-75% of that

Ceiling for a model with 16 GB of weights (e.g. 8B at FP16):

GPUBandwidthBatch-1 ceilingRealistic
L40.30 TB/s~19 tok/s~12-14
A10G0.60 TB/s~38 tok/s~23-28
L40S0.86 TB/s~54 tok/s~32-40
RTX 40901.01 TB/s~63 tok/s~38-47
RTX 50901.79 TB/s~112 tok/s~67-84
A100 80GB2.04 TB/s~127 tok/s~76-95
H100 SXM3.35 TB/s~209 tok/s~125-157
H2004.8 TB/s~300 tok/s~180-225
B200 / B300~8 TB/s~500 tok/s~300-375
Rubin~22 TB/s~1,375 tok/s~825-1,030 (predicted; measure it)

Three ways to beat the ceiling, all of which are “move fewer bytes per token”:

flowchart LR
  C["Batch-1 ceiling<br/>bandwidth / weight bytes"] --> B["Batch more sequences<br/>one weight read serves many tokens"]
  C --> Q["Quantize weights<br/>fewer bytes to read"]
  C --> S["Speculative decoding<br/>several tokens per weight read"]
  B --> R["Until the ridge point:<br/>then you are compute-bound"]

  class C neutral
  class B,Q,S memory
  class R compute

The ridge point (peak FLOPs / bandwidth) for these cards and what it means for batch size is worked through in VI.04 and V.15.


6. The paths bytes take: interconnect speeds#

flowchart LR
  DISK["NVMe SSD"] -->|"3 - 14 GB/s"| RAM["Host RAM"]
  NET["Network storage"] -->|"0.1 - 5 GB/s"| RAM
  RAM -->|"PCIe Gen4 x16: ~25 GB/s real<br/>Gen5 x16: ~50 GB/s real"| G0[("GPU 0 HBM<br/>2 - 22 TB/s internal")]
  G0 <-->|"NVLink: 600 - 3,600 GB/s"| G1[("GPU 1 HBM")]
  G0 -->|"RDMA NIC: 12 - 100 GB/s"| NODE["Other nodes"]

  class G0,G1 memory
  class DISK,NET,NODE io
  class RAM neutral
LinkTheoreticalWhat it means for inference
HBM inside one GPU2,000 - 22,000 GB/sThe speed everything else is compared to
NVLink 3 / 4 / 5 / 6 (per GPU)600 / 900 / 1,800 / 3,600 GB/sMakes tensor parallelism viable (IX.03)
PCIe Gen3 x16~16 GB/sT4-era hosts; slow model loads
PCIe Gen4 x16~32 GB/s~25 GB/s with pinned memory
PCIe Gen5 x16~64 GB/sCurrent servers
100 / 200 / 400 GbE or InfiniBand12.5 / 25 / 50 GB/sKV transfer for disaggregation (XIII.06)
NVMe Gen4 / Gen5~7 / ~14 GB/sCold-start weight loading (II.07)

The gap between the first row and the rest — two orders of magnitude — is why anything that leaves the GPU on the per-token path is a design error, and why swapping KV to host memory is often slower than recomputing it.

Load-time arithmetic: 140 GB of weights ÷ 7 GB/s NVMe ≈ 20 s at best; from network storage at 1 GB/s, over two minutes. That is your cold start floor before the first kernel runs.


7. Inside a GPU server#

flowchart TB
  subgraph NODE["8-GPU server (HGX / DGX class)"]
    direction TB
    subgraph CPUS["Host"]
      C0["CPU socket 0<br/>+ DRAM (NUMA node 0)"]
      C1["CPU socket 1<br/>+ DRAM (NUMA node 1)"]
    end
    subgraph PX["PCIe switches"]
      S0["Switch A"]
      S1["Switch B"]
    end
    subgraph GP["GPUs"]
      direction LR
      G0["GPU 0-3"]
      G1["GPU 4-7"]
    end
    NVS["NVSwitch fabric<br/>every GPU to every GPU at full NVLink speed"]
    NIC["RDMA NICs / DPUs"]
    SSD["NVMe"]
    C0 --> S0 --> G0
    C1 --> S1 --> G1
    G0 <--> NVS
    G1 <--> NVS
    S0 --> NIC
    S1 --> SSD
  end

  class G0,G1 compute
  class NVS memory
  class NIC,SSD,S0,S1 io
  class C0,C1 neutral
SystemGPUsTotal GPU memoryGPU fabricPower
Single-GPU cloud VM (T4, L4, A10G, L40S)116-48 GBnone< 0.5 kW
8× A100 80GB server8640 GBNVSwitch, 600 GB/s~6.5 kW
8× H100 server (DGX H100 class)8640 GBNVSwitch, 900 GB/s~10 kW
8× H200 server8~1.1 TBNVSwitch, 900 GB/s~10 kW
GB200 NVL72 rack72~13 TBNVLink switch spine, 1.8 TB/s, one domain~120 kW, liquid cooled
GB300 NVL72 rack (Blackwell Ultra)72~20.7 TBNVLink switch spine, 1.8 TB/s, one domain~120-140 kW, liquid cooled
Vera Rubin NVL72 rack72~20.7 TB (HBM4)NVLink 6, 3.6 TB/s, one domainroughly 200 kW or more (reported), liquid cooled

Host-side sizing that people get wrong:

  • System RAM ≥ total VRAM as a floor (more if you offload KV or keep several models page-cached). Loading goes disk → page cache → GPU.
  • NUMA locality. A GPU hangs off one socket. Pinning the process to the other socket’s memory costs transfer bandwidth (II.03). nvidia-smi topo -m shows the mapping.
  • CPU cores. Tokenization, scheduling, and detokenization are CPU work on the critical path. Budget several fast cores per GPU; a starved host shows up as an idle GPU.
  • Power and cooling are the real datacenter limits. Tens of kW per air-cooled chassis, over 100 kW per liquid-cooled rack. A thermally throttled GPU silently loses clock speed.

8. Reading GPU memory: the commands#

# The overview
nvidia-smi

# The fields that matter, machine-readable, every second
nvidia-smi --query-gpu=index,name,memory.total,memory.used,memory.free,utilization.gpu,utilization.memory,temperature.gpu,power.draw,clocks_throttle_reasons.active \
           --format=csv -l 1

# Who is using the memory
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv

# Rolling per-second device stats (power, util, clocks, memory, PCIe throughput)
nvidia-smi dmon -s pucvmet

# Topology: which GPU sits on which CPU socket / NUMA node, and GPU-GPU link types
nvidia-smi topo -m

# NVLink status and error counters
nvidia-smi nvlink -s
nvidia-smi nvlink -e

# Memory health
nvidia-smi -q -d ECC,ROW_REMAPPER,PAGE_RETIREMENT

# MIG layout (if enabled)
nvidia-smi -L

# Driver-level errors
dmesg -T | grep -i xid
import torch
free, total = torch.cuda.mem_get_info()        # what the DRIVER says is free (bytes)
torch.cuda.memory_allocated()                  # bytes in live tensors
torch.cuda.memory_reserved()                   # bytes PyTorch holds from the driver
torch.cuda.max_memory_allocated()              # peak — the number you size against
print(torch.cuda.memory_summary())             # allocator breakdown

How the three views relate:

flowchart LR
  T["Live tensors<br/>memory_allocated"] --> R["PyTorch cache<br/>memory_reserved<br/>(allocated + free blocks it keeps)"]
  R --> P["Process total in nvidia-smi<br/>(reserved + CUDA context + libraries)"]
  P --> G["GPU memory.used<br/>(sum over all processes)"]

  class T,R memory
  class P,G neutral

Reading traps

  • utilization.gpu is “percent of time at least one kernel was running”, not how hard the GPU worked (X.04).
  • utilization.memory is “percent of time memory was being read or written” — not how full the memory is. That is memory.used / memory.total.
  • memory.used near 90% on a vLLM/SGLang node is by design (pre-allocated KV pool). The useful signal is the engine’s own KV-cache-usage metric.
  • reserved much greater than allocated → fragmentation or a cache that never shrinks, not a leak in your tensors.

9. Sharing one GPU#

MechanismMemory isolationFault isolationGranularityUse it for
Whole GPU per podFullFull1 GPUProduction LLM serving (default)
MIGHardware-enforced slicesYesUp to 7 fixed slices (e.g. 10 GB each on H100)Small models, strict multi-tenancy
MPSOptional per-client limits, not enforced by hardwareNo — one client’s fault can take down the othersArbitrarySeveral cooperating processes of one tenant
Time-slicingNoneNoArbitrary replicasDev/test, bursty notebooks

A MIG slice gets a fixed share of memory and of memory bandwidth and SMs, so the bandwidth arithmetic in section 5 scales down with the slice. Kubernetes exposure of each mode is in XII.03.


10. When the hardware misbehaves#

SymptomLikely causeFirst check
Tokens/s dropped 20-40% with no deployThermal or power throttlingclocks_throttle_reasons.active, temperature, power cap
Xid 48, Xid 94/95 in dmesgUncorrectable / contained ECC memory errorsnvidia-smi -q -d ECC,ROW_REMAPPER; drain the node
Xid 63/64Row remapping event / failureReset or RMA if remap failed
Xid 79GPU fell off the busPower, seating, PCIe errors; node needs a reboot
Xid 31GPU memory page faultUsually a software bug (bad pointer in a kernel)
Xid 74NVLink errornvidia-smi nvlink -e; affects TP jobs first
One GPU of eight is slowerPCIe link trained down, or wrong NUMA nodelspci -vv link speed/width, nvidia-smi topo -m
OOM at steady load after hoursFragmentation or slow leakmemory_reserved vs memory_allocated over time (X.06)
Slow model loadNetwork storage or cold page cacheMeasure read GB/s; cache weights on local NVMe

Corrected single-bit errors are normal at fleet scale. Rising counts on one device are a reason to drain it before it becomes an uncorrectable one in the middle of someone’s request.


11. Picking a GPU#

flowchart TB
  S["Start: model + precision + context + concurrency"] --> W["Weight bytes + KV bytes<br/>= memory needed"]
  W --> F{"Fits on one GPU<br/>with KV headroom?"}
  F -->|"no"| Q{"Can you quantize<br/>within your quality budget?"}
  Q -->|"yes"| W
  Q -->|"no"| TP["Bigger-memory GPU (H200 / B200 / B300 / Rubin / MI3xx-4xx)<br/>or tensor parallelism over NVLink"]
  F -->|"yes"| L{"Latency SLO<br/>tight on tokens/s per user?"}
  L -->|"yes"| HB["HBM card<br/>bandwidth sets per-user speed"]
  L -->|"no, throughput per dollar"| GD["Cheapest card where it fits<br/>(L40S / A10G / L4) + more replicas"]
  TP --> CHK["Benchmark on the real workload"]
  HB --> CHK
  GD --> CHK

  class W,HB,TP memory
  class GD compute
  class F,Q,L queue
  class S,CHK neutral

Rules of thumb:

  1. Size memory first. If it does not fit with KV room, nothing else matters.
  2. Per-user speed is bandwidth; fleet throughput per dollar is often a cheaper card, replicated. Replicas beat tensor parallelism whenever the model fits on one device (IX.12).
  3. Price per GB of VRAM and price per TB/s are more useful columns than price per TFLOP.
  4. Check the exact SKU (SXM vs PCIe, 40 vs 80 GB) and whether the instance exposes NVLink.
  5. Then measure. Datasheet → prediction → benchmark, in that order.

12. Hands-on#

A. Inventory. On any GPU machine, run every command in section 8. Record model, VRAM, driver, NUMA node, PCIe link speed and width into numbers.md.

B. Budget. Start vLLM (or your Project 08 engine) with a small model. Before sending traffic, account for every GB in nvidia-smi: weights, KV pool, context, the unreserved 10%. Your sum should land within ~1 GB.

C. Predict then measure. Using sections 4 and 5, predict batch-1 tokens/s for your model on your GPU. Measure it. Explain the gap.

D. Fill the table. Add a row to section 3 for a GPU not listed, from the vendor datasheet, and compute its ridge point and its batch-1 ceiling for a 16 GB model.

E. Break it on purpose. Run two processes on one GPU and watch per-process memory. Then set a power cap (nvidia-smi -pl) and re-measure tokens/s.


13. Interview questions#

  1. Why does an H200 serve LLMs faster than an H100 when both have the same FLOPs?
  2. What is HBM, and why do datacenter GPUs use it instead of GDDR?
  3. An 80 GB GPU runs a 16 GB model. Where does the rest of the memory go, and why does nvidia-smi show it full with no traffic?
  4. Estimate batch-1 tokens/s for a 70B FP8 model on an H200. Show the arithmetic.
  5. What is the difference between memory_allocated, memory_reserved, and what nvidia-smi reports?
  6. utilization.memory reads 35%. Is the GPU’s memory 35% full?
  7. MIG vs MPS vs time-slicing — which gives memory isolation, and which would you allow for two different customers?
  8. Why is PCIe bandwidth irrelevant to steady-state decode but critical to cold start and KV offloading?
  9. You see Xid 79 on one node of a tensor-parallel job. What happened and what do you do?
  10. Two cards have the same VRAM. One costs half as much. What single spec do you check before buying it for chat serving?

14. Further reading#

↑↓ navigate ↵ open