The idea in one minute#
A GPU does not float in space. It is attached to a host through PCIe, it draws hundreds of watts, and every one of those watts becomes heat that must be removed. These “boring” physical facts set real limits: how fast you can load a model, how many GPUs fit in a server, and whether a GPU quietly slows itself down under load.
A picture#
flowchart LR
subgraph HOST["Host"]
APP["Go program<br/>user space"] --> DRV["NVIDIA driver<br/>kernel space"]
RAM[("Pinned host memory")]
end
PCIE["PCIe x16<br/>about 32 to 64 GB/s"]
subgraph DEV["GPU"]
CQ["Command queue"] --> SMS["SMs"]
HBM[("HBM")]
end
DRV -->|"commands"| PCIE --> CQ
RAM <-->|"DMA: data"| PCIE
PCIE <--> HBM
PSU["Power: 300 to 1000 W"] --> DEV
DEV --> HEAT["Heat out<br/>air or liquid"]
class APP,DRV,SMS compute
class RAM,HBM memory
class PCIE,CQ queue
class PSU,HEAT warnHow it really works#
PCIe: the host link#
PCIe (Peripheral Component Interconnect Express) is the bus that plugs the GPU into the motherboard. Its speed depends on the generation and the number of lanes (x16 is standard for GPUs):
| Generation | x16 raw speed | Seen in |
|---|---|---|
| PCIe 3.0 | ~16 GB/s | Older servers, T4 |
| PCIe 4.0 | ~32 GB/s | A100 era |
| PCIe 5.0 | ~64 GB/s | H100 era |
| PCIe 6.0 | ~128 GB/s | Newest platforms |
Real copies reach perhaps half to three quarters of the raw figure. Compare with the 3,350 GB/s inside an H100 and the rule from II.02 is obvious: PCIe is the narrow bridge.
Two details matter in practice:
- DMA and pinned memory. The GPU reads host memory directly (direct memory access) without the CPU copying bytes. That only works if the operating system promises not to move those pages — the memory must be pinned (page-locked). Copies from ordinary pageable memory need an extra hidden copy and run about half as fast.
- Commands cross too. Every kernel launch is a small message through the driver and over PCIe. It costs a few microseconds. Tiny kernels launched thousands of times can spend more time on launches than on work (IV.04).
Power#
| GPU | Power (TDP) |
|---|---|
| T4 | 70 W |
| L4 | 72 W |
| RTX 4090 | 450 W |
| A100 | 400 W |
| H100 | 700 W |
| B200 | ~1,000 W |
| B300 (Blackwell Ultra) | ~1,400 W |
| Rubin | ~1,800–2,300 W (reported; liquid-cooled only) |
An eight-GPU H100 server draws around 10 kW — as much as several homes. Power, more than money or floor space, is now the limiting resource for building AI data centers.
Performance per watt is therefore a first-class metric. Small cards like the L4 are slower than an H100 but do more work per joule on small models, which is why they exist.
Heat and throttling#
All input power becomes heat. If cooling cannot keep up, the GPU protects itself by lowering its clock speed: thermal throttling. Nothing crashes. The GPU simply becomes 10–40% slower, often only under sustained load, often only on the GPUs in the hottest part of the chassis.
This is a classic source of “the same code is slower on node 7” mysteries. The signs are in the monitoring data: temperature near the limit and clock frequency below nominal.
Above roughly 40 kW per rack, air cannot carry the heat away and data centers switch to liquid cooling (cold plates on the chips). Current flagship racks draw well over 100 kW and are liquid-cooled by design (VII.03).
Reliability#
GPU memory uses ECC (error-correcting codes) to fix single-bit errors caused by radiation and wear. Uncorrectable errors and “fell off the bus” link failures do happen at fleet scale. NVIDIA reports them as numbered XID events in the kernel log; a production system watches for them and drains the affected GPU.
Code#
Estimate load time and running cost from the physical numbers.
// physical.go — load time and electricity, from the facts on the label.
package main
import "fmt"
func main() {
// 1. How long to get a model onto the GPU?
const modelGB = 140.0 // a 70B model in FP16
for _, link := range []struct {
name string
gbps float64
}{{"NVMe -> RAM", 5}, {"RAM -> GPU, pageable", 12}, {"RAM -> GPU, pinned (PCIe 5)", 40}} {
fmt.Printf("%-30s %6.1f s\n", link.name, modelGB/link.gbps)
}
// 2. What does a server cost to power?
const (
gpus, wattsPerGPU = 8, 700.0
otherWatts = 2500.0 // CPUs, fans, drives, network
pue = 1.3 // data-center overhead: cooling, power conversion
dollarsPerKWh = 0.10
)
kw := (gpus*wattsPerGPU + otherWatts) * pue / 1000
fmt.Printf("\nserver draw incl. cooling: %.1f kW\n", kw)
fmt.Printf("electricity per year: $%.0f\n", kw*24*365*dollarsPerKWh)
}The slowest step in loading is the disk, not PCIe — a hint about where cold-start time goes in real serving systems.
Remember this#
- PCIe is tens of GB/s; HBM is thousands. Data should cross once, from pinned memory.
- Every kernel launch crosses the driver and PCIe: a few microseconds each.
- Power is the scarce resource of AI data centers; performance per watt matters.
- Hot GPUs throttle silently. Check temperature and clocks before blaming software.
Try it#
- Run
physical.gowith your region’s electricity price. - If you have a GPU: run
nvidia-smi -qand find temperature, power draw, the PCIe link generation and width, and any ECC error counts. - A rack can supply 40 kW. How many 8×H100 servers fit, electrically?
Check yourself#
- Why must host memory be pinned for the fastest transfers?
- What happens when a GPU overheats? How would you notice?
- Which is the slowest step when cold-loading a large model?