1. What is it?#
The physical links between GPUs, and between nodes. These numbers determine which parallelism strategies are viable.
Link Bandwidth (per GPU) Latency Scope
Within-GPU HBM 3,350 GB/s ~0.5 µs —
NVLink 4 (H100) 900 GB/s bidir ~2 µs within a node
NVLink 3 (A100) 600 GB/s bidir ~2 µs within a node
NVSwitch (all-to-all NVLink) full bandwidth ~2 µs within a node
PCIe Gen5 x16 128 GB/s bidir ~5-10 µs within a node
PCIe Gen4 x16 64 GB/s bidir ~5-10 µs within a node
InfiniBand NDR (400 Gb/s) 50 GB/s ~2-3 µs across nodes
InfiniBand HDR (200 Gb/s) 25 GB/s ~2-3 µs across nodes
RoCE (100 GbE) 12.5 GB/s ~5-10 µs across nodes
Ethernet (25 GbE, no RDMA) 3 GB/s ~50-100 µs across nodesRatio HBM : NVLink : PCIe : IB ≈ 67 : 18 : 1.3 : 1 (using Gen4 as the unit).
2. Why the numbers matter#
Because each parallelism strategy has a communication volume, and dividing volume by bandwidth gives you time. That time either fits in your budget or doesn’t.
Llama-3-70B decode, batch 32, TP=8, per token:
communication volume: 143 MB (Section IX.03)
compute+weight-read time: 5.3 ms
NVLink 4: 143 MB / 450 GB/s = 0.32 ms → 6% overhead ✓
PCIe Gen5: 143 MB / 64 GB/s = 2.2 ms → 42% overhead ✗
PCIe Gen4: 143 MB / 32 GB/s = 4.5 ms → 85% overhead ✗✗
InfiniBand: 143 MB / 50 GB/s = 2.9 ms → 55% overhead ✗One table entry decides whether TP=8 is a good idea. This is why “check the topology” appears in every file of this section.
3. Simple analogy#
Roads between workshops.
- HBM — the workbench. Instant.
- NVLink — a wide corridor between adjacent rooms. Fast enough that cooperating on a single task is practical.
- PCIe — a normal doorway with a queue. You can pass things through, but not continuously.
- InfiniBand — a road between buildings. Fine for delivering completed work; hopeless for passing a tool back and forth 160 times per minute.
- Ethernet without RDMA — the same road, but every delivery requires unloading into a warehouse first and reloading (the host memory bounce).
The rule: cooperate tightly over NVLink; hand off completed work over the network. That’s TP within a node, PP or DP across nodes.
4. NVLink and NVSwitch#
NVLink is a point-to-point link. On a DGX/HGX system, NVSwitch chips provide
ALL-TO-ALL connectivity at full bandwidth:
Without NVSwitch (e.g. 4-GPU PCIe servers with NVLink bridges):
GPU 0 ←→ GPU 1 direct
GPU 0 → GPU 2 must route through... nothing. Falls back to PCIe.
With NVSwitch (DGX H100, 8 GPUs):
every GPU ←→ every GPU at 900 GB/s
→ any TP grouping works equally wellCheck which you have. nvidia-smi topo -m:
GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7
GPU0 X NV18 NV18 NV18 NV18 NV18 NV18 NV18 ← NVSwitch: all NV18
GPU1 NV18 X NV18 NV18 NV18 NV18 NV18 NV18
...versus:
GPU0 GPU1 GPU2 GPU3
GPU0 X NV4 SYS SYS ← only GPU0-GPU1 have NVLink
GPU1 NV4 X SYS SYS GPU0-GPU2 goes through the CPU (SYS)
GPU2 SYS SYS X NV4
GPU3 SYS SYS NV4 XIn the second case, TP=4 is a bad idea — half the communication crosses the CPU. TP=2 within each NVLink pair, plus DP=2, is much better.
Legend:
NV# NVLink with # links
PIX same PCIe switch (good)
PXB multiple PCIe switches (ok)
PHB through the PCIe host bridge (worse)
NODE across NUMA nodes within a socket
SYS across the CPU interconnect (worst)5. Across nodes: RDMA and GPUDirect#
WITHOUT RDMA (plain TCP/IP):
GPU memory → host memory (PCIe)
→ kernel network stack (CPU)
→ NIC → network → NIC
→ kernel stack (CPU)
→ host memory
→ GPU memory (PCIe)
Two PCIe crossings, two memory copies, CPU involvement on both ends.
Effective bandwidth: maybe 40% of link rate. Latency: 50-100 µs.
WITH RDMA (InfiniBand or RoCE):
GPU memory → NIC (via GPUDirect RDMA, direct DMA over PCIe)
→ network → NIC → GPU memory
No host memory, no CPU. Effective bandwidth: 85-95% of link rate.
Latency: 2-5 µs.GPUDirect RDMA is not optional for multi-node inference. Verify it’s active:
# Is the module loaded?
lsmod | grep nvidia_peermem # or nv_peer_mem on older systems
# Does NCCL use it?
NCCL_DEBUG=INFO python your_job.py 2>&1 | grep -i "GDRDMA\|GPU Direct"
# Look for: "NCCL INFO ... [receive] via NET/IB/0/GDRDMA"
# ^^^^^^ this is what you want
If you see via NET/IB/0 without GDRDMA, you’re bouncing through host memory.
InfiniBand vs RoCE vs Ethernet#
InfiniBand purpose-built, lowest latency, native RDMA, expensive
→ the default for GPU clusters
RoCE v2 RDMA over Converged Ethernet. Cheaper, uses Ethernet
switches, but REQUIRES lossless configuration (PFC/ECN).
Misconfigured RoCE performs terribly and is hard to debug.
Plain Ethernet no RDMA. 3-5x worse effective bandwidth, 20x worse latency.
Acceptable for DP (no inter-GPU communication) only.If your cluster is plain Ethernet, do not attempt cross-node TP or EP. Use DP across nodes and keep tight coupling within them.
6. The design rules that follow#
RULE 1: Tensor parallelism ONLY within an NVLink domain.
TP across PCIe or network: 40-85% overhead.
RULE 2: Pipeline parallelism across nodes.
Its communication is 100x smaller and tolerates network latency.
RULE 3: Data parallelism anywhere.
Zero inter-replica communication.
RULE 4: Expert parallelism prefers within-node.
All-to-All across nodes is the dominant cost in large MoE.
RULE 5: Verify GPUDirect RDMA for any cross-node collective.
RULE 6: The canonical layout for large models:
TP = (GPUs per NVLink domain, usually 8)
PP = across nodes, as few stages as possible
DP = replicate the whole TP×PP unit7. Performance — measuring your links#
# GPU-to-GPU bandwidth and latency (from cuda-samples)
./p2pBandwidthLatencyTest
# Host-device
./bandwidthTest --memory=pinned --mode=range --start=1024 --end=1073741824
# Collectives (from nccl-tests)
./build/all_reduce_perf -b 8 -e 1G -f 2 -g 8
# Cross-node bandwidth (InfiniBand)
ib_write_bw -d mlx5_0 -a # on one node
ib_write_bw -d mlx5_0 -a <host> # on the other
# Cross-node latency
ib_write_lat -d mlx5_0
Record everything in numbers.md. These are the constants for every distributed capacity
estimate you’ll make.
Representative results on a DGX H100:
p2p bandwidth (NVLink): ~370 GB/s unidirectional per pair
p2p latency: ~1.8 µs
H2D pinned (PCIe Gen5): ~55 GB/s
AllReduce 512 KB, 8 GPUs: ~35 µs
AllReduce 64 MB, 8 GPUs: ~880 µs (146 GB/s busbw)
IB write bandwidth (NDR): ~48 GB/s
IB write latency: ~1.9 µs8. Production implications#
- Verify the topology on every new instance type. Cloud SKUs with the same GPU can have wildly different interconnects. An 8×A100 instance may or may not have NVSwitch.
- Check for degraded PCIe links (
nvidia-smi -q | grep -A4 "GPU Link Info"). A GPU negotiated at x8 instead of x16 halves its host bandwidth silently. - In Kubernetes, use the Topology Manager with
single-numa-nodepolicy so a pod’s GPUs come from one NVLink domain (Section XII.03). - Budget for the network in multi-node deployments. InfiniBand adds real cost; plain Ethernet constrains your architecture.
- Test GPUDirect RDMA explicitly after any driver or kernel update. It silently regresses.
- Document the topology assumptions in your deployment configuration. “This config requires NVSwitch” should be written down.
9. Common mistakes#
Assuming all 8 GPUs in a node are NVLink-connected. Many aren’t.
TP across a SYS link. 85% overhead.
Multi-node without GPUDirect RDMA. Half the bandwidth, 20x the latency.
Misconfigured RoCE. Without lossless configuration it performs worse than TCP.
Not noticing a degraded PCIe link.
Deploying the same parallelism config across heterogeneous instance types.
Ignoring NUMA affinity for the pinned buffers used in host-device transfers (Section II.03).
10. Hands-on exercise#
A. Map your hardware. Run nvidia-smi topo -m, p2pBandwidthLatencyTest, and
all_reduce_perf. Draw the topology diagram and annotate it with measured bandwidths. Save
this — it’s the reference for every distributed decision you make on this hardware.
B. Verify link health. Check every GPU’s PCIe link width and generation. Are any degraded?
C. Quantify NVLink’s value. Run TP=2 on an NVLink-connected pair and on a SYS-connected
pair (or simulate with NCCL_P2P_DISABLE=1). Measure ITL and throughput for both.
D. Multi-node. If you have access, measure ib_write_bw between nodes, then run a TP job
across nodes and compare to within-node. Confirm the design rule.
E. GPUDirect verification. Run a multi-node NCCL job with NCCL_DEBUG=INFO. Confirm GDRDMA
is in use. Then disable nvidia_peermem (on a test system) and measure the difference.
11. Interview questions#
- Rank the interconnects by bandwidth and give approximate numbers.
- Why can’t you do tensor parallelism across nodes?
- What is NVSwitch and why does its presence change your TP strategy?
- What is GPUDirect RDMA and what does it save?
- How do you verify that NVLink is actually being used between two GPUs?
- What does
SYSmean innvidia-smi topo -mand what are its implications? - Given a cluster with plain Ethernet between nodes, how would you deploy a 70B model?
12. Further reading#
- [REFERENCE] NVIDIA NVLink and NVSwitch technical briefs
- [REFERENCE] NVIDIA GPUDirect RDMA documentation
- [REFERENCE]
nvidia-smi topo, cuda-samplesp2pBandwidthLatencyTest,nccl-tests,perftest(ib_write_bw) - Next: 09 — Communication cost math