Below the API

RAM, DRAM, and NUMA

Foundations Intermediate 1h Difficulty 3/5

Prerequisites 02


1. What is it?#

DRAM is main memory: cheap, large, and slow relative to caches. NUMA (Non-Uniform Memory Access) is what happens on multi-socket servers: each CPU socket has its own memory controller and its own attached DRAM, and reaching the other socket’s memory costs roughly 1.5-2x more latency and less bandwidth.

   ┌─────────────┐  UPI/Infinity Fabric  ┌─────────────┐
   │  Socket 0   │◄─────────────────────►│  Socket 1   │
   │  32 cores   │      ~30-60 GB/s      │  32 cores   │
   └──────┬──────┘                       └──────┬──────┘
          │ ~300 GB/s, 80 ns                    │ ~300 GB/s, 80 ns
    ┌─────▼──────┐                        ┌─────▼──────┐
    │ 512 GB DRAM│                        │ 512 GB DRAM│
    │  (node 0)  │                        │  (node 1)  │
    └────────────┘                        └────────────┘

   Socket 0 reading node 1 memory: ~130 ns, lower bandwidth

GPUs are attached to specific sockets too. Getting this wrong costs you real throughput on every host-to-device transfer.


2. Why does it exist?#

One memory controller cannot feed 128 cores. So each socket gets its own, and the sockets are connected. The result is a shared address space with non-uniform costs — convenient to program, easy to get wrong.

For inference, NUMA matters in three places:

  1. Host-to-device transfers (pinned buffers should be on the GPU’s local NUMA node).
  2. CPU-side work (tokenization pools, the engine loop) touching remote memory.
  3. CPU inference and CPU-offloaded KV cache, which are bandwidth-bound and therefore very sensitive to remote access.

3. Simple analogy#

A two-building office. Your team’s filing cabinet is in your building (30 seconds away). The other team’s is across the courtyard (90 seconds). Everything works either way, but if half your files end up in the wrong building, your day takes twice as long — and nothing in the org chart tells you it’s happening.

NUMA problems are exactly like that: invisible in code, visible only in a profiler or a throughput number that is mysteriously half of what you expected.


4. Tiny example#

# What does the machine look like?
numactl --hardware
lscpu | grep -i numa

# Which NUMA node is each GPU attached to?
nvidia-smi topo -m
# Look for the "NUMA Affinity" column

Typical nvidia-smi topo -m output on an 8-GPU node:

        GPU0  GPU1  GPU2  GPU3  GPU4  GPU5  GPU6  GPU7  CPU Affinity  NUMA
GPU0     X    NV18  NV18  NV18  SYS   SYS   SYS   SYS   0-31,64-95    0
GPU1    NV18   X    NV18  NV18  SYS   SYS   SYS   SYS   0-31,64-95    0
...
GPU4    SYS   SYS   SYS   SYS    X    NV18  NV18  NV18  32-63,96-127  1

Read it: GPUs 0-3 are on NUMA node 0, GPUs 4-7 on node 1. NV18 = 18 NVLink connections; SYS = the traffic must cross the CPU interconnect. This single table tells you how to place your processes and how tensor parallelism will perform (Section IX.08).

Now measure the penalty:

# Local:   run on node 0, allocate on node 0
numactl --cpunodebind=0 --membind=0 ./bandwidth_test
# Remote:  run on node 0, allocate on node 1
numactl --cpunodebind=0 --membind=1 ./bandwidth_test

Expect 1.3-2x worse bandwidth and ~1.6x worse latency in the remote case.


5. Technical explanation#

DRAM basics that matter#

DRAM is organized into channels, ranks, banks, rows. A read activates a whole row (2-8 KB) into a row buffer; subsequent reads in the same row are fast (“row hit”), reads to a different row in the same bank require a precharge + activate (“row miss”, ~2-3x slower).

Practical consequences:

  • Sequential access is fast; random access across banks is much slower than the raw latency number suggests.
  • Bandwidth scales with channels. A CPU with 12 DDR5 channels populated has ~2x the bandwidth of one with 6. Under-populating DIMM slots is a common and expensive procurement mistake in inference hosts, especially for CPU offload workloads.

NUMA policies#

default / first-touch  page is allocated on the node of the thread that FIRST WRITES it
bind                   allocate only on specified nodes (fail if full)
interleave             round-robin pages across nodes (halves worst case, kills best case)
preferred              try one node, fall back

First-touch is the default and the source of most surprises. If a single initialization thread touches all your memory, everything lands on one node, and the other socket’s threads all run remote. Fix: parallelize initialization with the same thread affinity used later.

Pinned memory and GPU transfers#

cudaHostAlloc / torch.empty(pin_memory=True) allocates page-locked host memory, which DMA engines can read directly. If that buffer is on the wrong NUMA node relative to the GPU’s PCIe root complex, every transfer crosses the socket interconnect:

Correct:   pinned buffer on node 0 → PCIe root on node 0 → GPU0        ~25 GB/s
Wrong:     pinned buffer on node 1 → UPI → PCIe root node 0 → GPU0     ~12-18 GB/s

For weight loading (140 GB) that’s the difference between 6 s and 11 s. For per-step transfers in a CPU-offload design, it’s the difference between viable and not.


6. Under the hood#

Check where a running process’s memory actually is:

# Per-node memory usage of a process
numastat -p $(pgrep -f vllm)

# Page-level detail
cat /proc/<pid>/numa_maps | head

# Automatic NUMA balancing (kernel migrates pages; can cause latency spikes)
cat /proc/sys/kernel/numa_balancing

numa_balancing migrates pages toward the accessing thread. Good for long-running steady workloads; a source of periodic latency spikes for latency-sensitive ones. Many HPC and low-latency deployments disable it and pin explicitly instead.


7. Performance implications#

ScenarioPenalty for getting NUMA wrong
Host→device weight load1.5-2x slower cold start
Pinned-buffer streaming (KV offload)1.5-2x less effective bandwidth
CPU inference (bandwidth-bound)up to 2x slower
Tokenizer thread pool1.1-1.3x (small but free to fix)
NCCL over PCIe across socketssignificant; often the reason TP=8 underperforms

8. Production implications#

  • Pin one engine process per NUMA node when running multiple model replicas on a multi-socket, multi-GPU host:
    numactl --cpunodebind=0 --membind=0 python -m vllm.entrypoints.openai.api_server --port 8000 ...
    numactl --cpunodebind=1 --membind=1 python -m vllm.entrypoints.openai.api_server --port 8001 ...
    
  • Keep a tensor-parallel group within one NUMA node / NVLink domain where possible. TP across sockets over PCIe is a common cause of disappointing multi-GPU scaling.
  • Check nvidia-smi topo -m on every new instance type. Cloud SKUs differ, and the topology determines your parallelism plan.
  • In Kubernetes, enable the Topology Manager with single-numa-node policy so pods get CPU, memory, and GPU from the same node (Section XII.03).
  • Populate all memory channels when specifying hosts.

9. Common mistakes#

Ignoring NUMA entirely on a 2-socket box. The most common one. Costs 20-50% on memory-bound host work, silently.

Initializing all memory from one thread. First-touch puts it all on one node.

Interleaving as a default “fix.” It bounds the worst case but prevents the best case. Bind properly instead.

Assuming the container sees the topology. Containers inherit the host’s NUMA layout but may be restricted to CPUs on one node while memory is allocated elsewhere.

Forgetting that GPUs have NUMA affinity too. They hang off a PCIe root complex attached to one socket.


10. Hands-on exercise#

A. Map your machine. Run numactl --hardware, lstopo --output-format txt, nvidia-smi topo -m. Draw the topology on paper: sockets, memory, PCIe roots, GPUs, NVLink.

B. Measure the penalty. Write a simple memory-bandwidth benchmark (STREAM-like triad). Run it with --membind=0 --cpunodebind=0 and --membind=1 --cpunodebind=0. Report the ratio.

C. Measure H2D transfer with NUMA. Allocate pinned host memory on each node (use numactl --membind), transfer 1 GB to GPU 0, and compare bandwidth. Record in numbers.md.

D. Fix a real deployment. If you have a multi-socket GPU box, run a model server with and without numactl binding and compare cold-start time and steady-state throughput.


11. Interview questions#

  1. What is NUMA and what is the typical remote-access penalty?
  2. What is first-touch allocation and how does it cause NUMA problems?
  3. How do you determine which NUMA node a GPU is attached to?
  4. Why does NUMA affect host-to-device transfer bandwidth?
  5. You are running 2 model replicas on a dual-socket 8-GPU box. How do you place them?
  6. When would you use --interleave and why is it usually not the right answer?

12. Further reading#

  • [REFERENCE] numactl(8), numastat(8), man 7 numa
  • [FUNDAMENTAL] Drepper, “What Every Programmer Should Know About Memory,” part 5
  • [REFERENCE] Kubernetes Topology Manager documentation
  • Next: 04 — SIMD and vectorization

↑↓ navigate ↵ open