Below the API

CPU vs GPU

Foundations Beginner 1h Difficulty 2/5

Prerequisites 07, 08


1. What is it?#

Two fundamentally different bets about what a processor is for.

CPU: a few very smart cores that finish one task as fast as possible
GPU: thousands of simple cores that finish enormous amounts of uniform work
        CPU (e.g. 32-core server)          GPU (e.g. H100)
        ┌────┬────┬────┬────┐              ┌──────────────────────────┐
        │ big│ big│ big│ big│              │▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪│
        │core│core│core│core│              │▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪│
        ├────┴────┴────┴────┤              │▪▪▪▪▪▪▪ 16,896 ▪▪▪▪▪▪▪▪▪▪│
        │   huge caches     │              │▪▪▪▪ CUDA cores ▪▪▪▪▪▪▪▪▪│
        │   branch predict  │              │▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪│
        │   out-of-order    │              ├──────────────────────────┤
        └───────────────────┘              │  small caches, huge BW   │
         ~0.3 TB/s to DRAM                 └──────────────────────────┘
                                            3.35 TB/s to HBM

For inference, the GPU wins — but understanding why tells you when it doesn’t.


2. Why does it exist?#

Historical accident turned into a deliberate divergence.

CPUs were designed for general programs: unpredictable branches, pointer chasing, system calls, one thing at a time as fast as possible. Enormous transistor budgets went into making a single instruction stream fast: branch predictors, out-of-order execution, speculative execution, deep caches. Perhaps 1% of a CPU’s die area is arithmetic units.

GPUs were designed to color millions of pixels with the same shader. Every pixel is independent and does the same operations. So: no branch prediction, no out-of-order, tiny caches per core, and spend the entire transistor budget on arithmetic and memory bandwidth. Perhaps 30-50% of a GPU’s die is arithmetic.

Neural networks turned out to be shaped exactly like graphics: enormous amounts of identical, independent arithmetic on large arrays. That coincidence is the reason NVIDIA is the most valuable company in the industry.


3. Simple analogy#

A brilliant professor versus a stadium full of students with calculators.

Ask “what’s the derivative of this unfamiliar function?” — the professor answers instantly, the stadium is useless (they’d have to coordinate, and only one person can think about it).

Ask “multiply these 10 million pairs of numbers” — the professor takes weeks, the stadium finishes in seconds.

Now the crucial extra detail that most versions of this analogy omit: the stadium’s real constraint is not thinking, it’s the aisles. Getting 10 million numbers to 50,000 students requires enormous throughput of paper. If the aisles can only carry 100 sheets/second, the students sit idle no matter how fast they compute. That’s memory bandwidth, and it is why Sections VII and XIII exist.


4. Tiny example#

// matmul.go — how far one CPU core gets on the GPU's favourite workload.
package main

import (
	"fmt"
	"time"
)

func main() {
	const n = 1024
	a, b, c := make([]float32, n*n), make([]float32, n*n), make([]float32, n*n)
	for i := range a {
		a[i], b[i] = float32(i%7), float32(i%5)
	}
	t0 := time.Now()
	for i := 0; i < n; i++ {
		for k := 0; k < n; k++ {
			aik := a[i*n+k]
			for j := 0; j < n; j++ {
				c[i*n+j] += aik * b[k*n+j]
			}
		}
	}
	dt := time.Since(t0).Seconds()
	fmt.Printf("go, 1 core, float32: %.0f ms  %.4f TFLOP/s\n", dt*1e3, 2*n*n*n/dt/1e12)
}

That prints roughly 0.007 TFLOP/s. Now the same multiplication (at 8192×8192) through an optimized BLAS library on a 32-core Xeon, and on an A100:

go, 1 core, plain loops      0.007 TFLOP/s
cpu, 32 cores, AVX-512 BLAS   0.60 TFLOP/s      86x
gpu, float32                 37.3  TFLOP/s    5300x
gpu, float16 (tensor cores) 250.0  TFLOP/s   35700x

Three lessons in one table. Library-grade CPU code (SIMD + all cores, Section II.04) is ~90x faster than a simple loop. The GPU is ~60x faster again at the same precision, and using the GPU’s specialized units (tensor cores, via FP16) gives another 7x on top. Buying a GPU and running FP32 leaves most of the value on the table — a mistake real teams make.

Now the reverse experiment:

// A branchy, sequential, data-dependent workload
func collatzSteps(n uint64) int {
	c := 0
	for n != 1 {
		if n%2 == 0 {
			n /= 2
		} else {
			n = 3*n + 1
		}
		c++
	}
	return c
}

Try to make that fast on a GPU. You can’t, meaningfully: every element takes a different number of iterations (warp divergence, Section VI.10), and there’s no arithmetic density. The CPU wins easily. GPUs are not “faster computers.” They are faster at one specific shape of work.


5. Technical explanation#

The comparison table#

PropertyServer CPU (2024)H100 SXM
Cores32-128 heavyweight132 SMs × 128 lanes = 16,896
Threads in flight64-256~250,000
Clock2.5-3.8 GHz1.4-1.8 GHz
Peak FP32~2-4 TFLOP/s67 TFLOP/s
Peak FP16 (tensor)~4-8 TFLOP/s (AMX)990 TFLOP/s
Memory0.5-2 TB DDR580 GB HBM3
Memory bandwidth~0.3-0.5 TB/s3.35 TB/s
Cache100-300 MB L350 MB L2
Latency to memory~80 ns~500 ns (hidden by threads)
Branch predictionexcellentnone
Cost$5-15k$25-40k
Power250-350 W700 W

The two numbers that matter for inference: bandwidth (7-10x) and tensor FLOPs (100x+).

Latency hiding vs latency avoidance#

This is the deepest architectural difference and worth understanding precisely.

CPU strategy — AVOID latency:
  big caches so most accesses never reach DRAM
  prefetchers that predict what you'll need
  out-of-order execution to keep going past a stall
  → optimizes for one thread that must not stall

GPU strategy — HIDE latency:
  when a warp stalls on memory, instantly switch to another warp
  keep 64 warps resident per SM so there is always one ready
  → optimizes for aggregate throughput; individual threads stall constantly

Consequence: a GPU needs enough parallel work to hide its latency. Give it a batch of 1 and you get a fraction of its capability — not because the hardware is broken, but because you gave it nothing to switch to. This is the architectural reason batching matters, complementing the bandwidth argument from file 06.

When the CPU is actually the right answer#

Genuine cases, not consolation prizes:

CaseWhy
Small models (< 1B params) at low QPSGPU fixed overheads (launch, transfer) dominate
Sparse / irregular models (trees, GBDT)branchy; GPUs hate it
Very strict cost floor, low traffica GPU idle at 2% utilization is pure waste
Embedding lookups on huge tablesrandom access to hundreds of GB; capacity beats bandwidth
Edge / on-device with no GPUno choice
Preprocessing, tokenization, sampling logicinherently serial and cheap
Very large models with offloading, batch-1when the model doesn’t fit, DRAM capacity wins over HBM speed

Modern CPUs with AMX (Intel) or SVE (ARM) plus INT8 do respectably on small transformers. Llama.cpp on an Apple M-series chip is a real, useful system — Apple silicon’s unified memory gives ~400-800 GB/s, closer to a low-end GPU than to a typical server CPU.

The transfer tax#

The GPU is behind a bus:

CPU DRAM ──── PCIe Gen4 x16 (~25 GB/s real) ────► GPU HBM (3350 GB/s)
                        ↑
                134x slower than HBM

Moving 1 GB to the GPU costs ~40 ms — longer than several decode steps. Rules that follow:

  1. Load weights once, keep them resident. Never stream weights per request.
  2. Transfer only tokens, not text. Tokenize on CPU; send small integer arrays.
  3. Overlap transfers with compute using pinned memory and separate CUDA streams.
  4. Never round-trip mid-model. A CPU-side operation in the middle of a forward pass costs two transfers plus a synchronization; it can easily dominate the entire step.

That last one is a real and common performance bug: a custom logits processor implemented in Python forces a device→host copy and a sync every single decode step.


6. Under the hood#

A GPU kernel’s life:

1. CPU: cudaLaunchKernel(...)         ~5 µs of CPU work
2. Command written to a queue in pinned memory
3. GPU front-end fetches it, allocates blocks to SMs
4. Each SM launches warps; warp scheduler picks ready warps every cycle
5. Warps issue loads; stall; scheduler switches to other warps
6. Results written back to HBM
7. (optional) completion signal / event

Step 1 is why decode is launch-bound: a 70B model runs ~500-2,000 kernels per token, and at 5 µs each that is 2.5-10 ms of pure CPU launch overhead — potentially more than the GPU work itself at small batch. CUDA graphs (Section VI.07) collapse those thousands of launches into one, and routinely give 20-40% end-to-end speedups on small-batch decode. That is one of the highest return-on-effort optimizations available, and it exists purely because of the CPU-GPU boundary.


7. Performance implications#

For a 7B FP16 model, single request:

                     CPU (32-core)      GPU (A100)
Weights memory       14 GB DDR          14 GB HBM
Bandwidth            0.35 TB/s          2.0 TB/s
Decode floor         40 ms/token        7 ms/token
Realistic            60-100 ms/token    10-15 ms/token
Tokens/sec           10-16              65-100
Max useful batch     ~4-8               ~256
Cost/hour            $1.50              $2.50
Cost per 1M tokens   ~$30               ~$8

The GPU is ~6x faster per user and ~4x cheaper per token even though the hardware costs more — because it is far better utilized. That is the general shape of the argument.

Where CPUs close the gap: small models and low traffic. At 0.5 requests/sec with a 300M model, the GPU sits idle 99% of the time and the CPU wins on cost outright.


8. Production implications#

  • Right-size the accelerator to the model and traffic. An L4 or A10G serving a 7B model at moderate traffic is often better economics than an H100. Not everything needs the flagship.
  • Keep the CPU busy with CPU work. Tokenization, detokenization, request parsing, JSON schema validation for structured output, image preprocessing. Under-provisioning CPU on a GPU node is a classic bottleneck — you will see GPU utilization dips that correlate with CPU saturation.
  • Watch the CPU→GPU boundary in profiles. Nsight Systems showing gaps between kernels with busy CPU means you are launch-bound or Python-bound.
  • Consider CPU inference for the long tail. A platform serving 200 models can put the 150 rarely-used small ones on CPU nodes and reserve GPUs for the hot ones (Section XII).
  • Unified-memory systems (Apple, Grace-Hopper) change the calculus by removing or widening the transfer bottleneck. Grace-Hopper’s NVLink-C2C gives ~900 GB/s CPU↔GPU, making CPU-offloaded KV cache genuinely viable (Section XIII.07).

9. Common mistakes#

“GPU is always faster.” Not for small models, small batches, branchy code, or random access. Measure.

Running FP32 on a GPU with tensor cores. You are leaving 5-15x on the table. Use FP16/BF16 at minimum.

Under-provisioning CPU or host RAM on GPU nodes. Tokenization, the scheduler, and the HTTP stack all run on CPU. A saturated CPU starves an expensive GPU.

Transferring data every request. Weights resident, tokens transferred. If you see multi-MB H2D copies per request in a profile, something is wrong.

Ignoring launch overhead. At batch 1 it can exceed compute. Use CUDA graphs.

Comparing a tuned GPU path against an untuned CPU path (or vice versa). CPU inference with oneDNN/AMX/INT8 is 5-10x faster than naive PyTorch FP32 on CPU. Compare tuned to tuned or your conclusion is meaningless.


10. Hands-on exercise#

A. Benchmark both. Run the matmul benchmark on CPU FP32, GPU FP32, GPU FP16, GPU BF16. Compute achieved TFLOP/s and the fraction of spec peak for each. Record in numbers.md.

B. Measure the transfer tax. Time tensor.cuda() for sizes from 1 KB to 1 GB, with and without pinned memory (torch.empty(..., pin_memory=True)). Plot GB/s vs size. At what size do you reach peak PCIe bandwidth? What is the fixed overhead per transfer?

C. Find the crossover. Run a small model (e.g. a 125M-parameter transformer) on CPU and GPU at batch sizes 1, 4, 16, 64. At which batch size does the GPU win? Why isn’t it batch 1?

D. Launch overhead. Time 1,000 tiny kernel launches (x + 1 on a 16-element tensor). Divide by 1,000. That is your launch overhead. Multiply by the number of kernels in a decode step for a 32-layer model (~200-400). Compare to your measured ITL.


11. Interview questions#

  1. Why are GPUs faster at inference than CPUs? Give two independent reasons.
  2. Explain latency hiding vs latency avoidance. Which does each architecture use and why?
  3. When would you choose CPU inference? Give three concrete scenarios.
  4. What is the cost of a CPU↔GPU transfer, and how do you avoid paying it per request?
  5. Why does a GPU need large batches to be efficient — give both the bandwidth and the occupancy explanation.
  6. A 7B model runs at 12 tokens/sec on GPU at batch 1. Is that good? How would you tell?
  7. Your GPU utilization dips to 40% under load while CPU is pegged. Diagnose.

12. Further reading#

  • [FUNDAMENTAL] Kirk & Hwu, Programming Massively Parallel Processors, ch. 1-4
  • [REFERENCE] NVIDIA H100 architecture whitepaper
  • [FUNDAMENTAL] “Computer Architecture: A Quantitative Approach,” ch. 4 (data-level parallelism)
  • [ESTABLISHED] llama.cpp — read its CPU kernels to see how far CPU inference can be pushed
  • Next: 10 — Why inference is not normal backend engineering

↑↓ navigate ↵ open