Below the API

CPU Architecture

Foundations Beginner 1h Difficulty 2/5

Prerequisites I.07, I.09


1. What is it?#

A CPU is a machine that fetches instructions, decodes them, executes them, and writes results. Modern CPUs do all four stages for dozens of instructions simultaneously, out of order, speculatively, across many cores.

For an inference engineer, the CPU is not where the model runs — it is where everything else runs: the HTTP server, the tokenizer, the scheduler, the sampling logic, the Python interpreter driving the GPU. A saturated CPU starves a $30,000 GPU, and that is one of the most common avoidable production problems.


2. Why does it exist?#

Because programs are sequential and unpredictable, and memory is slow. Everything complicated about a modern CPU is a response to one of those two facts:

ProblemCPU’s answer
Instructions depend on each otherpipelining + out-of-order execution
Branches are unpredictablebranch prediction + speculation
Memory is 100x slower than the corecaches, prefetchers
One instruction per value is slowSIMD (file 04)
One thread can’t fill the machineSMT/hyperthreading, multicore

3. Simple analogy#

An assembly line staffed by someone who guesses.

The line has stages: fetch the part, read the blueprint, do the work, inspect, ship. Five items are in the line at once, each at a different stage — that’s pipelining.

At a fork in the instructions (“if the part is red, do X”), the worker doesn’t wait to find out; they guess red, and keep the line moving. If the guess was wrong, everything downstream is thrown away and restarted — a pipeline flush, costing 15-20 cycles. Modern predictors guess right 95-99% of the time, which is why they exist.

The relevance to you: tokenizers and samplers are branch-heavy, so they run far below peak, and the Python interpreter is the worst case — a giant switch statement with unpredictable targets and pointer chasing everywhere.


4. Tiny example#

Branch prediction, measured:

// branch.go — the same loop over the same numbers, in a different order.
package main

import (
	"fmt"
	"math/rand"
	"slices"
	"time"
)

//go:noinline
func work(total *int64, v int32) { *total += int64(v) } // a real call, so the `if` stays a branch

func sumIfGreater(data []int32) (total int64) {
	for _, v := range data {
		if v > 128 {
			work(&total, v)
		}
	}
	return total
}

func main() {
	data := make([]int32, 50_000_000)
	for i := range data {
		data[i] = rand.Int31n(256)
	}

	// Unsorted: the branch is unpredictable
	t0 := time.Now()
	a := sumIfGreater(data)
	unsorted := time.Since(t0)

	slices.Sort(data) // sorted: a long run of "no", then a long run of "yes"
	t0 = time.Now()
	b := sumIfGreater(data)
	sorted := time.Since(t0)

	fmt.Printf("unsorted %v   sorted %v   ratio %.2fx   (same sum: %v)\n",
		unsorted.Round(time.Millisecond), sorted.Round(time.Millisecond),
		float64(unsorted)/float64(sorted), a == b)
}

In Go (or C) this shows a 3-6x difference: same data, same instructions, and the only change is whether the CPU can guess the branch. Write the same loop in Python and the interpreter overhead masks it (you may see 1.1-1.3x) — which is its own lesson: an interpreter is so slow that hardware effects disappear beneath it. A tokenizer written in pure Python is leaving 10-50x on the floor. (This is why Hugging Face’s tokenizers library is written in Rust.)


5. Technical explanation#

The pipeline#

IF  →  ID  →  RENAME  →  DISPATCH  →  EXECUTE  →  WRITEBACK  →  RETIRE
fetch  decode  register    into        in ALUs/    results       in-order
              renaming    scheduler    FPUs/       to regs       commit
                                       load-store

Key properties of a modern core (e.g. Intel Golden Cove, AMD Zen 4):

  • Fetch/decode ~6 instructions per cycle
  • Reorder buffer of 300-500 instructions in flight
  • 8-12 execution ports (several ALUs, 2 FMA/vector units, 2-3 load, 2 store)
  • Peak ~4-6 instructions retired per cycle (IPC)

Real-world IPC: numerical code 2-4; Python interpreter 0.5-1.2; pointer-chasing 0.2-0.5. IPC is the fastest diagnostic of what kind of work you’re doing — perf stat reports it.

Superscalar and out-of-order#

a = b + c;     ← independent
d = e * f;     ← independent      these three can execute in parallel
g = h - i;     ← independent

x = a + d;     ← depends on both; must wait

The scheduler finds this parallelism automatically. Your job is to provide it: avoid long dependency chains, avoid unnecessary serialization.

SMT / Hyperthreading#

Two hardware threads share one core’s execution units. When one stalls on memory, the other uses the idle ports. Typical gain: 15-30% on mixed workloads; negative on bandwidth-saturated or cache-sensitive numerical code, because the two threads evict each other’s cache lines.

For inference hosts: SMT usually helps (the workload is a mix of I/O, Python, and syscalls). For pure CPU inference kernels, test with it off. lscpu shows the topology.

Clocks, turbo, and thermal reality#

A CPU rated at 3.8 GHz may run at 2.4 GHz when all cores are busy with AVX-512, because wide-vector instructions draw more power and trigger frequency offsets. Consequences:

  • Benchmarks run for 2 seconds see turbo; production sees sustained clocks. Always warm up and measure steady state.
  • A microbenchmark of an AVX-512 kernel in isolation can be misleading when the same code runs alongside 60 other threads.

6. Under the hood#

What actually happens when your Python code calls model.generate():

CPU thread:
  interpret bytecode          ← low IPC, branch-heavy, pointer chasing
  build kernel arguments
  cudaLaunchKernel syscall-ish path (user-space, but with driver work)
  ... repeat 200-2000 times per token ...
  occasionally: cudaMemcpy, cudaStreamSynchronize (blocks!)

Every one of those steps is CPU work, and it is on the critical path of every token. This is why:

  • A slow CPU shows up as low GPU utilization with gaps between kernels.
  • CUDA graphs help so much (they collapse thousands of launches into one).
  • Engines move the hot loop out of Python where possible.

Check it in practice: run your server and watch top. If a Python thread is at 100% of a core while GPU utilization is 60%, you are CPU-bound on the launch path.


7. Performance implications#

CPU-side task in an inference serverTypical costRisk
HTTP parse + JSON decode50-500 µshigh at high QPS
Tokenization10 µs - 5 mshigh for long prompts
Scheduler bookkeeping50-500 µs/stephigh at large batch
Kernel launches (no CUDA graph)1-10 ms/stepvery high at small batch
Sampling (if in Python)100 µs - 2 mshigh
Detokenization + SSE framing20-200 µs/tokenhigh at high token rate
Logging / metrics10-100 µsmedium

Add these up for a system emitting 5,000 tokens/sec and you can easily need 4-8 dedicated CPU cores per GPU. Nodes provisioned with 8 CPUs for 8 GPUs will be CPU-starved.


8. Production implications#

  • Provision 8-16 CPU cores per GPU for LLM serving. Check your cloud instance types; some GPU SKUs are CPU-poor.
  • Pin threads. Let the tokenizer pool, the engine loop, and the NCCL threads live on separate, NUMA-local cores (file 03).
  • Use compiled tokenizers (tokenizers Rust library, not a Python loop).
  • Keep the per-token Python path minimal. Custom logits processors written in Python are a classic ITL killer.
  • Monitor CPU alongside GPU. A dashboard with only GPU metrics will not show you the bottleneck.

9. Common mistakes#

Assuming the CPU doesn’t matter because “the model runs on the GPU.” The CPU drives the GPU.

Benchmarking with turbo and no warmup. Overstates sustained performance by 20-40%.

Enabling every thread on every core for a numeric workload. Oversubscription causes context switching and cache thrashing. Set OMP_NUM_THREADS deliberately.

Ignoring hyperthreading effects. Test both ways for CPU-heavy paths.

Writing hot-path logic in Python. A per-token Python callback executed 5,000 times/sec at 100 µs each consumes half a core and adds directly to ITL.


10. Hands-on exercise#

A. Read your CPU. Run lscpu, cat /proc/cpuinfo | head -30, lstopo (from hwloc). Write down: cores, threads/core, sockets, L1/L2/L3 sizes, and supported vector ISAs (avx2, avx512, amx).

B. Measure IPC.

perf stat -e cycles,instructions,branch-misses,cache-misses python your_script.py

Do it for (i) a pure Python loop, (ii) a NumPy operation, (iii) your model server under load. Compare IPC. Explain the differences.

C. Branch prediction in C. Write the sorted/unsorted branch benchmark in C, compile with -O2, and measure. Confirm the 3-6x. Then check perf stat -e branch-misses.

D. Find CPU starvation. Run a model server under load with nvidia-smi dmon and top side by side. Is GPU utilization limited by CPU? How would you prove it?


11. Interview questions#

  1. Why does a fast GPU need a fast CPU for LLM inference?
  2. What is IPC and what does a low IPC tell you about a workload?
  3. Explain branch prediction and give an inference-relevant example where it matters.
  4. When does hyperthreading hurt?
  5. How many CPU cores would you provision per GPU for LLM serving, and why?
  6. You see gaps between CUDA kernels in an Nsight timeline. List three causes.

12. Further reading#

↑↓ navigate ↵ open