Below the API

Latency vs Throughput

Foundations Beginner 40 min Difficulty 2/5

Prerequisites 01

The idea in one minute#

Chip designers get a fixed budget of transistors and watts. A CPU spends that budget on making one instruction stream finish as soon as possible (low latency). A GPU spends it on doing as much total arithmetic per second as possible (high throughput).

Neither is “better”. They are answers to different questions.

An analogy#

A sports car carries two people at 250 km/h. A bus carries sixty people at 80 km/h.

To move one person across town, take the car. To move a stadium, the bus wins by a wide margin even though every individual passenger travels slower.

A CPU core is the car. A GPU is a fleet of buses.

A picture#

flowchart TB
  subgraph CPU["CPU core: spend transistors on being clever"]
    direction LR
    BP["Branch<br/>predictor"] --> OOO["Out-of-order<br/>engine"] --> ALU1["A few<br/>arithmetic units"]
    CACHE[("Big private caches")]
  end
  subgraph GPU["GPU block: spend transistors on arithmetic"]
    direction LR
    CTRL["One small<br/>control unit"] --> L1["lane"] & L2["lane"] & L3["lane"] & L4["... x128"]
  end
  class BP,OOO,CTRL neutral
  class ALU1,L1,L2,L3,L4 compute
  class CACHE memory

How it really works#

Where a CPU core’s transistors go#

Most of a modern CPU core is not arithmetic. It is machinery for guessing:

  • A branch predictor guesses which way every if will go, so work can start early.
  • An out-of-order engine looks hundreds of instructions ahead and runs whichever are ready.
  • Large caches keep recently used data a nanosecond away.

All of it exists to make unpredictable, branchy, pointer-chasing code — a web server, a compiler, a database — run fast on one thread.

Where a GPU’s transistors go#

A GPU deletes almost all of that. No sophisticated prediction, no reordering, small caches. The saved space is filled with arithmetic units and the wiring to feed them. One control unit steers many lanes at once, because all lanes run the same instruction (module II explains this as SIMT).

Server CPU (64 cores)Data-center GPU (H100)
Independent control units64132
Arithmetic lanes64 × 16 (AVX-512) ≈ 1,000132 × 128 ≈ 17,000
Clock speed~3.5 GHz~1.7 GHz
Peak FP32 arithmetic~3 TFLOP/s~67 TFLOP/s
Peak with AI-specific units~10 TFLOP/s~1,000 TFLOP/s (FP16 tensor cores)
Memory bandwidth~0.3 TB/s~3.35 TB/s
Handles branchy codeExcellentPoor

Read the table as a trade, not a ranking. The GPU’s lanes are individually slower and cannot make independent decisions cheaply. It wins only when all the lanes have identical work.

Amdahl’s law: the part you cannot parallelize#

If a fraction p of a job can be spread over n workers and the rest cannot:

speedup = 1 / ((1 - p) + p / n)

With p = 0.95 and unlimited workers, the best possible speedup is 1 / 0.05 = 20x. The serial 5% becomes the whole runtime. This is why a GPU program is always a partnership: the CPU does the serial part, and you fight to keep that part tiny.

Code#

// amdahl.go — why 17,000 lanes do not give a 17,000x speedup.
package main

import "fmt"

func speedup(parallelFraction float64, workers float64) float64 {
	return 1 / ((1 - parallelFraction) + parallelFraction/workers)
}

func main() {
	fmt.Println("parallel%   8 workers   64 workers   17000 workers")
	for _, p := range []float64{0.50, 0.90, 0.99, 0.999, 0.9999} {
		fmt.Printf("%7.2f%%  %9.1fx  %10.1fx  %13.1fx\n",
			p*100, speedup(p, 8), speedup(p, 64), speedup(p, 17000))
	}
}

Notice the last column: going from 99% to 99.99% parallel changes the GPU speedup from about 100x to about 6,300x. On a GPU, the small serial remainder is everything.

Remember this#

  • CPU = minimum time for one stream of work. GPU = maximum total work per second.
  • A GPU gets its arithmetic density by removing the “clever” parts of a CPU core.
  • A GPU needs uniform, independent, arithmetic-heavy work to win.
  • The serial fraction of your job caps the speedup (Amdahl’s law).

Try it#

  1. Run amdahl.go. What parallel fraction do you need before 17,000 lanes beat 64 workers by at least 10x?
  2. Think of a program you wrote recently. Estimate its parallel fraction. Would a GPU help?
  3. From the table, compute bytes of memory bandwidth available per FLOP for the CPU and for the GPU. Which one is more “starved” for memory? Keep the answer — module IV is built on it.

Check yourself#

  1. What does a CPU spend most of its transistors on, and why?
  2. Why is a single GPU lane slower than a CPU core?
  3. A job is 90% parallel. What is the maximum speedup with infinite hardware?

↑↓ navigate ↵ open