The idea in one minute#
Chip designers get a fixed budget of transistors and watts. A CPU spends that budget on making one instruction stream finish as soon as possible (low latency). A GPU spends it on doing as much total arithmetic per second as possible (high throughput).
Neither is “better”. They are answers to different questions.
An analogy#
A sports car carries two people at 250 km/h. A bus carries sixty people at 80 km/h.
To move one person across town, take the car. To move a stadium, the bus wins by a wide margin even though every individual passenger travels slower.
A CPU core is the car. A GPU is a fleet of buses.
A picture#
flowchart TB
subgraph CPU["CPU core: spend transistors on being clever"]
direction LR
BP["Branch<br/>predictor"] --> OOO["Out-of-order<br/>engine"] --> ALU1["A few<br/>arithmetic units"]
CACHE[("Big private caches")]
end
subgraph GPU["GPU block: spend transistors on arithmetic"]
direction LR
CTRL["One small<br/>control unit"] --> L1["lane"] & L2["lane"] & L3["lane"] & L4["... x128"]
end
class BP,OOO,CTRL neutral
class ALU1,L1,L2,L3,L4 compute
class CACHE memoryHow it really works#
Where a CPU core’s transistors go#
Most of a modern CPU core is not arithmetic. It is machinery for guessing:
- A branch predictor guesses which way every
ifwill go, so work can start early. - An out-of-order engine looks hundreds of instructions ahead and runs whichever are ready.
- Large caches keep recently used data a nanosecond away.
All of it exists to make unpredictable, branchy, pointer-chasing code — a web server, a compiler, a database — run fast on one thread.
Where a GPU’s transistors go#
A GPU deletes almost all of that. No sophisticated prediction, no reordering, small caches. The saved space is filled with arithmetic units and the wiring to feed them. One control unit steers many lanes at once, because all lanes run the same instruction (module II explains this as SIMT).
| Server CPU (64 cores) | Data-center GPU (H100) | |
|---|---|---|
| Independent control units | 64 | 132 |
| Arithmetic lanes | 64 × 16 (AVX-512) ≈ 1,000 | 132 × 128 ≈ 17,000 |
| Clock speed | ~3.5 GHz | ~1.7 GHz |
| Peak FP32 arithmetic | ~3 TFLOP/s | ~67 TFLOP/s |
| Peak with AI-specific units | ~10 TFLOP/s | ~1,000 TFLOP/s (FP16 tensor cores) |
| Memory bandwidth | ~0.3 TB/s | ~3.35 TB/s |
| Handles branchy code | Excellent | Poor |
Read the table as a trade, not a ranking. The GPU’s lanes are individually slower and cannot make independent decisions cheaply. It wins only when all the lanes have identical work.
Amdahl’s law: the part you cannot parallelize#
If a fraction p of a job can be spread over n workers and the rest cannot:
speedup = 1 / ((1 - p) + p / n)With p = 0.95 and unlimited workers, the best possible speedup is 1 / 0.05 = 20x. The
serial 5% becomes the whole runtime. This is why a GPU program is always a partnership: the
CPU does the serial part, and you fight to keep that part tiny.
Code#
// amdahl.go — why 17,000 lanes do not give a 17,000x speedup.
package main
import "fmt"
func speedup(parallelFraction float64, workers float64) float64 {
return 1 / ((1 - parallelFraction) + parallelFraction/workers)
}
func main() {
fmt.Println("parallel% 8 workers 64 workers 17000 workers")
for _, p := range []float64{0.50, 0.90, 0.99, 0.999, 0.9999} {
fmt.Printf("%7.2f%% %9.1fx %10.1fx %13.1fx\n",
p*100, speedup(p, 8), speedup(p, 64), speedup(p, 17000))
}
}
Notice the last column: going from 99% to 99.99% parallel changes the GPU speedup from about 100x to about 6,300x. On a GPU, the small serial remainder is everything.
Remember this#
- CPU = minimum time for one stream of work. GPU = maximum total work per second.
- A GPU gets its arithmetic density by removing the “clever” parts of a CPU core.
- A GPU needs uniform, independent, arithmetic-heavy work to win.
- The serial fraction of your job caps the speedup (Amdahl’s law).
Try it#
- Run
amdahl.go. What parallel fraction do you need before 17,000 lanes beat 64 workers by at least 10x? - Think of a program you wrote recently. Estimate its parallel fraction. Would a GPU help?
- From the table, compute bytes of memory bandwidth available per FLOP for the CPU and for the GPU. Which one is more “starved” for memory? Keep the answer — module IV is built on it.
Check yourself#
- What does a CPU spend most of its transistors on, and why?
- Why is a single GPU lane slower than a CPU core?
- A job is 90% parallel. What is the maximum speedup with infinite hardware?