The idea in one minute#
A GPU (graphics processing unit) is a processor built to do the same simple calculation on a very large amount of data at once. A CPU has a few powerful cores that each finish one task quickly. A GPU has thousands of small arithmetic units that work in step with each other, plus its own very fast memory.
It cannot run a program by itself. A CPU hands it data and a small function, the GPU applies that function to all the data in parallel, and hands the result back.
An analogy#
A CPU is a small team of master chefs. Each can cook any dish, improvise, and change plans halfway. Ask for one complicated meal and it arrives quickly.
A GPU is a factory line with ten thousand workers who each know one step. It is useless for a single custom meal — but ask for a million identical sandwiches and nothing else comes close.
Most of computing is custom meals. Graphics and AI happen to be sandwiches.
A picture#
flowchart LR
subgraph HOST["Host: the computer"]
CPU["CPU<br/>runs your Go program"]
RAM[("System RAM")]
CPU --- RAM
end
subgraph CARD["GPU: the device"]
CORES["Thousands of<br/>arithmetic lanes"]
VRAM[("GPU memory<br/>VRAM")]
CORES --- VRAM
end
CPU -->|"1. copy data in"| VRAM
CPU -->|"2. start a kernel"| CORES
VRAM -->|"3. copy results out"| RAM
class CPU,CORES compute
class RAM,VRAM memoryThree words you will use for the rest of the course:
- Host — the CPU and system RAM. Your Go program lives here.
- Device — the GPU and its own memory.
- Kernel — the small function the GPU runs on every piece of data. (Nothing to do with an operating-system kernel.)
How it really works#
A GPU card contains four things:
| Part | What it does | Example (NVIDIA H100) |
|---|---|---|
| The GPU chip | Thousands of arithmetic units grouped into blocks | 132 blocks (“SMs”), 80 billion transistors |
| GPU memory (VRAM) | Holds the data being worked on; separate from system RAM | 80 GB |
| The link to the host | Carries data and commands to and from the CPU | PCIe 5.0, about 64 GB/s |
| Power and cooling | A GPU turns a lot of electricity into heat | up to 700 W |
Two facts shape everything else:
- GPU memory is separate. Data in system RAM is invisible to the GPU until you copy it across the link, and the link is 50 times slower than the GPU’s own memory. Good GPU programs copy data in once and keep it there.
- The GPU only speeds up work that is the same for every element. “Multiply these two billion numbers pairwise” is perfect. “Parse this JSON” is hopeless.
Code#
You do not need a GPU to feel the idea. This program does one job — brighten every pixel of an image — first as a single worker, then as many workers doing identical steps side by side.
// brighten.go — the shape of GPU work: one tiny function, applied to everything, in parallel.
package main
import (
"fmt"
"runtime"
"sync"
"time"
)
// the "kernel": what happens to ONE element
func brighten(p uint8) uint8 {
return uint8(min(int(p)+40, 255))
}
func main() {
pixels := make([]uint8, 3840*2160*3*20) // twenty 4K frames, 3 bytes per pixel
out := make([]uint8, len(pixels))
// CPU style: one worker walks the whole array
t0 := time.Now()
for i, p := range pixels {
out[i] = brighten(p)
}
serial := time.Since(t0)
// GPU style: many workers, each takes a slice, all run the same kernel
workers := runtime.NumCPU()
chunk := (len(pixels) + workers - 1) / workers
t0 = time.Now()
var wg sync.WaitGroup
for w := 0; w < workers; w++ {
lo, hi := w*chunk, min((w+1)*chunk, len(pixels))
wg.Add(1)
go func() {
defer wg.Done()
for i := lo; i < hi; i++ {
out[i] = brighten(pixels[i])
}
}()
}
wg.Wait()
parallel := time.Since(t0)
fmt.Printf("1 worker: %v\n%d workers: %v (%.1fx)\n",
serial, workers, parallel, float64(serial)/float64(parallel))
}Your laptop has perhaps 8–16 workers. A GPU runs the same pattern with tens of thousands of lanes. The program’s structure — a tiny function, a huge array, no worker needs another worker’s result — is exactly what a GPU kernel looks like.
Remember this#
- A GPU applies one small function (a kernel) to huge amounts of data in parallel.
- It has its own memory. Moving data between host and device is slow; do it rarely.
- It is a helper to a CPU, never a replacement.
- It only wins on work that is uniform across elements.
Try it#
- Run
brighten.go. Is the speedup equal to your core count? If not, guess why. (You will find out in module II: memory, not arithmetic, is the limit.) - Change the kernel so each pixel depends on the previous output pixel. Can you still split the work across workers? What does that tell you about which problems suit a GPU?
- If you have an NVIDIA GPU, run
nvidia-smiand find: the model, total memory, and power draw.
Check yourself#
- What are “host”, “device” and “kernel”?
- Why can a GPU not replace the CPU?
- Name one workload that suits a GPU and one that does not, and say why.