PidokuInfra

What Is a GPU?

Foundations Beginner 30 min Difficulty 1/5 Topic 01 of 04

The idea in one minute#

A GPU (graphics processing unit) is a processor built to do the same simple calculation on a very large amount of data at once. A CPU has a few powerful cores that each finish one task quickly. A GPU has thousands of small arithmetic units that work in step with each other, plus its own very fast memory.

It cannot run a program by itself. A CPU hands it data and a small function, the GPU applies that function to all the data in parallel, and hands the result back.

An analogy#

A CPU is a small team of master chefs. Each can cook any dish, improvise, and change plans halfway. Ask for one complicated meal and it arrives quickly.

A GPU is a factory line with ten thousand workers who each know one step. It is useless for a single custom meal — but ask for a million identical sandwiches and nothing else comes close.

Most of computing is custom meals. Graphics and AI happen to be sandwiches.

A picture#

flowchart LR
  subgraph HOST["Host: the computer"]
    CPU["CPU<br/>runs your Go program"]
    RAM[("System RAM")]
    CPU --- RAM
  end
  subgraph CARD["GPU: the device"]
    CORES["Thousands of<br/>arithmetic lanes"]
    VRAM[("GPU memory<br/>VRAM")]
    CORES --- VRAM
  end
  CPU -->|"1. copy data in"| VRAM
  CPU -->|"2. start a kernel"| CORES
  VRAM -->|"3. copy results out"| RAM
  class CPU,CORES compute
  class RAM,VRAM memory

Three words you will use for the rest of the course:

  • Host — the CPU and system RAM. Your Go program lives here.
  • Device — the GPU and its own memory.
  • Kernel — the small function the GPU runs on every piece of data. (Nothing to do with an operating-system kernel.)

How it really works#

A GPU card contains four things:

PartWhat it doesExample (NVIDIA H100)
The GPU chipThousands of arithmetic units grouped into blocks132 blocks (“SMs”), 80 billion transistors
GPU memory (VRAM)Holds the data being worked on; separate from system RAM80 GB
The link to the hostCarries data and commands to and from the CPUPCIe 5.0, about 64 GB/s
Power and coolingA GPU turns a lot of electricity into heatup to 700 W

Two facts shape everything else:

  1. GPU memory is separate. Data in system RAM is invisible to the GPU until you copy it across the link, and the link is 50 times slower than the GPU’s own memory. Good GPU programs copy data in once and keep it there.
  2. The GPU only speeds up work that is the same for every element. “Multiply these two billion numbers pairwise” is perfect. “Parse this JSON” is hopeless.

Code#

You do not need a GPU to feel the idea. This program does one job — brighten every pixel of an image — first as a single worker, then as many workers doing identical steps side by side.

Go
// brighten.go — the shape of GPU work: one tiny function, applied to everything, in parallel.
package main

import (
	"fmt"
	"runtime"
	"sync"
	"time"
)

// the "kernel": what happens to ONE element
func brighten(p uint8) uint8 {
	return uint8(min(int(p)+40, 255))
}

func main() {
	pixels := make([]uint8, 3840*2160*3*20) // twenty 4K frames, 3 bytes per pixel
	out := make([]uint8, len(pixels))

	// CPU style: one worker walks the whole array
	t0 := time.Now()
	for i, p := range pixels {
		out[i] = brighten(p)
	}
	serial := time.Since(t0)

	// GPU style: many workers, each takes a slice, all run the same kernel
	workers := runtime.NumCPU()
	chunk := (len(pixels) + workers - 1) / workers
	t0 = time.Now()
	var wg sync.WaitGroup
	for w := 0; w < workers; w++ {
		lo, hi := w*chunk, min((w+1)*chunk, len(pixels))
		wg.Add(1)
		go func() {
			defer wg.Done()
			for i := lo; i < hi; i++ {
				out[i] = brighten(pixels[i])
			}
		}()
	}
	wg.Wait()
	parallel := time.Since(t0)

	fmt.Printf("1 worker:   %v\n%d workers: %v  (%.1fx)\n",
		serial, workers, parallel, float64(serial)/float64(parallel))
}

Your laptop has perhaps 8–16 workers. A GPU runs the same pattern with tens of thousands of lanes. The program’s structure — a tiny function, a huge array, no worker needs another worker’s result — is exactly what a GPU kernel looks like.

Remember this#

  • A GPU applies one small function (a kernel) to huge amounts of data in parallel.
  • It has its own memory. Moving data between host and device is slow; do it rarely.
  • It is a helper to a CPU, never a replacement.
  • It only wins on work that is uniform across elements.

Try it#

  1. Run brighten.go. Is the speedup equal to your core count? If not, guess why. (You will find out in module II: memory, not arithmetic, is the limit.)
  2. Change the kernel so each pixel depends on the previous output pixel. Can you still split the work across workers? What does that tell you about which problems suit a GPU?
  3. If you have an NVIDIA GPU, run nvidia-smi and find: the model, total memory, and power draw.

Check yourself#

  1. What are “host”, “device” and “kernel”?
  2. Why can a GPU not replace the CPU?
  3. Name one workload that suits a GPU and one that does not, and say why.

↑↓ navigate↵ openesc close