Below the API

From Pixels to Tensors

Foundations Beginner 35 min Difficulty 2/5

Prerequisites 01, 02

The idea in one minute#

GPUs were invented to draw video-game frames. Drawing a frame means running the same small calculation for millions of pixels, sixty times a second. Neural networks turned out to need exactly the same thing: the same small calculation over millions of numbers. The hardware built for games was, almost by accident, the right hardware for AI.

Knowing the history explains why GPU vocabulary (shaders, textures, kernels) and GPU design (massive parallel arithmetic, huge memory bandwidth) look the way they do.

A picture#

flowchart LR
  A["1990s<br/>Fixed-function<br/>3D cards"] --> B["2001<br/>Programmable<br/>shaders"]
  B --> C["2006<br/>CUDA: general<br/>purpose GPU"]
  C --> D["2012<br/>AlexNet trained<br/>on two GPUs"]
  D --> E["2017<br/>Tensor cores<br/>and transformers"]
  E --> F["2022 onward<br/>LLMs: GPUs become<br/>AI factories"]
  class A,B neutral
  class C,D compute
  class E,F io

How it really works#

Step 1: graphics is a parallel problem#

To draw a 3D scene, a GPU runs a pipeline:

flowchart LR
  V["Vertices<br/>3D points"] --> VS["Vertex shader<br/>move each point"]
  VS --> R["Rasterizer<br/>which pixels does<br/>each triangle cover"]
  R --> FS["Fragment shader<br/>colour each pixel"]
  FS --> FB[("Framebuffer<br/>the image")]
  class V,FB memory
  class VS,FS compute
  class R neutral

A 4K frame has 8.3 million pixels. Each pixel’s colour depends on its own inputs and on read-only data (textures, lights) — never on a neighbouring pixel’s result. So all of them can be computed at once. That property has a name: the work is data-parallel.

Step 2: shaders made GPUs programmable#

Early cards had the maths hard-wired. Around 2001 they began accepting small programs — shaders — that ran once per vertex or per pixel. A shader is a tiny function applied to millions of independent elements. That is a kernel in everything but name.

Researchers noticed, and started disguising scientific calculations as “drawing” to borrow the speed.

Step 3: CUDA removed the disguise#

In 2006 NVIDIA released CUDA: a way to write those per-element functions in ordinary C and run them on arrays of numbers, no graphics involved. This is GPGPU — general-purpose GPU computing. Open alternatives followed (OpenCL, later Vulkan compute, Metal, WebGPU), but CUDA’s libraries and tooling made it the default for scientific and AI work, and it still is.

Step 4: neural networks are matrix multiplication#

A neural network layer computes output = activation(W · input). The expensive part is the matrix multiply: every output number is a sum of products, and no output depends on another. Data-parallel again.

In 2012 a network called AlexNet, trained on two gaming GPUs, won the ImageNet image recognition contest by a wide margin. Training that would have taken months on CPUs took days. The field switched to GPUs almost overnight.

Step 5: GPUs redesigned for AI#

Since then the design has bent toward AI:

  • Tensor cores (2017): circuits that do nothing but small matrix multiplies, at reduced precision. They multiplied AI speed again by 4–8x.
  • High-bandwidth memory: large language models are limited by how fast weights can be read, so memory speed became as important as arithmetic.
  • Fast GPU-to-GPU links: models outgrew one GPU.

A modern data-center “GPU” often has no display output at all. The name stuck; the job changed.

Vocabulary that survives from graphics#

Graphics wordMeans in general computing
ShaderKernel: a function run per element
TextureRead-only array sampled by the kernel
FramebufferOutput array
VRAM (video RAM)GPU memory

Remember this#

  • Graphics and neural networks share one property: huge numbers of independent, identical calculations.
  • CUDA (2006) made that parallel hardware usable for any computation.
  • AlexNet (2012) proved GPUs transform deep learning; hardware has specialized for AI ever since.
  • A data-center GPU today is an AI accelerator that kept its old name.

Try it#

  1. A 4K frame at 60 frames per second: how many pixel calculations per second? Compare that to a language model with 8 billion weights producing 50 tokens per second, touching each weight once per token.
  2. Write a Go function shade(x, y int) color that returns a colour from the pixel coordinates only (for example a gradient). Render a 1920×1080 image with one goroutine per row. Why is no locking needed?

Check yourself#

  1. What does “data-parallel” mean?
  2. Why did shaders make GPUs interesting to scientists?
  3. Which two hardware features were added to GPUs specifically because of AI?

↑↓ navigate ↵ open