Below the API

Convolutions

Basic Intermediate 45 min Difficulty 2/5

Prerequisites 03

LLMs don’t use convolutions. This file exists because (a) vision encoders in multimodal models do, (b) speech models do, (c) convolution is where many inference optimization techniques were invented, and (d) you will be asked about it.


1. What is it?#

A convolution slides a small learned filter over an input, computing a dot product at each position.

Input (5×5)          Kernel (3×3)      Output (3×3)
1 2 3 0 1              1 0 1
0 1 2 3 0              0 1 0     →     (each output = sum of 9 products)
1 0 1 2 3              1 0 1
2 1 0 1 2
0 2 1 0 1

Key properties: weight sharing (the same filter everywhere) and locality (each output depends only on a small neighborhood).


2. Why does it exist?#

Because for images, the same feature (an edge, a texture) is meaningful anywhere in the frame. A fully connected layer would learn a separate detector for every position — enormously wasteful. Convolution shares one detector across all positions.

For inference: convolutions have very high arithmetic intensity (each input pixel is used by many output positions), which makes them compute-bound and GPU-friendly — the opposite of decode.


3. Simple analogy#

A rubber stamp. You have one stamp (the filter) and you press it at every position on the page. The stamp doesn’t change; only where you press it does. That’s weight sharing.


4. Tiny example#

package main

import "fmt"

func main() {
	x := [3][3]float64{{1, 2, 3}, {4, 5, 6}, {7, 8, 9}}
	k := [2][2]float64{{1, 0}, {0, -1}}

	var out [2][2]float64
	for i := 0; i < 2; i++ {
		for j := 0; j < 2; j++ {
			for di := 0; di < 2; di++ { // slide the 2x2 kernel over the input
				for dj := 0; dj < 2; dj++ {
					out[i][j] += x[i+di][j+dj] * k[di][dj]
				}
			}
		}
	}
	fmt.Println(out)
	// [[1*1 + 5*(-1), 2*1 + 6*(-1)],      = [[-4 -4]
	//  [4*1 + 8*(-1), 5*1 + 9*(-1)]]         [-4 -4]]
}

FLOPs: 2 × H_out × W_out × C_in × C_out × K_h × K_w.

For a ResNet-50 first layer (224×224 input, 7×7 kernel, 3→64 channels, stride 2):

2 × 112 × 112 × 3 × 64 × 7 × 7 = 236 MFLOPs      for one layer, one image

5. Technical explanation#

How convolution becomes GEMM#

The dominant implementation strategy is im2col: rearrange overlapping patches into a matrix, then call GEMM.

Input (C_in, H, W) → im2col → (C_in·K_h·K_w, H_out·W_out)
Weights (C_out, C_in, K_h, K_w) → reshape → (C_out, C_in·K_h·K_w)

Convolution = Weights_matrix @ im2col_matrix     ← just a GEMM

Cost: the im2col matrix duplicates data by a factor of K_h·K_w (9x for a 3×3 kernel). Memory- hungry but lets you reuse the world’s best GEMM kernels.

Alternatives:

  • Implicit GEMM — compute the im2col indices on the fly inside the kernel. No extra memory. This is what cuDNN and CUTLASS actually do.
  • FFT convolution — for large kernels, O(n log n) instead of O(n·k). Rarely used for the 3×3 kernels that dominate.
  • Winograd — reduces multiplications for small kernels (3×3) by ~2.25x at the cost of more additions and reduced numerical precision. Widely used.
  • Direct — specialized kernels for specific shapes; often best for depthwise.

Arithmetic intensity#

3×3 conv, C_in=C_out=256, H=W=56:
  FLOPs = 2 × 56×56 × 256×256 × 9 = 3.7 GFLOP
  Bytes = input 56×56×256×2 + weights 256×256×9×2 + output 56×56×256×2
        = 1.6 MB + 1.2 MB + 1.6 MB = 4.4 MB
  Intensity = 3.7e9 / 4.4e6 = 841 FLOP/byte     ← very compute-bound

Compare to decode’s intensity of 1. Convolutions are the friendliest workload GPUs have. This is why image models hit 80%+ of peak FLOPs while LLM decode hits 0.3%.

Convolution variants worth knowing#

Standard         C_in × C_out × K × K params
Depthwise        C_in × K × K   (one filter per channel; very cheap, memory-bound)
Pointwise (1×1)  C_in × C_out   (exactly a GEMM over channels)
Depthwise-sep    depthwise + pointwise (MobileNet); 8-9x fewer FLOPs
Grouped          channels split into G groups (ResNeXt)
Dilated          gaps in the kernel; larger receptive field, same params
Transposed       "upsampling" convolution (decoders, diffusion)

Depthwise convolutions are memory-bound despite being convolutions — very few FLOPs per byte. They are a known optimization headache: mobile architectures reduce FLOPs but the wall-clock improvement is much smaller than the FLOP reduction suggests. A useful lesson that transfers directly to LLM work.

Where you’ll meet convolutions in an LLM stack#

Vision encoders in VLMs   ViT patch embedding is a strided conv; some use CNN backbones
Whisper / audio models    conv frontend before the transformer
Diffusion models          U-Net is convolution-heavy
Speculative decoding      some draft models use conv layers

For a multimodal serving stack, the vision encoder is often a separate, compute-bound stage with completely different batching characteristics from the LLM — a real scheduling complication (Section VIII.09).


6. Under the hood#

cuDNN picks among dozens of algorithms per convolution shape:

torch.backends.cudnn.benchmark = True   # autotune: try algorithms, cache the best

This helps a lot with fixed shapes and hurts with varying shapes (re-benchmarks on every new shape). For VLM serving with variable image sizes, either bucket the sizes or turn it off.


7. Performance implications#

  • Convolutions are compute-bound — the opposite of LLM decode. Optimize FLOPs, use tensor cores, use lower precision for the compute benefit.
  • Depthwise convs are memory-bound. FLOP reductions don’t translate to speedups.
  • Vision encoders in VLMs are a fixed cost per image, independent of the text length. At small text lengths they can dominate.
  • Channels-last (NHWC) layout is required for tensor cores on convolutions. NCHW forces a transpose or a slower kernel. (model.to(memory_format=torch.channels_last).)

8. Production implications#

  • In a VLM, profile the vision encoder separately. It may be 30-60% of prefill time for image-heavy requests.
  • Batch images across requests. The encoder is compute-bound; batching helps.
  • Cache image embeddings. If the same image appears in multiple turns of a conversation, encode once. This is often a large and easy win.
  • Use channels-last for any conv model.

9. Common mistakes#

Using NCHW with tensor cores. Silently falls back to a slower path.

Enabling cudnn.benchmark with dynamic shapes. Re-benchmarks constantly.

Assuming FLOP reduction = speedup for depthwise convolutions. It doesn’t.

Forgetting the vision encoder in capacity planning for VLMs.


10. Hands-on exercise#

A. Implement im2col. Write a convolution using im2col + GEMM in NumPy. Verify against a reference. Measure the memory blowup for a 3×3 kernel.

B. Intensity comparison. Compute arithmetic intensity for: a 3×3 conv, a depthwise 3×3 conv, a 1×1 conv, and a decode GEMV. Rank them. Explain the ranking.

C. Layout. Time a ResNet block in NCHW and channels-last on GPU in FP16. Report the ratio.

D. VLM profile. If you have access to a vision-language model, profile a request with one image. What fraction of prefill is the vision encoder?


11. Interview questions#

  1. How is a convolution turned into a GEMM, and what does that cost?
  2. Why are convolutions compute-bound while LLM decode is memory-bound?
  3. Why are depthwise convolutions slower than their FLOP count suggests?
  4. What is Winograd convolution and what does it trade?
  5. Where do convolutions appear in an LLM serving stack?
  6. Why does channels-last matter?

12. Further reading#

  • [REFERENCE] NVIDIA cuDNN developer guide; “Convolution Algorithms” section
  • [ESTABLISHED] Lavin & Gray, “Fast Algorithms for Convolutional Neural Networks” (Winograd)
  • [REFERENCE] PyTorch channels-last memory format tutorial
  • Next: 05 — Normalization layers

↑↓ navigate ↵ open