Below the API

Occupancy and Divergence

Intermediate Intermediate 55 min Difficulty 4/5

Prerequisites II.01, II.02

The idea in one minute#

An SM hides memory waits by switching between many resident warps. Occupancy is how many warps a kernel actually keeps resident, as a fraction of the SM’s maximum. Each thread needs registers and each block may need shared memory; a kernel that is greedy with either leaves room for fewer warps, and the SM runs out of ready work.

Divergence is the other way to waste an SM: warps whose threads disagree at branches execute both paths (II.01). This lesson shows how to compute the first and reduce the second.

A picture#

flowchart LR
  subgraph LOW["Low occupancy"]
    direction TB
    A1["warp: waiting on memory"]
    A2["warp: waiting on memory"]
    A3["SM idle: nothing ready"]
  end
  subgraph HIGH["High occupancy"]
    direction TB
    B1["warp: waiting"]
    B2["warp: waiting"]
    B3["warp: waiting"]
    B4["warp: ready, runs now"]
    B5["warp: ready next"]
  end
  class A1,A2,B1,B2,B3 queue
  class A3 warn
  class B4,B5 compute

How it really works#

What limits resident warps#

An SM has fixed pools (H100 figures):

ResourcePer SMConsumed by
Warp slots64Every warp
Registers65,536registers per thread × threads
Shared memoryup to ~228 KBshared memory per block × blocks
Block slots32Every block

The number of blocks that fit is the minimum allowed by each resource. Occupancy is:

occupancy = resident warps ÷ 64

Example: a kernel using 80 registers per thread, block size 256 (8 warps), no shared memory.

by registers: 65,536 ÷ (80 × 256) = 3 blocks  → 24 warps
by warp slots: 64 ÷ 8             = 8 blocks  → 64 warps
resident = 24 warps  → occupancy 37.5%

Registers are the limit. Trimming the kernel to 64 registers per thread gives 4 blocks, 32 warps, 50%.

Higher is not always better#

Occupancy exists to hide memory latency. A kernel that mostly waits on memory benefits from many warps. A kernel that is pure arithmetic on data already in registers is saturated by a few warps; raising occupancy does nothing. Past roughly 50%, gains are usually small.

Treat occupancy as a diagnostic: very low occupancy (under ~25%) in a memory-heavy kernel is a finding. Chasing 100% is not a goal.

Reducing divergence#

Three standard moves:

  1. Replace the branch with arithmetic. max(x, 0) instead of if x < 0 { x = 0 }. The compiler turns simple conditionals into a “select” instruction that every thread runs identically.
  2. Group similar work. Sort or partition elements so each warp’s threads take the same path (the trick from II.01’s simulator).
  3. Move the decision outward. If a condition is the same for a whole batch, choose between two kernels on the host instead of branching inside one.

Loops with data-dependent trip counts are divergence too: a warp runs until its slowest thread finishes, with the finished ones idle.

Why neural networks are friendly#

Matrix multiplies, convolutions and activations do the same thing for every element: no data-dependent branches, no variable-length loops. That uniformity is a large part of why AI workloads fit GPUs so well. Divergence shows up in the parts around the model — sampling, tokenization, sparse or ragged data — which is why those often stay on the CPU.

Code#

An occupancy calculator. Real toolchains print the same numbers; computing them yourself makes the limits obvious.

// occupancy.go — how many warps can an SM keep resident for a given kernel?
package main

import "fmt"

type SM struct{ WarpSlots, Registers, SharedBytes, BlockSlots int }

type Kernel struct {
	Name           string
	BlockThreads   int
	RegsPerThread  int
	SharedPerBlock int // bytes
}

func (sm SM) Occupancy(k Kernel) (pct float64, limiter string) {
	warpsPerBlock := (k.BlockThreads + 31) / 32
	blocks, limiter := sm.WarpSlots/warpsPerBlock, "warp slots"
	if b := sm.Registers / (k.RegsPerThread * k.BlockThreads); b < blocks {
		blocks, limiter = b, "registers"
	}
	if k.SharedPerBlock > 0 {
		if b := sm.SharedBytes / k.SharedPerBlock; b < blocks {
			blocks, limiter = b, "shared memory"
		}
	}
	if sm.BlockSlots < blocks {
		blocks, limiter = sm.BlockSlots, "block slots"
	}
	return 100 * float64(blocks*warpsPerBlock) / float64(sm.WarpSlots), limiter
}

func main() {
	h100 := SM{WarpSlots: 64, Registers: 65536, SharedBytes: 228 * 1024, BlockSlots: 32}
	for _, k := range []Kernel{
		{"light elementwise", 256, 16, 0},
		{"register-hungry", 256, 80, 0},
		{"trimmed to 64 regs", 256, 64, 0},
		{"big shared tile", 256, 32, 96 * 1024},
		{"tiny blocks", 32, 16, 0},
	} {
		pct, lim := h100.Occupancy(k)
		fmt.Printf("%-20s occupancy %5.1f%%  limited by %s\n", k.Name, pct, lim)
	}
}

Remember this#

  • Occupancy = resident warps ÷ maximum. Limited by registers, shared memory, or slot counts.
  • It matters for kernels that wait on memory; it is irrelevant for pure-arithmetic kernels.
  • Very low occupancy is a red flag; maximum occupancy is not a target.
  • Cut divergence by using arithmetic instead of branches, grouping similar work, and deciding on the host.

Try it#

  1. Run occupancy.go. Why is “tiny blocks” limited to 50%? What block size fixes it?
  2. For the “big shared tile” kernel, what is the largest tile that still gives 50% occupancy?
  3. Rewrite if x < 0 { y = 0 } else { y = x } and if a > b { m = a } else { m = b } without branches.

Check yourself#

  1. Name three resources that can limit occupancy.
  2. When does raising occupancy not help?
  3. Give two ways to reduce divergence.

↑↓ navigate ↵ open