Below the API

The Garbage Collector

Intermediate Advanced 1h 10m Difficulty 4/5

Prerequisites 01, 02, 03

The idea in one minute#

Go’s garbage collector finds heap objects that nothing can reach any more and frees them. It is a concurrent mark-and-sweep collector: it traces everything reachable from the roots (goroutine stacks and global variables) while your program keeps running, then reclaims whatever was not reached. It does not move objects, and it has no generations.

It stops the world only twice per cycle, each for well under a millisecond. The cost you actually pay is CPU: roughly proportional to how much live data has pointers and to how fast you allocate. Two knobs set the trade-off — GOGC (how much the heap may grow before the next cycle) and GOMEMLIMIT (a soft ceiling on total memory).

An analogy#

Clearing a huge warehouse while it stays open. Staff start from the order desk and tag every item that some open order still refers to, following references from item to item. Meanwhile workers keep moving stock — so whenever a worker re-shelves a reference, they must tell the taggers (the write barrier), or a needed item might be missed. When tagging is done, everything untagged goes in the skip. Tagging costs staff time in proportion to the tagged stock; how often you run the clear-out depends on how fast new stock arrives.

A picture#

flowchart TB
  subgraph CYCLE["One GC cycle"]
    direction TB
    ST1["Sweep termination<br/>STOP THE WORLD<br/>tens of microseconds"] --> MARK["Mark<br/>CONCURRENT<br/>write barrier on,<br/>~25% of CPUs"]
    MARK --> ST2["Mark termination<br/>STOP THE WORLD<br/>tens of microseconds"]
    ST2 --> SWEEP["Sweep<br/>CONCURRENT<br/>free unmarked slots,<br/>lazily, as memory is needed"]
  end
  ROOTS["Roots:<br/>goroutine stacks, globals"] --> MARK
  SWEEP --> WAIT["Mutator runs.<br/>Heap grows until it reaches the goal"]
  WAIT -->|"heap goal = live heap x (1 + GOGC/100)"| ST1
  class ST1,ST2 warn
  class MARK,SWEEP memory
  class ROOTS neutral
  class WAIT compute

How it really works#

Tri-colour marking#

Every object is, conceptually, one of three colours:

ColourMeaning
WhiteNot yet seen. At the end of marking, white means garbage
GreySeen, but its pointers have not been followed yet
BlackSeen, and all its pointers followed

Start with the roots grey. Repeatedly take a grey object, follow each pointer in it (greying any white object found), and blacken it. When no grey objects remain, everything reachable is black. Objects with no pointers (noscan — II.04) go straight to black: there is nothing to follow, which is why pointer-free data is nearly free to the collector.

Running concurrently: the write barrier#

Because your goroutines keep running, one could move a pointer to a white object into a black object and erase the only other path to it — the collector would then free a live object. To prevent this, while marking is active the compiler-inserted write barrier intercepts every pointer write to the heap and greys the objects involved. Pointer writes are a little slower during marking; outside it the barrier is a single not-taken branch.

Newly allocated objects during marking are born black.

Who does the work#

  • Dedicated background workers use about 25% of GOMAXPROCS during the mark phase.
  • If goroutines allocate faster than marking progresses, the allocating goroutine is drafted to do marking itself — a mark assist. Assists are how the collector keeps up, and they are how GC shows up as latency in a request: the request that allocates pays.
  • Sweeping is done lazily, a span at a time, mostly when the allocator needs a span.

The pacer: when a cycle starts#

heap goal = live heap × (1 + GOGC/100)  +  (stacks + globals) × GOGC/100

With the default GOGC=100, a collection is triggered so that it finishes by the time the heap has doubled relative to what survived the last one. So:

SettingEffect
GOGC=100 (default)Heap up to ~2× live; baseline GC CPU
GOGC=200Heap up to ~3× live; about half the GC cycles
GOGC=50Heap up to ~1.5× live; about twice the GC cycles
GOGC=offNo collections triggered by growth (only by the memory limit)

GC CPU cost ≈ (allocation rate ÷ free headroom) × cost to mark the live heap. Halving your allocation rate halves the number of cycles. Doubling GOGC does the same, for memory.

GOMEMLIMIT: a soft ceiling#

GOMEMLIMIT=4GiB (or debug.SetMemoryLimit) tells the runtime to keep total Go memory under that figure: as the limit approaches, it collects more often regardless of GOGC and returns memory to the OS more eagerly.

The pattern for containers: set GOMEMLIMIT to about 90% of the container’s memory limit and leave GOGC at 100 or higher. You use the memory you are paying for, and the collector gets aggressive only when it must — instead of the kernel’s OOM killer deciding.

It is a soft limit. If live data genuinely exceeds it, the collector would run continuously; the runtime caps GC at roughly half the CPU and lets memory exceed the limit rather than grind to a halt. A memory limit does not fix a leak.

Green Tea#

From Go 1.26 the default collector is Green Tea. The algorithm above is unchanged; what changed is the order of work. The classic implementation followed pointers one object at a time, hopping randomly across memory — most of marking time was cache misses. Green Tea queues whole spans of small objects and scans the marked objects in a span together, so memory is touched in runs and the work can use vector instructions. The Go team reports 10–40% less GC CPU in allocation-heavy programs. No code change is needed; it is one more reason to keep the toolchain current.

What Go’s collector does not do#

  • It does not move objects, so it does not compact. Fragmentation is handled by size classes (lesson 03). In exchange, addresses are stable and C code can hold Go pointers during a call.
  • It is not generational. Most collectors bet that young objects die young and collect them cheaply; Go instead leans on escape analysis to keep short-lived values off the heap entirely.

Reading a GC trace#

$ GODEBUG=gctrace=1 ./app
gc 14 @2.104s 3%: 0.031+4.2+0.022 ms clock, 0.25+1.1/8.0/0+0.18 ms cpu, 96->98->49 MB, 100 MB goal, 0 MB stacks, 0 MB globals, 8 P
FieldMeaning
gc 14 @2.104s14th cycle, 2.1 s after start
3%Share of available CPU spent in GC since the program started
0.031+4.2+0.022 ms clockSTW sweep termination + concurrent mark + STW mark termination
96->98->49 MBHeap at start → heap at end of mark → live heap after
100 MB goalThe pacer’s target for this cycle
8 PProcessors used

The two pauses are the first and third numbers. The live heap is the last of the three sizes. If the percentage is above ~10%, lesson 05 is for you.

Finalizers, cleanups and weak pointers#

  • runtime.AddCleanup(ptr, fn, arg) (Go 1.24) runs a function after an object becomes unreachable — for releasing a non-memory resource. It replaces SetFinalizer, which had sharp edges. Neither is a destructor: timing is not guaranteed, so close files explicitly.
  • weak.Pointer[T] (Go 1.24) refers to an object without keeping it alive — the building block for caches that should not pin their entries.

How programs “leak” in a garbage-collected language#

The collector frees what is unreachable. A leak is something still reachable that you no longer want:

LeakCause
A map or slice that only growsA cache with no eviction; a registry nobody deletes from
Goroutines blocked foreverEach pins its stack and everything it references (IV.06)
A small slice or substring of a large bufferPins the whole buffer (II.01, II.02)
Tickers and timers never stoppedHeld by the runtime’s timer heap
defer in a long loopDeferred calls accumulate until the function returns

A heap profile shows what is live and who allocated it (V.01).

Code#

// gc.go — how GOGC and live-heap shape change GC frequency, CPU share and pauses.
package main

import (
	"fmt"
	"runtime"
	"runtime/debug"
	"time"
)

type Node struct {
	Next *Node
	Val  [6]int64
}

var live any // keeps the long-lived data reachable

// churn allocates short-lived garbage for a fixed amount of work.
func churn() {
	var keep [][]byte
	for i := 0; i < 400000; i++ {
		b := make([]byte, 256)
		if i%64 == 0 {
			keep = append(keep, b) // a little survives for a while
		}
		if len(keep) > 512 {
			keep = keep[:0]
		}
	}
	runtime.KeepAlive(keep)
}

func run(name string, gogc int) {
	debug.SetGCPercent(gogc)
	runtime.GC()
	var a, b runtime.MemStats
	runtime.ReadMemStats(&a)
	start := time.Now()
	churn()
	elapsed := time.Since(start)
	runtime.ReadMemStats(&b)

	cycles := b.NumGC - a.NumGC
	pause := time.Duration(b.PauseTotalNs - a.PauseTotalNs)
	avg := time.Duration(0)
	if cycles > 0 {
		avg = pause / time.Duration(cycles)
	}
	fmt.Printf("%-30s GOGC=%-4d %3d cycles  %6.1f ms  avg pause %5v  next GC at %4d MB\n",
		name, gogc, cycles, float64(elapsed.Microseconds())/1000, avg.Round(time.Microsecond), b.NextGC/1e6)
}

func main() {
	fmt.Println("allocating ~100 MB of short-lived garbage in each run")

	live = nil
	run("small live heap", 100)
	run("small live heap", 400)

	// 60 MB of live data FULL of pointers: every cycle must walk all of it.
	var head *Node
	for i := 0; i < 1_000_000; i++ {
		head = &Node{Next: head}
	}
	live = head
	run("60 MB live, pointer-rich", 100)
	run("60 MB live, pointer-rich", 400)

	// The same amount of live data with NO pointers: marked without being scanned.
	live = make([]int64, 7_000_000)
	head = nil
	run("56 MB live, pointer-free", 100)

	// A soft memory limit makes the collector work harder as the ceiling approaches.
	debug.SetGCPercent(400)
	debug.SetMemoryLimit(80 << 20)
	run("56 MB live, GOMEMLIMIT=80MB", 400)
	debug.SetMemoryLimit(1 << 62)
}

Read the output as four comparisons. Raising GOGC cuts the number of cycles. A larger live heap means fewer cycles, because the goal is further away. Each of those cycles is far more expensive when the live data is full of pointers than when it is pointer-free — compare the elapsed times of the two ~60 MB runs at GOGC=100. And a memory limit forces extra cycles that GOGC alone would not have run.

Remember this#

  • Concurrent mark and sweep: trace from stacks and globals while the program runs; two sub-millisecond pauses per cycle.
  • Cost is CPU, driven by allocation rate and by how much live data contains pointers.
  • GOGC trades memory for CPU. GOMEMLIMIT is a soft ceiling; set it to ~90% of the container limit.
  • Non-moving, non-generational. Since 1.26, Green Tea scans span by span for better locality.
  • A “leak” is reachable memory you forgot about.

Try it#

  1. Run gc.go, then again with GODEBUG=gctrace=1. Find the live heap and the two pause times in one trace line.
  2. Change Node so it holds an index (int32) into a slice instead of a pointer. What happens to the pointer-rich runs?
  3. Write a program with a deliberately leaking map and watch the live heap in the trace grow after every cycle.

Check yourself#

  1. Why does the collector need a write barrier?
  2. With GOGC=100 and 500 MB live, at roughly what heap size does the next cycle finish?
  3. Why does a []float32 of a gigabyte cost the collector almost nothing?

↑↓ navigate ↵ open