The idea in one minute#
Go’s garbage collector finds heap objects that nothing can reach any more and frees them. It is a concurrent mark-and-sweep collector: it traces everything reachable from the roots (goroutine stacks and global variables) while your program keeps running, then reclaims whatever was not reached. It does not move objects, and it has no generations.
It stops the world only twice per cycle, each for well under a millisecond. The cost you
actually pay is CPU: roughly proportional to how much live data has pointers and to how
fast you allocate. Two knobs set the trade-off — GOGC (how much the heap may grow before
the next cycle) and GOMEMLIMIT (a soft ceiling on total memory).
An analogy#
Clearing a huge warehouse while it stays open. Staff start from the order desk and tag every item that some open order still refers to, following references from item to item. Meanwhile workers keep moving stock — so whenever a worker re-shelves a reference, they must tell the taggers (the write barrier), or a needed item might be missed. When tagging is done, everything untagged goes in the skip. Tagging costs staff time in proportion to the tagged stock; how often you run the clear-out depends on how fast new stock arrives.
A picture#
flowchart TB
subgraph CYCLE["One GC cycle"]
direction TB
ST1["Sweep termination<br/>STOP THE WORLD<br/>tens of microseconds"] --> MARK["Mark<br/>CONCURRENT<br/>write barrier on,<br/>~25% of CPUs"]
MARK --> ST2["Mark termination<br/>STOP THE WORLD<br/>tens of microseconds"]
ST2 --> SWEEP["Sweep<br/>CONCURRENT<br/>free unmarked slots,<br/>lazily, as memory is needed"]
end
ROOTS["Roots:<br/>goroutine stacks, globals"] --> MARK
SWEEP --> WAIT["Mutator runs.<br/>Heap grows until it reaches the goal"]
WAIT -->|"heap goal = live heap x (1 + GOGC/100)"| ST1
class ST1,ST2 warn
class MARK,SWEEP memory
class ROOTS neutral
class WAIT computeHow it really works#
Tri-colour marking#
Every object is, conceptually, one of three colours:
| Colour | Meaning |
|---|---|
| White | Not yet seen. At the end of marking, white means garbage |
| Grey | Seen, but its pointers have not been followed yet |
| Black | Seen, and all its pointers followed |
Start with the roots grey. Repeatedly take a grey object, follow each pointer in it (greying any white object found), and blacken it. When no grey objects remain, everything reachable is black. Objects with no pointers (noscan — II.04) go straight to black: there is nothing to follow, which is why pointer-free data is nearly free to the collector.
Running concurrently: the write barrier#
Because your goroutines keep running, one could move a pointer to a white object into a black object and erase the only other path to it — the collector would then free a live object. To prevent this, while marking is active the compiler-inserted write barrier intercepts every pointer write to the heap and greys the objects involved. Pointer writes are a little slower during marking; outside it the barrier is a single not-taken branch.
Newly allocated objects during marking are born black.
Who does the work#
- Dedicated background workers use about 25% of
GOMAXPROCSduring the mark phase. - If goroutines allocate faster than marking progresses, the allocating goroutine is drafted to do marking itself — a mark assist. Assists are how the collector keeps up, and they are how GC shows up as latency in a request: the request that allocates pays.
- Sweeping is done lazily, a span at a time, mostly when the allocator needs a span.
The pacer: when a cycle starts#
heap goal = live heap × (1 + GOGC/100) + (stacks + globals) × GOGC/100With the default GOGC=100, a collection is triggered so that it finishes by the time the
heap has doubled relative to what survived the last one. So:
| Setting | Effect |
|---|---|
GOGC=100 (default) | Heap up to ~2× live; baseline GC CPU |
GOGC=200 | Heap up to ~3× live; about half the GC cycles |
GOGC=50 | Heap up to ~1.5× live; about twice the GC cycles |
GOGC=off | No collections triggered by growth (only by the memory limit) |
GC CPU cost ≈ (allocation rate ÷ free headroom) × cost to mark the live heap. Halving
your allocation rate halves the number of cycles. Doubling GOGC does the same, for memory.
GOMEMLIMIT: a soft ceiling#
GOMEMLIMIT=4GiB (or debug.SetMemoryLimit) tells the runtime to keep total Go memory under
that figure: as the limit approaches, it collects more often regardless of GOGC and returns
memory to the OS more eagerly.
The pattern for containers: set GOMEMLIMIT to about 90% of the container’s memory limit
and leave GOGC at 100 or higher. You use the memory you are paying for, and the collector
gets aggressive only when it must — instead of the kernel’s OOM killer deciding.
It is a soft limit. If live data genuinely exceeds it, the collector would run continuously; the runtime caps GC at roughly half the CPU and lets memory exceed the limit rather than grind to a halt. A memory limit does not fix a leak.
Green Tea#
From Go 1.26 the default collector is Green Tea. The algorithm above is unchanged; what changed is the order of work. The classic implementation followed pointers one object at a time, hopping randomly across memory — most of marking time was cache misses. Green Tea queues whole spans of small objects and scans the marked objects in a span together, so memory is touched in runs and the work can use vector instructions. The Go team reports 10–40% less GC CPU in allocation-heavy programs. No code change is needed; it is one more reason to keep the toolchain current.
What Go’s collector does not do#
- It does not move objects, so it does not compact. Fragmentation is handled by size classes (lesson 03). In exchange, addresses are stable and C code can hold Go pointers during a call.
- It is not generational. Most collectors bet that young objects die young and collect them cheaply; Go instead leans on escape analysis to keep short-lived values off the heap entirely.
Reading a GC trace#
$ GODEBUG=gctrace=1 ./app
gc 14 @2.104s 3%: 0.031+4.2+0.022 ms clock, 0.25+1.1/8.0/0+0.18 ms cpu, 96->98->49 MB, 100 MB goal, 0 MB stacks, 0 MB globals, 8 P| Field | Meaning |
|---|---|
gc 14 @2.104s | 14th cycle, 2.1 s after start |
3% | Share of available CPU spent in GC since the program started |
0.031+4.2+0.022 ms clock | STW sweep termination + concurrent mark + STW mark termination |
96->98->49 MB | Heap at start → heap at end of mark → live heap after |
100 MB goal | The pacer’s target for this cycle |
8 P | Processors used |
The two pauses are the first and third numbers. The live heap is the last of the three sizes. If the percentage is above ~10%, lesson 05 is for you.
Finalizers, cleanups and weak pointers#
runtime.AddCleanup(ptr, fn, arg)(Go 1.24) runs a function after an object becomes unreachable — for releasing a non-memory resource. It replacesSetFinalizer, which had sharp edges. Neither is a destructor: timing is not guaranteed, so close files explicitly.weak.Pointer[T](Go 1.24) refers to an object without keeping it alive — the building block for caches that should not pin their entries.
How programs “leak” in a garbage-collected language#
The collector frees what is unreachable. A leak is something still reachable that you no longer want:
| Leak | Cause |
|---|---|
| A map or slice that only grows | A cache with no eviction; a registry nobody deletes from |
| Goroutines blocked forever | Each pins its stack and everything it references (IV.06) |
| A small slice or substring of a large buffer | Pins the whole buffer (II.01, II.02) |
| Tickers and timers never stopped | Held by the runtime’s timer heap |
defer in a long loop | Deferred calls accumulate until the function returns |
A heap profile shows what is live and who allocated it (V.01).
Code#
// gc.go — how GOGC and live-heap shape change GC frequency, CPU share and pauses.
package main
import (
"fmt"
"runtime"
"runtime/debug"
"time"
)
type Node struct {
Next *Node
Val [6]int64
}
var live any // keeps the long-lived data reachable
// churn allocates short-lived garbage for a fixed amount of work.
func churn() {
var keep [][]byte
for i := 0; i < 400000; i++ {
b := make([]byte, 256)
if i%64 == 0 {
keep = append(keep, b) // a little survives for a while
}
if len(keep) > 512 {
keep = keep[:0]
}
}
runtime.KeepAlive(keep)
}
func run(name string, gogc int) {
debug.SetGCPercent(gogc)
runtime.GC()
var a, b runtime.MemStats
runtime.ReadMemStats(&a)
start := time.Now()
churn()
elapsed := time.Since(start)
runtime.ReadMemStats(&b)
cycles := b.NumGC - a.NumGC
pause := time.Duration(b.PauseTotalNs - a.PauseTotalNs)
avg := time.Duration(0)
if cycles > 0 {
avg = pause / time.Duration(cycles)
}
fmt.Printf("%-30s GOGC=%-4d %3d cycles %6.1f ms avg pause %5v next GC at %4d MB\n",
name, gogc, cycles, float64(elapsed.Microseconds())/1000, avg.Round(time.Microsecond), b.NextGC/1e6)
}
func main() {
fmt.Println("allocating ~100 MB of short-lived garbage in each run")
live = nil
run("small live heap", 100)
run("small live heap", 400)
// 60 MB of live data FULL of pointers: every cycle must walk all of it.
var head *Node
for i := 0; i < 1_000_000; i++ {
head = &Node{Next: head}
}
live = head
run("60 MB live, pointer-rich", 100)
run("60 MB live, pointer-rich", 400)
// The same amount of live data with NO pointers: marked without being scanned.
live = make([]int64, 7_000_000)
head = nil
run("56 MB live, pointer-free", 100)
// A soft memory limit makes the collector work harder as the ceiling approaches.
debug.SetGCPercent(400)
debug.SetMemoryLimit(80 << 20)
run("56 MB live, GOMEMLIMIT=80MB", 400)
debug.SetMemoryLimit(1 << 62)
}
Read the output as four comparisons. Raising GOGC cuts the number of cycles. A larger live
heap means fewer cycles, because the goal is further away. Each of those cycles is far more
expensive when the live data is full of pointers than when it is pointer-free — compare the
elapsed times of the two ~60 MB runs at GOGC=100. And a memory limit forces extra cycles
that GOGC alone would not have run.
Remember this#
- Concurrent mark and sweep: trace from stacks and globals while the program runs; two sub-millisecond pauses per cycle.
- Cost is CPU, driven by allocation rate and by how much live data contains pointers.
GOGCtrades memory for CPU.GOMEMLIMITis a soft ceiling; set it to ~90% of the container limit.- Non-moving, non-generational. Since 1.26, Green Tea scans span by span for better locality.
- A “leak” is reachable memory you forgot about.
Try it#
- Run
gc.go, then again withGODEBUG=gctrace=1. Find the live heap and the two pause times in one trace line. - Change
Nodeso it holds an index (int32) into a slice instead of a pointer. What happens to the pointer-rich runs? - Write a program with a deliberately leaking map and watch the live heap in the trace grow after every cycle.
Check yourself#
- Why does the collector need a write barrier?
- With
GOGC=100and 500 MB live, at roughly what heap size does the next cycle finish? - Why does a
[]float32of a gigabyte cost the collector almost nothing?