Below the API

Writing Allocation-Aware Code

Intermediate Advanced 55 min Difficulty 3/5

Prerequisites 02, 03, 04

The idea in one minute#

Most Go performance work is allocation work. The method is always the same: measure which allocations are on the hot path, then remove them with a small set of techniques — preallocate, reuse buffers, keep values off the heap, avoid pointers in large data, and pool what cannot be avoided.

Two warnings before any technique. Do it only where a profile says it matters; clear code that allocates is better than clever code that does not, everywhere except the hot path. And reused memory is shared memory: every technique here trades safety margin for speed.

An analogy#

A café that hands every customer a new ceramic mug and throws it away afterwards would spend its day buying mugs and emptying bins. Washing and reusing mugs is cheaper — but now someone must make sure a mug is not handed to the next customer while the first is still drinking.

A picture#

flowchart TB
  PROF["Profile: where are the allocations?<br/>-benchmem, heap profile, gctrace"] --> HOT{"On a hot path?"}
  HOT -->|"no"| LEAVE["Leave it. Readability wins."]
  HOT -->|"yes"| T1["Can it stay on the stack?<br/>values, constant sizes, no interfaces"]
  T1 -->|"no"| T2["Can the size be known?<br/>preallocate with make(T, 0, n)"]
  T2 -->|"no"| T3["Can the caller own the buffer?<br/>AppendX(dst, ...) APIs"]
  T3 -->|"no"| T4["Can it be reused across calls?<br/>sync.Pool, per-worker scratch"]
  T4 --> T5["Is the live data pointer-heavy?<br/>flat slices and indexes"]
  T5 --> VERIFY["Benchmark again.<br/>Keep it only if it helped."]
  class PROF,VERIFY queue
  class HOT queue
  class LEAVE neutral
  class T1,T2,T3 compute
  class T4,T5 memory

How it really works#

1. Measure first#

go test -bench . -benchmem                     # B/op and allocs/op per benchmark
go test -bench X -memprofile mem.out           # then: go tool pprof -sample_index=alloc_space mem.out
GODEBUG=gctrace=1 ./app                        # how much CPU the collector is using
go build -gcflags=-m ./...                     # why something escapes

In pprof, alloc_space shows where bytes are allocated (the pressure on the GC) and inuse_space shows what is live (your footprint). They answer different questions.

2. Preallocate#

out := make([]Result, 0, len(inputs))      // one allocation instead of ~log n
m := make(map[string]int, expected)
var sb strings.Builder; sb.Grow(n)

3. Let the caller own the buffer#

func AppendTokens(dst []int, text string) []int      // appends and returns dst
buf = AppendTokens(buf[:0], text)                    // caller reuses one buffer for every call

buf[:0] keeps the capacity and resets the length: the idiom for reuse. The standard library follows the same shape: strconv.AppendInt, fmt.Appendf, time.Time.AppendFormat, binary.Append, hex.AppendEncode.

4. Keep values on the stack#

  • Return small structs by value.
  • Use fixed-size arrays for scratch space: var tmp [64]byte.
  • Avoid any, fmt and reflection in inner loops; strconv and slog’s typed attributes exist for this.
  • Do not take the address of a value unless you need to.

5. Avoid pointers in large live data#

From II.04 and lesson 04: the collector’s work is proportional to pointer-carrying live memory.

Instead ofUse
[]*Item[]Item
map[string]*Item with millions of entriesmap[string]int32 indexing a []Item; or sorted keys and binary search
A tree of nodes with child pointersA slice of nodes with int32 child indexes
Millions of small stringsOne []byte plus offsets
time.Time in huge slices (it contains a pointer)int64 Unix nanoseconds

An arena-style layout — one big slice from which you carve by index and which you drop all at once — makes allocation a pointer bump and deallocation free.

6. sync.Pool#

A pool holds temporary objects for reuse across goroutines:

var bufPool = sync.Pool{New: func() any { return new(bytes.Buffer) }}

b := bufPool.Get().(*bytes.Buffer)
b.Reset()
defer bufPool.Put(b)

How it behaves:

  • Each P has a private slot and a local list, so Get/Put are usually lock-free.
  • The collector empties pools: objects survive at most two GC cycles. A pool is a cache that smooths allocation between collections, not a guaranteed store.
  • Store pointers (*bytes.Buffer, *[]byte): putting a slice value in boxes it into an interface, which itself allocates.
  • Reset before use, and do not return a buffer to the pool while anything still references it — that is a use-after-free, the kind of bug Go otherwise spares you.
  • Cap what you return. One request that grew a buffer to 50 MB will otherwise keep 50 MB alive in the pool. Drop buffers above a threshold.

Use a pool when objects are expensive to allocate, short-lived and used at a high rate: request buffers, encoders, scratch tensors. For a fixed number of workers, per-worker scratch space — a struct of buffers owned by each goroutine — is simpler and faster than a pool.

7. Strings and bytes#

Avoid round trips between string and []byte (each one copies). Work in []byte through the pipeline and convert once at the edge. Intern repeated strings with unique.Make. The zero-copy conversions unsafe.String and unsafe.Slice exist; they are correct only if the bytes are never modified afterwards, and they belong behind a small, tested API.

8. Tuning last#

After the code is reasonable, GOGC and GOMEMLIMIT (lesson 04) trade spare memory for less GC CPU with no code at all. A heap ballast — a large unused allocation to inflate the GC goal — was a common trick before GOMEMLIMIT existed; it is no longer needed.

What it looks like in AI code#

Module VI’s inner loops are matrix multiplies and attention over []float32. The rules there: allocate tensors once per model or per request, never per token; give every worker its own scratch buffers; keep weights in flat pointer-free slices so that a multi-gigabyte model costs the collector nothing; and make the per-token decode step allocation-free, because it runs thousands of times per second.

Code#

The same task — tokenize lines and count token IDs — written naively, then with each technique.

// allocs.go — one task, four implementations, measured.
package main

import (
	"fmt"
	"strings"
	"sync"
	"testing"
)

var corpus = strings.Repeat("the quick brown fox jumps over the lazy dog\n", 200)

func hash(w []byte) int {
	h := 0
	for _, c := range w {
		h = h*31 + int(c)
	}
	return h & 1023
}

// 1. Naive: Split allocates every line and word; strings are converted; the slice grows.
func naive(text string) []int {
	var ids []int
	for _, line := range strings.Split(text, "\n") {
		for _, w := range strings.Fields(line) {
			ids = append(ids, hash([]byte(w)))
		}
	}
	return ids
}

// 2. Preallocated and scanning bytes in place: one allocation for the result.
func prealloc(text string) []int {
	ids := make([]int, 0, len(text)/4)
	return appendIDs(ids, text)
}

// appendIDs is the allocation-free core: the caller owns dst.
func appendIDs(dst []int, text string) []int {
	start := -1
	h := 0
	for i := 0; i < len(text); i++ {
		c := text[i]
		if c == ' ' || c == '\n' {
			if start >= 0 {
				dst = append(dst, h&1023)
				start, h = -1, 0
			}
			continue
		}
		if start < 0 {
			start = i
		}
		h = h*31 + int(c)
	}
	if start >= 0 {
		dst = append(dst, h&1023)
	}
	return dst
}

// 3. Caller-owned buffer, reused across calls.
type worker struct{ ids []int }

func (w *worker) run(text string) []int {
	w.ids = appendIDs(w.ids[:0], text)
	return w.ids
}

// 4. sync.Pool, for when callers are many goroutines.
var pool = sync.Pool{New: func() any { s := make([]int, 0, 4096); return &s }}

func pooled(text string) int {
	p := pool.Get().(*[]int)
	ids := appendIDs((*p)[:0], text)
	n := len(ids)
	*p = ids
	pool.Put(p)
	return n
}

func main() {
	w := &worker{}
	impls := []struct {
		name string
		f    func()
	}{
		{"1 naive", func() { _ = naive(corpus) }},
		{"2 preallocate, scan bytes", func() { _ = prealloc(corpus) }},
		{"3 reuse a worker buffer", func() { _ = w.run(corpus) }},
		{"4 sync.Pool", func() { _ = pooled(corpus) }},
	}
	fmt.Println("implementation                 ns/op     B/op  allocs/op")
	for _, im := range impls {
		r := testing.Benchmark(func(b *testing.B) {
			b.ReportAllocs()
			for i := 0; i < b.N; i++ {
				im.f()
			}
		})
		fmt.Printf("%-28s %8d %8d %10d\n", im.name, r.NsPerOp(), r.AllocedBytesPerOp(), r.AllocsPerOp())
	}
	fmt.Println("\nSame output, same algorithm. The difference is who allocates, and how often.")
}

Remember this#

  • Profile first; optimize allocations only on the hot path.
  • Preallocate, let the caller own buffers (AppendX(dst, …)), keep values on the stack.
  • Large live data should be flat and pointer-free.
  • sync.Pool smooths allocation between collections; reset, cap, and never use after Put.
  • Tune GOGC/GOMEMLIMIT after the code, not instead of it.

Try it#

  1. Run allocs.go. Which single change removed the most allocations? The most time?
  2. Introduce a bug: have worker.run return w.ids and call it twice, keeping both results. What happens to the first result? This is the price of reuse.
  3. Take a function of your own, run it under -benchmem, and remove one allocation using -gcflags=-m to find it.

Check yourself#

  1. What is the difference between alloc_space and inuse_space in a heap profile?
  2. Why should a sync.Pool hold pointers rather than slice values?
  3. Why does replacing []*Item with []Item help the garbage collector?

↑↓ navigate ↵ open