The idea in one minute#
Most Go performance work is allocation work. The method is always the same: measure which allocations are on the hot path, then remove them with a small set of techniques — preallocate, reuse buffers, keep values off the heap, avoid pointers in large data, and pool what cannot be avoided.
Two warnings before any technique. Do it only where a profile says it matters; clear code that allocates is better than clever code that does not, everywhere except the hot path. And reused memory is shared memory: every technique here trades safety margin for speed.
An analogy#
A café that hands every customer a new ceramic mug and throws it away afterwards would spend its day buying mugs and emptying bins. Washing and reusing mugs is cheaper — but now someone must make sure a mug is not handed to the next customer while the first is still drinking.
A picture#
flowchart TB
PROF["Profile: where are the allocations?<br/>-benchmem, heap profile, gctrace"] --> HOT{"On a hot path?"}
HOT -->|"no"| LEAVE["Leave it. Readability wins."]
HOT -->|"yes"| T1["Can it stay on the stack?<br/>values, constant sizes, no interfaces"]
T1 -->|"no"| T2["Can the size be known?<br/>preallocate with make(T, 0, n)"]
T2 -->|"no"| T3["Can the caller own the buffer?<br/>AppendX(dst, ...) APIs"]
T3 -->|"no"| T4["Can it be reused across calls?<br/>sync.Pool, per-worker scratch"]
T4 --> T5["Is the live data pointer-heavy?<br/>flat slices and indexes"]
T5 --> VERIFY["Benchmark again.<br/>Keep it only if it helped."]
class PROF,VERIFY queue
class HOT queue
class LEAVE neutral
class T1,T2,T3 compute
class T4,T5 memoryHow it really works#
1. Measure first#
go test -bench . -benchmem # B/op and allocs/op per benchmark
go test -bench X -memprofile mem.out # then: go tool pprof -sample_index=alloc_space mem.out
GODEBUG=gctrace=1 ./app # how much CPU the collector is using
go build -gcflags=-m ./... # why something escapes
In pprof, alloc_space shows where bytes are allocated (the pressure on the GC) and
inuse_space shows what is live (your footprint). They answer different questions.
2. Preallocate#
out := make([]Result, 0, len(inputs)) // one allocation instead of ~log n
m := make(map[string]int, expected)
var sb strings.Builder; sb.Grow(n)
3. Let the caller own the buffer#
func AppendTokens(dst []int, text string) []int // appends and returns dst
buf = AppendTokens(buf[:0], text) // caller reuses one buffer for every call
buf[:0] keeps the capacity and resets the length: the idiom for reuse. The standard library
follows the same shape: strconv.AppendInt, fmt.Appendf, time.Time.AppendFormat,
binary.Append, hex.AppendEncode.
4. Keep values on the stack#
- Return small structs by value.
- Use fixed-size arrays for scratch space:
var tmp [64]byte. - Avoid
any,fmtand reflection in inner loops;strconvandslog’s typed attributes exist for this. - Do not take the address of a value unless you need to.
5. Avoid pointers in large live data#
From II.04 and lesson 04: the collector’s work is proportional to pointer-carrying live memory.
| Instead of | Use |
|---|---|
[]*Item | []Item |
map[string]*Item with millions of entries | map[string]int32 indexing a []Item; or sorted keys and binary search |
| A tree of nodes with child pointers | A slice of nodes with int32 child indexes |
Millions of small strings | One []byte plus offsets |
time.Time in huge slices (it contains a pointer) | int64 Unix nanoseconds |
An arena-style layout — one big slice from which you carve by index and which you drop all at once — makes allocation a pointer bump and deallocation free.
6. sync.Pool#
A pool holds temporary objects for reuse across goroutines:
var bufPool = sync.Pool{New: func() any { return new(bytes.Buffer) }}
b := bufPool.Get().(*bytes.Buffer)
b.Reset()
defer bufPool.Put(b)
How it behaves:
- Each P has a private slot and a local list, so
Get/Putare usually lock-free. - The collector empties pools: objects survive at most two GC cycles. A pool is a cache that smooths allocation between collections, not a guaranteed store.
- Store pointers (
*bytes.Buffer,*[]byte): putting a slice value in boxes it into an interface, which itself allocates. - Reset before use, and do not return a buffer to the pool while anything still references it — that is a use-after-free, the kind of bug Go otherwise spares you.
- Cap what you return. One request that grew a buffer to 50 MB will otherwise keep 50 MB alive in the pool. Drop buffers above a threshold.
Use a pool when objects are expensive to allocate, short-lived and used at a high rate: request buffers, encoders, scratch tensors. For a fixed number of workers, per-worker scratch space — a struct of buffers owned by each goroutine — is simpler and faster than a pool.
7. Strings and bytes#
Avoid round trips between string and []byte (each one copies). Work in []byte through the
pipeline and convert once at the edge. Intern repeated strings with unique.Make. The
zero-copy conversions unsafe.String and unsafe.Slice exist; they are correct only if the
bytes are never modified afterwards, and they belong behind a small, tested API.
8. Tuning last#
After the code is reasonable, GOGC and GOMEMLIMIT (lesson 04) trade spare memory for less
GC CPU with no code at all. A heap ballast — a large unused allocation to inflate the GC
goal — was a common trick before GOMEMLIMIT existed; it is no longer needed.
What it looks like in AI code#
Module VI’s inner loops are matrix multiplies and attention over []float32. The rules there:
allocate tensors once per model or per request, never per token; give every worker its own
scratch buffers; keep weights in flat pointer-free slices so that a multi-gigabyte model costs
the collector nothing; and make the per-token decode step allocation-free, because it runs
thousands of times per second.
Code#
The same task — tokenize lines and count token IDs — written naively, then with each technique.
// allocs.go — one task, four implementations, measured.
package main
import (
"fmt"
"strings"
"sync"
"testing"
)
var corpus = strings.Repeat("the quick brown fox jumps over the lazy dog\n", 200)
func hash(w []byte) int {
h := 0
for _, c := range w {
h = h*31 + int(c)
}
return h & 1023
}
// 1. Naive: Split allocates every line and word; strings are converted; the slice grows.
func naive(text string) []int {
var ids []int
for _, line := range strings.Split(text, "\n") {
for _, w := range strings.Fields(line) {
ids = append(ids, hash([]byte(w)))
}
}
return ids
}
// 2. Preallocated and scanning bytes in place: one allocation for the result.
func prealloc(text string) []int {
ids := make([]int, 0, len(text)/4)
return appendIDs(ids, text)
}
// appendIDs is the allocation-free core: the caller owns dst.
func appendIDs(dst []int, text string) []int {
start := -1
h := 0
for i := 0; i < len(text); i++ {
c := text[i]
if c == ' ' || c == '\n' {
if start >= 0 {
dst = append(dst, h&1023)
start, h = -1, 0
}
continue
}
if start < 0 {
start = i
}
h = h*31 + int(c)
}
if start >= 0 {
dst = append(dst, h&1023)
}
return dst
}
// 3. Caller-owned buffer, reused across calls.
type worker struct{ ids []int }
func (w *worker) run(text string) []int {
w.ids = appendIDs(w.ids[:0], text)
return w.ids
}
// 4. sync.Pool, for when callers are many goroutines.
var pool = sync.Pool{New: func() any { s := make([]int, 0, 4096); return &s }}
func pooled(text string) int {
p := pool.Get().(*[]int)
ids := appendIDs((*p)[:0], text)
n := len(ids)
*p = ids
pool.Put(p)
return n
}
func main() {
w := &worker{}
impls := []struct {
name string
f func()
}{
{"1 naive", func() { _ = naive(corpus) }},
{"2 preallocate, scan bytes", func() { _ = prealloc(corpus) }},
{"3 reuse a worker buffer", func() { _ = w.run(corpus) }},
{"4 sync.Pool", func() { _ = pooled(corpus) }},
}
fmt.Println("implementation ns/op B/op allocs/op")
for _, im := range impls {
r := testing.Benchmark(func(b *testing.B) {
b.ReportAllocs()
for i := 0; i < b.N; i++ {
im.f()
}
})
fmt.Printf("%-28s %8d %8d %10d\n", im.name, r.NsPerOp(), r.AllocedBytesPerOp(), r.AllocsPerOp())
}
fmt.Println("\nSame output, same algorithm. The difference is who allocates, and how often.")
}
Remember this#
- Profile first; optimize allocations only on the hot path.
- Preallocate, let the caller own buffers (
AppendX(dst, …)), keep values on the stack. - Large live data should be flat and pointer-free.
sync.Poolsmooths allocation between collections; reset, cap, and never use afterPut.- Tune
GOGC/GOMEMLIMITafter the code, not instead of it.
Try it#
- Run
allocs.go. Which single change removed the most allocations? The most time? - Introduce a bug: have
worker.runreturnw.idsand call it twice, keeping both results. What happens to the first result? This is the price of reuse. - Take a function of your own, run it under
-benchmem, and remove one allocation using-gcflags=-mto find it.
Check yourself#
- What is the difference between
alloc_spaceandinuse_spacein a heap profile? - Why should a
sync.Poolhold pointers rather than slice values? - Why does replacing
[]*Itemwith[]Itemhelp the garbage collector?