PidokuInfra

Pointers and Struct Layout

Basic Intermediate 50 min Difficulty 3/5 Topic 04 of 06

Prerequisites I.02, 01

The idea in one minute#

A pointer is the memory address of a variable: eight bytes that say where something lives. Go pointers are safe — no arithmetic, never dangling, checked for nil — but they still decide what is shared and what the garbage collector has to follow.

A struct is its fields laid out in declaration order. Each field must start at an address that is a multiple of its alignment, so the compiler inserts padding. The same fields in a different order can make a struct 30% smaller. And a struct that contains no pointers is invisible to the garbage collector’s scan.

An analogy#

Packing a van with boxes that must sit on a grid: large boxes only fit at positions divisible by eight, small ones anywhere. Load a small box, then a large one, and you leave a gap. Load the large ones first and everything sits flush.

A picture#

flowchart TB
  subgraph BAD["type Bad struct: 24 bytes"]
    direction LR
    B0["A bool<br/>1 byte"] --- B1["padding<br/>7 bytes"] --- B2["B int64<br/>8 bytes"] --- B3["C bool<br/>1 byte"] --- B4["padding<br/>7 bytes"]
  end
  subgraph GOOD["type Good struct: 16 bytes"]
    direction LR
    G0["B int64<br/>8 bytes"] --- G1["A bool<br/>1"] --- G2["C bool<br/>1"] --- G3["padding<br/>6 bytes"]
  end
  BAD -->|"reorder: largest first"| GOOD
  class B0,B2,B3,G0,G1,G2 memory
  class B1,B4,G3 warn

How it really works#

Pointers#

Go
x := 42
p := &x        // p is a *int holding x's address
*p = 43        // write through it: x is now 43
var q *int     // nil; dereferencing it panics
n := new(int)  // a pointer to a fresh zeroed int
  • No pointer arithmetic. (unsafe.Pointer and unsafe.Add exist for the rare cases that need it; using them takes you outside the language’s guarantees.)
  • It is fine to return the address of a local variable. The compiler notices that the variable outlives the function and places it on the heap (III.02). There are no dangling pointers.
  • Since Go 1.26, new accepts an expression: new(42) is a *int pointing at 42 — convenient for optional fields: Timeout: new(30 * time.Second).

Pointer or value? Pass and store a pointer when the callee must mutate, when the value is large, or when it must have one identity (a mutex, a connection). Otherwise prefer values: they are simpler, need no nil checks, and often avoid a heap allocation and a GC scan.

Alignment and padding#

TypeSizeAlignment
bool, int8, uint811
int1622
int32, float3244
int64, float64, int, pointers88
string168
slice248
interface168

Rules, on 64-bit platforms:

  1. Each field starts at a multiple of its own alignment.
  2. The struct’s alignment is its largest field’s alignment.
  3. The struct’s size is rounded up to a multiple of its alignment (so arrays of it stay aligned).

Ordering fields from largest alignment to smallest minimizes padding. The Go compiler does not reorder for you. unsafe.Sizeof, unsafe.Alignof and unsafe.Offsetof report the layout; go vet -vettool with the fieldalignment analyzer suggests better orders.

Does it matter? For one struct, no. For a slice of ten million, 24 versus 16 bytes is 80 MB and a third more cache misses. And the allocator rounds each object up to a size class (III.03), so shaving a 33-byte struct to 32 saves 16 bytes per object, not one.

Embedding by value versus by pointer#

Go
type A struct { Meta Meta;  Data [4]float32 }   // one block of memory, one allocation
type B struct { Meta *Meta; Data []float32 }    // up to three blocks, two pointers to chase

A value field is inside the struct: adjacent in memory, no extra allocation, loaded with the same cache miss. A pointer field is an 8-byte address of something elsewhere: more allocations, more cache misses, more work for the GC. Go lets you choose per field — a real difference from languages where every object is a reference.

Pointer-free data and the garbage collector#

When the GC marks live memory it must look inside every object for pointers. For each type the compiler records which words are pointers. A type with none — []float32, []int64, a struct of numbers — is marked “no scan”: the collector notes it is alive and moves on without reading it. A 10 GB slice of floats costs the GC nothing per cycle; a 10 GB slice of *T or string must be walked word by word.

Design consequences for large in-memory data (indexes, embeddings, caches):

  • Store numbers in flat slices; refer to items by integer index, not pointer.
  • Keep strings in one big byte buffer with offsets rather than as millions of string values.
  • Put the pointer-carrying fields of a struct first: the GC stops scanning an object after its last pointer word.

Empty structs and zero-size fields#

struct{} occupies zero bytes: map[K]struct{} for sets, chan struct{} for signals. One caveat: a zero-size field at the end of a struct gets padding, so that a pointer to it cannot point past the object.

Code#

Go
// layout.go — field offsets, padding, and what field order costs across a large slice.
package main

import (
	"fmt"
	"runtime"
	"time"
	"unsafe"
)

type Bad struct {
	A bool
	B int64
	C bool
	D int32
	E bool
}

type Good struct {
	B int64
	D int32
	A bool
	C bool
	E bool
}

type WithPointers struct {
	ID   int64
	Name string // contains a pointer: the GC must scan this struct
}

type NoPointers struct {
	ID      int64
	NameOff uint32 // offset into a shared byte buffer instead
	NameLen uint32
}

func heap() uint64 {
	runtime.GC()
	var m runtime.MemStats
	runtime.ReadMemStats(&m)
	return m.HeapAlloc
}

func main() {
	var bad Bad
	var good Good
	fmt.Println("Bad: ", unsafe.Sizeof(bad), "bytes; offsets A,B,C,D,E =",
		unsafe.Offsetof(bad.A), unsafe.Offsetof(bad.B), unsafe.Offsetof(bad.C), unsafe.Offsetof(bad.D), unsafe.Offsetof(bad.E))
	fmt.Println("Good:", unsafe.Sizeof(good), "bytes; offsets B,D,A,C,E =",
		unsafe.Offsetof(good.B), unsafe.Offsetof(good.D), unsafe.Offsetof(good.A), unsafe.Offsetof(good.C), unsafe.Offsetof(good.E))

	const n = 5_000_000
	h0 := heap()
	bs := make([]Bad, n)
	h1 := heap()
	gs := make([]Good, n)
	h2 := heap()
	fmt.Printf("\n%d elements: []Bad %.0f MB, []Good %.0f MB\n", n, float64(h1-h0)/1e6, float64(h2-h1)/1e6)
	runtime.KeepAlive(bs)
	runtime.KeepAlive(gs)
	bs, gs = nil, nil

	// How long does a full GC cycle take when the data is full of pointers, or free of them?
	measure := func(name string, keep any) {
		runtime.GC() // settle
		start := time.Now()
		for i := 0; i < 5; i++ {
			runtime.GC()
		}
		fmt.Printf("%-28s 5 GC cycles: %6.1f ms\n", name, float64(time.Since(start).Microseconds())/1000)
		runtime.KeepAlive(keep)
	}
	wp := make([]WithPointers, n)
	for i := range wp {
		wp[i].Name = "tenant"
	}
	measure("[]WithPointers (scanned)", wp)
	wp = nil
	np := make([]NoPointers, n)
	measure("[]NoPointers (not scanned)", np)
}

Remember this#

  • A pointer is an address; Go’s are safe, and returning &local is fine.
  • Struct fields are laid out in order with padding for alignment; order large-to-small.
  • Value fields live inside the struct; pointer fields mean another allocation and another hop.
  • Pointer-free types are not scanned by the garbage collector. For big data, prefer flat slices of numbers and integer indexes.

Try it#

  1. Run layout.go. Add a field to Good that pushes it to the next multiple of eight. What does the slice cost now?
  2. Define a struct with fields of sizes 1, 8, 2, 4, 1, 8. Compute its size by hand for the given order and for the best order, then check with unsafe.Sizeof.
  3. Represent one million short names two ways: []string, and one []byte plus []uint32 offsets. Compare heap size.

Check yourself#

  1. Why does field order change a struct’s size?
  2. What is the difference in memory between a value field and a pointer field?
  3. Why does the garbage collector not need to read a []float32?

↑↓ navigate↵ openesc close