Below the API

cgo, SIMD and Assembly

Advanced Advanced 55 min Difficulty 4/5

Prerequisites 02, 03, IV.01

The idea in one minute#

Pure Go compiles to scalar code: one number per instruction. Numerical work — and AI is mostly numerical work — wants SIMD: instructions that process 4, 8 or 16 numbers at once. Go gives you three ways to get there, each with a price.

cgo calls C libraries (BLAS, llama.cpp, ONNX Runtime, CUDA): the fastest kernels that exist, at the cost of a slower call, a C toolchain, and losing easy cross-compilation. Go assembly is how the standard library vectorizes its own hot loops: full control, no dependencies, one implementation per CPU architecture. The simd packages, experimental in Go 1.26–1.27, expose vector instructions as ordinary Go functions.

The rule that decides between them: a foreign call is only worth it when it does a lot of work per call.

An analogy#

A specialist workshop across town. For a big job — machining a hundred parts — the trip is nothing compared with the time saved. For tightening one screw, you spend an hour travelling to save ten seconds. cgo is the trip; the question is always how much work you carry per journey.

A picture#

flowchart TB
  GO["Go code on a goroutine stack"] --> CALL{"C.sgemm(...)"}
  CALL --> SW["runtime.cgocall<br/>mark the goroutine as in a syscall,<br/>switch to the thread's system stack"]
  SW --> C["C function runs<br/>on a real OS thread stack"]
  C --> BACK["switch back, reacquire a P"]
  BACK --> GO2["Go continues"]
  SW -.->|"if the C call is slow"| P["the P is handed to another thread<br/>so other goroutines keep running"]
  subgraph COST["Cost per call"]
    N1["Go function call: ~1 ns"]
    N2["cgo call: tens of ns"]
    N3["Worth it when the C work is microseconds or more"]
  end
  class GO,GO2 compute
  class CALL,SW,BACK queue
  class C io
  class P neutral
  class N1,N2,N3 memory

How it really works#

cgo#

/*
#cgo LDFLAGS: -lopenblas
#include <cblas.h>
*/
import "C"

func Sgemm(m, n, k int, a, b, c []float32) {
    C.cblas_sgemm(C.CblasRowMajor, C.CblasNoTrans, C.CblasNoTrans,
        C.int(m), C.int(n), C.int(k), 1,
        (*C.float)(&a[0]), C.int(k), (*C.float)(&b[0]), C.int(n), 0,
        (*C.float)(&c[0]), C.int(n))
}

What a cgo call does: the goroutine leaves the Go scheduler’s world, switches from its small growable stack to the thread’s large fixed stack (C code cannot run on a stack that might move), runs the C function, and comes back. That costs tens of nanoseconds — roughly 30% less since Go 1.26 — against about one for a Go call.

PropertyConsequence
Per-call overheadCall C once per batch or per matrix, never once per element
The C call occupies an OS threadA thousand concurrent slow C calls means a thousand threads (IV.01)
Pointer-passing rulesYou may pass a Go pointer to C for the duration of the call; C must not keep it, and the memory must not itself contain Go pointers. runtime.Pinner pins objects for longer
C memory is invisible to the GCC.malloc needs C.free. A model loaded in C does not appear in Go’s heap statistics, and GOMEMLIMIT knows nothing about it
Builds need a C toolchainCross-compiling stops being two environment variables; static linking and container images get harder
Toolingpprof stops at the boundary; the race detector does not see C; a crash in C takes the process down

CGO_ENABLED=0 forces a pure-Go build; the standard library then uses its own DNS resolver and user lookup.

Avoiding cgo while still calling native code: purego-style libraries load a shared library and call it without a C compiler at build time (the calls are still foreign calls); running the native code as a separate process behind HTTP or gRPC — a model server — moves the boundary to the network and keeps your Go binary pure. For inference this is the common architecture: Go for the gateway, scheduler and control plane; a native engine for the kernels (Inference Engineering VIII.01).

SIMD#

A vector register holds several values: 128-bit NEON (ARM64) or SSE holds four float32; 256-bit AVX2 holds eight; 512-bit AVX-512 sixteen. One instruction adds or multiplies all of them. For a dot product that is up to a 4–16× speed-up over scalar code, on top of the multiple-accumulator trick from lesson 02.

The Go compiler does not auto-vectorize. Your options:

OptionStatus (October 2026)Trade-off
Hand-written Go assembly (.s files)Stable; what crypto, math/big and bytes useFastest; one file per architecture; hard to write and maintain
Generators such as avo, or translating C intrinsicsThird partyEasier to write; still per-architecture
simd/archsimd — architecture-specific vector types (Float32x8, …)Experimental since Go 1.26: GOEXPERIMENT=simdVector code in Go syntax; API not yet stable; you still write per-architecture code
simd — portable, vector-size-agnostic APIExperimental, added in Go 1.27: GOEXPERIMENT=simdOne source for all architectures; newest and least settled
A cgo or purego call to BLASStableBest kernels available; the costs listed above
Pure-Go numeric libraries (Gonum, and others with assembly kernels)StableConvenient; check which operations are actually vectorized

A sketch of what the experimental API looks like, to show the shape rather than to copy:

//go:build goexperiment.simd

import "simd/archsimd"

func dotAVX2(a, b []float32) float32 {
    var acc archsimd.Float32x8
    for len(a) >= 8 {
        va := archsimd.LoadFloat32x8Slice(a)
        vb := archsimd.LoadFloat32x8Slice(b)
        acc = acc.Add(va.Mul(vb))
        a, b = a[8:], b[8:]
    }
    // ...reduce acc to a scalar and handle the tail
}

Because it is experimental, names and signatures may change between releases; check the package documentation for your Go version before using it.

Runtime dispatch. Code that uses AVX2 crashes on a CPU without it. Libraries check golang.org/x/sys/cpu (cpu.X86.HasAVX2, cpu.ARM64.HasASIMD) once at start-up and pick an implementation. Building with GOAMD64=v3 lets the compiler use newer instructions everywhere, in exchange for refusing to start on old CPUs.

Go assembly, in one paragraph#

Go has its own assembler with a portable-looking syntax (Plan 9 style) and pseudo-registers (FP, SP, SB). A .s file next to a .go file with a body-less declaration (func dotAsm(a, b []float32) float32) links the two. Assembly functions cannot be inlined, so they too must do enough work per call. Read the standard library’s bytes/ or crypto/ directories for real examples; write your own only after a profile points at one loop and the pure-Go version has been unrolled and bounds-check-free.

Choosing#

Is it a standard dense kernel (matmul, convolution, a whole model)?
  → use a native library or a model server. Do not rewrite BLAS.
Is it a small custom loop on a hot path?
  → pure Go: flat slices, no bounds checks, 4-8 accumulators.     Often enough.
  → still hot? assembly or the simd experiment, with a pure-Go fallback and tests comparing both.
Is the call tiny and frequent?
  → keep it in Go. The boundary costs more than the work.

GPUs from Go#

There is no first-class GPU programming model in Go. Go reaches GPUs through cgo bindings to CUDA or vendor libraries, through go-nvml for management and telemetry, or — most commonly — by talking to a process that owns the GPU. GPU Engineering III.02 calls a CUDA kernel from Go with cgo and discusses the trade-offs.

Code#

No C toolchain needed: this program simulates the economics of a foreign-call boundary with a fixed per-call cost, and shows where batching makes it disappear.

// boundary.go — when does a call across an expensive boundary pay for itself?
package main

import (
	"fmt"
	"testing"
)

var sink float32

// kernel stands in for a vectorized native routine: 4 accumulators, no bounds checks.
func kernel(a, b []float32) float32 {
	b = b[:len(a)]
	var s0, s1, s2, s3 float32
	i := 0
	for ; i+4 <= len(a); i += 4 {
		s0 += a[i] * b[i]
		s1 += a[i+1] * b[i+1]
		s2 += a[i+2] * b[i+2]
		s3 += a[i+3] * b[i+3]
	}
	for ; i < len(a); i++ {
		s0 += a[i] * b[i]
	}
	return s0 + s1 + s2 + s3
}

func scalar(a, b []float32) (s float32) { // plain Go, one accumulator
	b = b[:len(a)]
	for i := range a {
		s += a[i] * b[i]
	}
	return
}

// boundary simulates the fixed cost of leaving Go: stack switch, scheduler bookkeeping.
//
//go:noinline
func boundary(spin int) int {
	x := 0
	for i := 0; i < spin; i++ {
		x += i
	}
	return x
}

var spinSink int

func viaBoundary(a, b []float32) float32 {
	spinSink += boundary(60) // roughly a cgo call's overhead on this machine class
	return kernel(a, b)
}

func main() {
	fmt.Println("vector length   pure Go scalar   'native' behind a boundary   winner")
	for _, n := range []int{4, 16, 64, 256, 4096, 65536} {
		a, b := make([]float32, n), make([]float32, n)
		for i := range a {
			a[i], b[i] = float32(i%7), 0.5
		}
		rs := testing.Benchmark(func(tb *testing.B) {
			for i := 0; i < tb.N; i++ {
				sink = scalar(a, b)
			}
		})
		rn := testing.Benchmark(func(tb *testing.B) {
			for i := 0; i < tb.N; i++ {
				sink = viaBoundary(a, b)
			}
		})
		s := float64(rs.T.Nanoseconds()) / float64(rs.N)
		v := float64(rn.T.Nanoseconds()) / float64(rn.N)
		winner := "pure Go"
		if v < s {
			winner = "native"
		}
		fmt.Printf("%13d   %11.1f ns   %22.1f ns   %s\n", n, s, v, winner)
	}
	fmt.Println("\nThe boundary is a fixed cost. Carry enough work across it, or stay in Go.")
}

Remember this#

  • Go code is scalar; SIMD comes from assembly, the experimental simd packages, or native libraries.
  • A cgo call costs tens of nanoseconds, ties up a thread, hides memory from the GC and complicates builds. Make each call do a lot.
  • For standard kernels use a native library or a separate model server; for small custom loops, well-written Go is often enough.
  • Vector code needs runtime CPU-feature checks and a tested pure-Go fallback.

Try it#

  1. Run boundary.go. At what length does crossing the boundary start to pay?
  2. If you have a C compiler, write the smallest cgo program (C.int add) and benchmark the call against a Go function. How many nanoseconds is the boundary on your machine?
  3. With Go 1.26 or newer and GOEXPERIMENT=simd, read the simd/archsimd documentation and write a vector dot product. Test it against the scalar version on random inputs.

Check yourself#

  1. Why must a cgo call switch stacks?
  2. Why is calling C once per matrix fine and once per element not?
  3. What does GOMEMLIMIT not know about in a program that loads a model through cgo?

↑↓ navigate ↵ open