The idea in one minute#
Pure Go compiles to scalar code: one number per instruction. Numerical work — and AI is mostly numerical work — wants SIMD: instructions that process 4, 8 or 16 numbers at once. Go gives you three ways to get there, each with a price.
cgo calls C libraries (BLAS, llama.cpp, ONNX Runtime, CUDA): the fastest kernels that
exist, at the cost of a slower call, a C toolchain, and losing easy cross-compilation.
Go assembly is how the standard library vectorizes its own hot loops: full control, no
dependencies, one implementation per CPU architecture. The simd packages, experimental in
Go 1.26–1.27, expose vector instructions as ordinary Go functions.
The rule that decides between them: a foreign call is only worth it when it does a lot of work per call.
An analogy#
A specialist workshop across town. For a big job — machining a hundred parts — the trip is nothing compared with the time saved. For tightening one screw, you spend an hour travelling to save ten seconds. cgo is the trip; the question is always how much work you carry per journey.
A picture#
flowchart TB
GO["Go code on a goroutine stack"] --> CALL{"C.sgemm(...)"}
CALL --> SW["runtime.cgocall<br/>mark the goroutine as in a syscall,<br/>switch to the thread's system stack"]
SW --> C["C function runs<br/>on a real OS thread stack"]
C --> BACK["switch back, reacquire a P"]
BACK --> GO2["Go continues"]
SW -.->|"if the C call is slow"| P["the P is handed to another thread<br/>so other goroutines keep running"]
subgraph COST["Cost per call"]
N1["Go function call: ~1 ns"]
N2["cgo call: tens of ns"]
N3["Worth it when the C work is microseconds or more"]
end
class GO,GO2 compute
class CALL,SW,BACK queue
class C io
class P neutral
class N1,N2,N3 memoryHow it really works#
cgo#
/*
#cgo LDFLAGS: -lopenblas
#include <cblas.h>
*/
import "C"
func Sgemm(m, n, k int, a, b, c []float32) {
C.cblas_sgemm(C.CblasRowMajor, C.CblasNoTrans, C.CblasNoTrans,
C.int(m), C.int(n), C.int(k), 1,
(*C.float)(&a[0]), C.int(k), (*C.float)(&b[0]), C.int(n), 0,
(*C.float)(&c[0]), C.int(n))
}
What a cgo call does: the goroutine leaves the Go scheduler’s world, switches from its small growable stack to the thread’s large fixed stack (C code cannot run on a stack that might move), runs the C function, and comes back. That costs tens of nanoseconds — roughly 30% less since Go 1.26 — against about one for a Go call.
| Property | Consequence |
|---|---|
| Per-call overhead | Call C once per batch or per matrix, never once per element |
| The C call occupies an OS thread | A thousand concurrent slow C calls means a thousand threads (IV.01) |
| Pointer-passing rules | You may pass a Go pointer to C for the duration of the call; C must not keep it, and the memory must not itself contain Go pointers. runtime.Pinner pins objects for longer |
| C memory is invisible to the GC | C.malloc needs C.free. A model loaded in C does not appear in Go’s heap statistics, and GOMEMLIMIT knows nothing about it |
| Builds need a C toolchain | Cross-compiling stops being two environment variables; static linking and container images get harder |
| Tooling | pprof stops at the boundary; the race detector does not see C; a crash in C takes the process down |
CGO_ENABLED=0 forces a pure-Go build; the standard library then uses its own DNS resolver and
user lookup.
Avoiding cgo while still calling native code: purego-style libraries load a shared
library and call it without a C compiler at build time (the calls are still foreign calls);
running the native code as a separate process behind HTTP or gRPC — a model server — moves
the boundary to the network and keeps your Go binary pure. For inference this is the common
architecture: Go for the gateway, scheduler and control plane; a native engine for the kernels
(Inference Engineering VIII.01).
SIMD#
A vector register holds several values: 128-bit NEON (ARM64) or SSE holds four float32;
256-bit AVX2 holds eight; 512-bit AVX-512 sixteen. One instruction adds or multiplies all of
them. For a dot product that is up to a 4–16× speed-up over scalar code, on top of the
multiple-accumulator trick from lesson 02.
The Go compiler does not auto-vectorize. Your options:
| Option | Status (October 2026) | Trade-off |
|---|---|---|
Hand-written Go assembly (.s files) | Stable; what crypto, math/big and bytes use | Fastest; one file per architecture; hard to write and maintain |
Generators such as avo, or translating C intrinsics | Third party | Easier to write; still per-architecture |
simd/archsimd — architecture-specific vector types (Float32x8, …) | Experimental since Go 1.26: GOEXPERIMENT=simd | Vector code in Go syntax; API not yet stable; you still write per-architecture code |
simd — portable, vector-size-agnostic API | Experimental, added in Go 1.27: GOEXPERIMENT=simd | One source for all architectures; newest and least settled |
| A cgo or purego call to BLAS | Stable | Best kernels available; the costs listed above |
| Pure-Go numeric libraries (Gonum, and others with assembly kernels) | Stable | Convenient; check which operations are actually vectorized |
A sketch of what the experimental API looks like, to show the shape rather than to copy:
//go:build goexperiment.simd
import "simd/archsimd"
func dotAVX2(a, b []float32) float32 {
var acc archsimd.Float32x8
for len(a) >= 8 {
va := archsimd.LoadFloat32x8Slice(a)
vb := archsimd.LoadFloat32x8Slice(b)
acc = acc.Add(va.Mul(vb))
a, b = a[8:], b[8:]
}
// ...reduce acc to a scalar and handle the tail
}
Because it is experimental, names and signatures may change between releases; check the package documentation for your Go version before using it.
Runtime dispatch. Code that uses AVX2 crashes on a CPU without it. Libraries check
golang.org/x/sys/cpu (cpu.X86.HasAVX2, cpu.ARM64.HasASIMD) once at start-up and pick an
implementation. Building with GOAMD64=v3 lets the compiler use newer instructions
everywhere, in exchange for refusing to start on old CPUs.
Go assembly, in one paragraph#
Go has its own assembler with a portable-looking syntax (Plan 9 style) and pseudo-registers
(FP, SP, SB). A .s file next to a .go file with a body-less declaration
(func dotAsm(a, b []float32) float32) links the two. Assembly functions cannot be inlined, so
they too must do enough work per call. Read the standard library’s bytes/ or crypto/
directories for real examples; write your own only after a profile points at one loop and the
pure-Go version has been unrolled and bounds-check-free.
Choosing#
Is it a standard dense kernel (matmul, convolution, a whole model)?
→ use a native library or a model server. Do not rewrite BLAS.
Is it a small custom loop on a hot path?
→ pure Go: flat slices, no bounds checks, 4-8 accumulators. Often enough.
→ still hot? assembly or the simd experiment, with a pure-Go fallback and tests comparing both.
Is the call tiny and frequent?
→ keep it in Go. The boundary costs more than the work.GPUs from Go#
There is no first-class GPU programming model in Go. Go reaches GPUs through cgo bindings to
CUDA or vendor libraries, through go-nvml for management and telemetry, or — most commonly —
by talking to a process that owns the GPU.
GPU Engineering III.02
calls a CUDA kernel from Go with cgo and discusses the trade-offs.
Code#
No C toolchain needed: this program simulates the economics of a foreign-call boundary with a fixed per-call cost, and shows where batching makes it disappear.
// boundary.go — when does a call across an expensive boundary pay for itself?
package main
import (
"fmt"
"testing"
)
var sink float32
// kernel stands in for a vectorized native routine: 4 accumulators, no bounds checks.
func kernel(a, b []float32) float32 {
b = b[:len(a)]
var s0, s1, s2, s3 float32
i := 0
for ; i+4 <= len(a); i += 4 {
s0 += a[i] * b[i]
s1 += a[i+1] * b[i+1]
s2 += a[i+2] * b[i+2]
s3 += a[i+3] * b[i+3]
}
for ; i < len(a); i++ {
s0 += a[i] * b[i]
}
return s0 + s1 + s2 + s3
}
func scalar(a, b []float32) (s float32) { // plain Go, one accumulator
b = b[:len(a)]
for i := range a {
s += a[i] * b[i]
}
return
}
// boundary simulates the fixed cost of leaving Go: stack switch, scheduler bookkeeping.
//
//go:noinline
func boundary(spin int) int {
x := 0
for i := 0; i < spin; i++ {
x += i
}
return x
}
var spinSink int
func viaBoundary(a, b []float32) float32 {
spinSink += boundary(60) // roughly a cgo call's overhead on this machine class
return kernel(a, b)
}
func main() {
fmt.Println("vector length pure Go scalar 'native' behind a boundary winner")
for _, n := range []int{4, 16, 64, 256, 4096, 65536} {
a, b := make([]float32, n), make([]float32, n)
for i := range a {
a[i], b[i] = float32(i%7), 0.5
}
rs := testing.Benchmark(func(tb *testing.B) {
for i := 0; i < tb.N; i++ {
sink = scalar(a, b)
}
})
rn := testing.Benchmark(func(tb *testing.B) {
for i := 0; i < tb.N; i++ {
sink = viaBoundary(a, b)
}
})
s := float64(rs.T.Nanoseconds()) / float64(rs.N)
v := float64(rn.T.Nanoseconds()) / float64(rn.N)
winner := "pure Go"
if v < s {
winner = "native"
}
fmt.Printf("%13d %11.1f ns %22.1f ns %s\n", n, s, v, winner)
}
fmt.Println("\nThe boundary is a fixed cost. Carry enough work across it, or stay in Go.")
}
Remember this#
- Go code is scalar; SIMD comes from assembly, the experimental
simdpackages, or native libraries. - A cgo call costs tens of nanoseconds, ties up a thread, hides memory from the GC and complicates builds. Make each call do a lot.
- For standard kernels use a native library or a separate model server; for small custom loops, well-written Go is often enough.
- Vector code needs runtime CPU-feature checks and a tested pure-Go fallback.
Try it#
- Run
boundary.go. At what length does crossing the boundary start to pay? - If you have a C compiler, write the smallest cgo program (
C.intadd) and benchmark the call against a Go function. How many nanoseconds is the boundary on your machine? - With Go 1.26 or newer and
GOEXPERIMENT=simd, read thesimd/archsimddocumentation and write a vector dot product. Test it against the scalar version on random inputs.
Check yourself#
- Why must a cgo call switch stacks?
- Why is calling C once per matrix fine and once per element not?
- What does
GOMEMLIMITnot know about in a program that loads a model through cgo?