Below the API

Go in the AI Stack

Expert Advanced 40 min Difficulty 2/5

Prerequisites modules I–V (skim is fine if you already write Go)

The idea in one minute#

An AI system has two very different kinds of code. The numerical core — training loops and tensor kernels — runs on GPUs through C++, CUDA and Python, and Go is not the language for it. Everything around the core is systems programming: API gateways, schedulers, batching servers, tokenization, retrieval pipelines, agents that call tools, operators that manage GPUs. That part is concurrent, long-running, network-heavy and latency-sensitive — exactly what Go is for.

So the honest answer to “Go for AI?” is: not for inventing models, very much for serving them and for building products on them.

An analogy#

A restaurant. The chef’s knives and ovens (the kernels) are specialist equipment, and you buy the best that exist. But most of what makes the restaurant run — taking orders, sequencing the kitchen, managing tables, deliveries, the till — is logistics. A great oven does not fix a kitchen that loses orders.

A picture#

flowchart TB
  subgraph GO["Where Go is strong"]
    APP["Applications and agents<br/>tool calls, retrieval, workflows"]
    GW["Gateways and routers<br/>auth, quotas, streaming, cache-aware routing"]
    SRV["Serving layer<br/>queues, batching, backpressure"]
    OPS["Platform<br/>Kubernetes operators, GPU device plugins, telemetry"]
    DATA["Data pipelines<br/>tokenize, chunk, embed, index"]
  end
  subgraph NATIVE["Where native code and Python dominate"]
    ENG["Inference engines<br/>vLLM, SGLang, llama.cpp, TensorRT"]
    KRN["Kernels<br/>CUDA, cuBLAS, FlashAttention"]
    TRN["Training and research<br/>PyTorch, JAX"]
  end
  APP --> GW --> SRV
  SRV -->|"HTTP, gRPC, or cgo"| ENG --> KRN
  OPS -.-> ENG
  DATA --> APP
  class APP,GW,SRV,OPS,DATA compute
  class ENG,KRN io
  class TRN neutral

How it really works#

Why the split falls where it does#

NeedGoPython + native
Many concurrent connections, streamingGoroutines and the netpoller (IV.01, V.05)Async frameworks; the GIL limits CPU-side parallelism
Predictable latency under loadCompiled, low-pause GC, cheap schedulingInterpreter overhead on the request path
DeploymentOne static binaryAn environment of packages and a runtime
GPU kernels, autodiffNone native; cgo or a separate processThe entire ecosystem
Research velocity, notebooksWeakThe standard
New model architectures on day oneNoYes

The usual production shape follows directly: Python or C++ engines own the GPU; Go owns the network and the control plane. Much of that control plane already is Go — Kubernetes, the NVIDIA GPU operator and device plugin, Prometheus, the llm-d router, Ollama’s server.

Three ways Go runs a model#

ApproachHowUse when
Call a model serverHTTP or gRPC to an engine in another processThe default. Clean boundary, independent scaling, any model
Bind a native runtimecgo or purego to llama.cpp, ONNX Runtime, a vendor libraryYou need one self-contained binary, or an edge deployment
Pure GoImplement the math in GoSmall models, embeddings on CPU, learning how it works

This module does the third to teach, and the first to build real things.

What exists (checked 3 October 2026)#

Names here are to orient you, not to endorse. Versions and maintenance status change; check each project before depending on it.

AreaGo options
Model provider clientsOfficial SDKs from OpenAI (openai-go), Anthropic (anthropic-sdk-go) and Google (google.golang.org/genai). Any OpenAI-compatible endpoint — vLLM, Ollama, llama.cpp’s server — works with the same client by changing the base URL
Tool and context protocolThe official Model Context Protocol Go SDK, for building MCP servers and clients
Application frameworksGenkit for Go (Google; 1.0 released), Google’s Agent Development Kit for Go, Eino (CloudWeGo/ByteDance), LangChainGo
Running models locallyOllama (a Go server around llama.cpp); bindings to llama.cpp and ONNX Runtime
NumericsGonum (matrices, statistics, optimization); GoMLX (tensors and autodiff with accelerated back ends)
Vector searchWeaviate and Milvus are written largely in Go; client libraries for pgvector, Qdrant and others; small embedded stores
GPU managementgo-nvml; the NVIDIA device plugin, GPU operator and DCGM exporter are Go programs
Serving and routingllm-d and the Gateway API Inference Extension endpoint picker; many API gateways

Two things to know about the frameworks. First, the model vendors’ own agent SDKs have been Python- and TypeScript-first, so in Go you either use Genkit, ADK or Eino, or — very commonly — write the loop yourself, because it is short (lesson 09). Second, an LLM API is JSON over HTTP with a streaming response; the standard library is enough to use one well (lesson 06), and understanding that makes every SDK easier to debug.

Recent Go features that matter here#

FeatureSinceWhy it matters for AI services
Container-aware GOMAXPROCS1.25Correct parallelism in CPU-limited pods without extra libraries
Green Tea garbage collector1.26Less GC CPU for allocation-heavy request handling
Cheaper cgo calls1.26Bindings to native runtimes cost less per call
encoding/json/v21.27Faster decoding and stricter defaults for the JSON every LLM API speaks
simd packages (experimental)1.26–1.27A path to vectorized kernels in Go itself
Iterators1.23Natural streaming APIs: for tok := range stream
testing/synctest1.25Deterministic tests for timeouts, retries and batching windows

The map of this module#

LessonBuildsLeans on
02 Tensors and matmulFlat storage, strides, a parallel blocked matmulII.01, V.02, V.03
03 Neural networkForward, backward, training loop02
04 TokenizerByte-pair encodingII.02, II.03
05 TransformerAttention, KV cache, sampling02, III.05
06 Calling LLMsStreaming client, retries, structured outputIV.03, V.05
07 Vector searchSimilarity, top-k, quantizationV.03, IV.06
08 ServingDynamic batching, backpressureIV.02, IV.06
09 Tool-calling loopThe agent loop06, IV.03

Code#

Before building a model, know what one costs. This is the arithmetic you will reuse for the rest of the module — and for every capacity conversation afterwards.

// modelcost.go — parameters, bytes and FLOPs per token for a decoder-only transformer.
package main

import "fmt"

type Config struct {
	Name                  string
	Layers, DModel, Heads int
	Vocab, Context        int
}

// params counts weights: embeddings, then per layer attention (4 d^2) and an MLP (8 d^2).
func (c Config) params() int64 {
	d := int64(c.DModel)
	embed := int64(c.Vocab) * d // tied with the output projection
	perLayer := 4*d*d + 8*d*d   // q,k,v,o projections + a 4x-wide MLP
	return embed + int64(c.Layers)*perLayer
}

// kvBytesPerToken is what the KV cache stores for every token of context.
func (c Config) kvBytesPerToken(bytesPerValue int) int64 {
	return 2 * int64(c.Layers) * int64(c.DModel) * int64(bytesPerValue)
}

func human(b float64) string {
	for _, u := range []string{"B", "kB", "MB", "GB", "TB"} {
		if b < 1024 {
			return fmt.Sprintf("%.1f %s", b, u)
		}
		b /= 1024
	}
	return fmt.Sprintf("%.1f PB", b)
}

func main() {
	models := []Config{
		{"this module's toy", 2, 64, 4, 512, 64},
		{"small (GPT-2 class)", 12, 768, 12, 50257, 1024},
		{"8B class", 32, 4096, 32, 128000, 8192},
		{"70B class", 80, 8192, 64, 128000, 8192},
	}
	fmt.Println("model                  params      FP32      FP16      INT4   KV/token   KV at full context   FLOPs/token")
	for _, m := range models {
		p := float64(m.params())
		kv := m.kvBytesPerToken(2)
		fmt.Printf("%-20s %8.1fM  %8s  %8s  %8s  %9s  %19s  %12.2e\n",
			m.Name, p/1e6, human(p*4), human(p*2), human(p*0.5),
			human(float64(kv)), human(float64(kv)*float64(m.Context)), 2*p)
	}
	fmt.Println("\nFLOPs per token ≈ 2 × parameters: every weight takes part in one multiply and one add.")
	fmt.Println("That is why decoding is limited by how fast memory can be read, not by arithmetic.")
}

Remember this#

  • Go’s place in AI is the system around the model: clients, gateways, serving, pipelines, agents, and the platform.
  • Three ways to run a model from Go: call a server (default), bind a native runtime, or pure Go (small models and learning).
  • An LLM API is JSON over HTTP with a streamed response; the standard library handles it.
  • Know the cost arithmetic: parameters × bytes, KV cache per token, ~2 FLOPs per parameter per token.

Try it#

  1. Run modelcost.go. Add a model you use; compare the parameter estimate with its published size. Which components did the simple formula leave out?
  2. List the components of an AI product you know and mark each as “numerical core” or “systems around it”. Which language is each written in today?
  3. Pick one library from the table and find its latest release date and open-issue count. Would you depend on it?

Check yourself#

  1. Why is Go a poor fit for training and a good fit for serving?
  2. What are the three ways a Go program can run a model?
  3. Roughly how many floating-point operations does one generated token cost?

↑↓ navigate ↵ open