The idea in one minute#
An AI system has two very different kinds of code. The numerical core — training loops and tensor kernels — runs on GPUs through C++, CUDA and Python, and Go is not the language for it. Everything around the core is systems programming: API gateways, schedulers, batching servers, tokenization, retrieval pipelines, agents that call tools, operators that manage GPUs. That part is concurrent, long-running, network-heavy and latency-sensitive — exactly what Go is for.
So the honest answer to “Go for AI?” is: not for inventing models, very much for serving them and for building products on them.
An analogy#
A restaurant. The chef’s knives and ovens (the kernels) are specialist equipment, and you buy the best that exist. But most of what makes the restaurant run — taking orders, sequencing the kitchen, managing tables, deliveries, the till — is logistics. A great oven does not fix a kitchen that loses orders.
A picture#
flowchart TB
subgraph GO["Where Go is strong"]
APP["Applications and agents<br/>tool calls, retrieval, workflows"]
GW["Gateways and routers<br/>auth, quotas, streaming, cache-aware routing"]
SRV["Serving layer<br/>queues, batching, backpressure"]
OPS["Platform<br/>Kubernetes operators, GPU device plugins, telemetry"]
DATA["Data pipelines<br/>tokenize, chunk, embed, index"]
end
subgraph NATIVE["Where native code and Python dominate"]
ENG["Inference engines<br/>vLLM, SGLang, llama.cpp, TensorRT"]
KRN["Kernels<br/>CUDA, cuBLAS, FlashAttention"]
TRN["Training and research<br/>PyTorch, JAX"]
end
APP --> GW --> SRV
SRV -->|"HTTP, gRPC, or cgo"| ENG --> KRN
OPS -.-> ENG
DATA --> APP
class APP,GW,SRV,OPS,DATA compute
class ENG,KRN io
class TRN neutralHow it really works#
Why the split falls where it does#
| Need | Go | Python + native |
|---|---|---|
| Many concurrent connections, streaming | Goroutines and the netpoller (IV.01, V.05) | Async frameworks; the GIL limits CPU-side parallelism |
| Predictable latency under load | Compiled, low-pause GC, cheap scheduling | Interpreter overhead on the request path |
| Deployment | One static binary | An environment of packages and a runtime |
| GPU kernels, autodiff | None native; cgo or a separate process | The entire ecosystem |
| Research velocity, notebooks | Weak | The standard |
| New model architectures on day one | No | Yes |
The usual production shape follows directly: Python or C++ engines own the GPU; Go owns the network and the control plane. Much of that control plane already is Go — Kubernetes, the NVIDIA GPU operator and device plugin, Prometheus, the llm-d router, Ollama’s server.
Three ways Go runs a model#
| Approach | How | Use when |
|---|---|---|
| Call a model server | HTTP or gRPC to an engine in another process | The default. Clean boundary, independent scaling, any model |
| Bind a native runtime | cgo or purego to llama.cpp, ONNX Runtime, a vendor library | You need one self-contained binary, or an edge deployment |
| Pure Go | Implement the math in Go | Small models, embeddings on CPU, learning how it works |
This module does the third to teach, and the first to build real things.
What exists (checked 3 October 2026)#
Names here are to orient you, not to endorse. Versions and maintenance status change; check each project before depending on it.
| Area | Go options |
|---|---|
| Model provider clients | Official SDKs from OpenAI (openai-go), Anthropic (anthropic-sdk-go) and Google (google.golang.org/genai). Any OpenAI-compatible endpoint — vLLM, Ollama, llama.cpp’s server — works with the same client by changing the base URL |
| Tool and context protocol | The official Model Context Protocol Go SDK, for building MCP servers and clients |
| Application frameworks | Genkit for Go (Google; 1.0 released), Google’s Agent Development Kit for Go, Eino (CloudWeGo/ByteDance), LangChainGo |
| Running models locally | Ollama (a Go server around llama.cpp); bindings to llama.cpp and ONNX Runtime |
| Numerics | Gonum (matrices, statistics, optimization); GoMLX (tensors and autodiff with accelerated back ends) |
| Vector search | Weaviate and Milvus are written largely in Go; client libraries for pgvector, Qdrant and others; small embedded stores |
| GPU management | go-nvml; the NVIDIA device plugin, GPU operator and DCGM exporter are Go programs |
| Serving and routing | llm-d and the Gateway API Inference Extension endpoint picker; many API gateways |
Two things to know about the frameworks. First, the model vendors’ own agent SDKs have been Python- and TypeScript-first, so in Go you either use Genkit, ADK or Eino, or — very commonly — write the loop yourself, because it is short (lesson 09). Second, an LLM API is JSON over HTTP with a streaming response; the standard library is enough to use one well (lesson 06), and understanding that makes every SDK easier to debug.
Recent Go features that matter here#
| Feature | Since | Why it matters for AI services |
|---|---|---|
Container-aware GOMAXPROCS | 1.25 | Correct parallelism in CPU-limited pods without extra libraries |
| Green Tea garbage collector | 1.26 | Less GC CPU for allocation-heavy request handling |
| Cheaper cgo calls | 1.26 | Bindings to native runtimes cost less per call |
encoding/json/v2 | 1.27 | Faster decoding and stricter defaults for the JSON every LLM API speaks |
simd packages (experimental) | 1.26–1.27 | A path to vectorized kernels in Go itself |
| Iterators | 1.23 | Natural streaming APIs: for tok := range stream |
testing/synctest | 1.25 | Deterministic tests for timeouts, retries and batching windows |
The map of this module#
| Lesson | Builds | Leans on |
|---|---|---|
| 02 Tensors and matmul | Flat storage, strides, a parallel blocked matmul | II.01, V.02, V.03 |
| 03 Neural network | Forward, backward, training loop | 02 |
| 04 Tokenizer | Byte-pair encoding | II.02, II.03 |
| 05 Transformer | Attention, KV cache, sampling | 02, III.05 |
| 06 Calling LLMs | Streaming client, retries, structured output | IV.03, V.05 |
| 07 Vector search | Similarity, top-k, quantization | V.03, IV.06 |
| 08 Serving | Dynamic batching, backpressure | IV.02, IV.06 |
| 09 Tool-calling loop | The agent loop | 06, IV.03 |
Code#
Before building a model, know what one costs. This is the arithmetic you will reuse for the rest of the module — and for every capacity conversation afterwards.
// modelcost.go — parameters, bytes and FLOPs per token for a decoder-only transformer.
package main
import "fmt"
type Config struct {
Name string
Layers, DModel, Heads int
Vocab, Context int
}
// params counts weights: embeddings, then per layer attention (4 d^2) and an MLP (8 d^2).
func (c Config) params() int64 {
d := int64(c.DModel)
embed := int64(c.Vocab) * d // tied with the output projection
perLayer := 4*d*d + 8*d*d // q,k,v,o projections + a 4x-wide MLP
return embed + int64(c.Layers)*perLayer
}
// kvBytesPerToken is what the KV cache stores for every token of context.
func (c Config) kvBytesPerToken(bytesPerValue int) int64 {
return 2 * int64(c.Layers) * int64(c.DModel) * int64(bytesPerValue)
}
func human(b float64) string {
for _, u := range []string{"B", "kB", "MB", "GB", "TB"} {
if b < 1024 {
return fmt.Sprintf("%.1f %s", b, u)
}
b /= 1024
}
return fmt.Sprintf("%.1f PB", b)
}
func main() {
models := []Config{
{"this module's toy", 2, 64, 4, 512, 64},
{"small (GPT-2 class)", 12, 768, 12, 50257, 1024},
{"8B class", 32, 4096, 32, 128000, 8192},
{"70B class", 80, 8192, 64, 128000, 8192},
}
fmt.Println("model params FP32 FP16 INT4 KV/token KV at full context FLOPs/token")
for _, m := range models {
p := float64(m.params())
kv := m.kvBytesPerToken(2)
fmt.Printf("%-20s %8.1fM %8s %8s %8s %9s %19s %12.2e\n",
m.Name, p/1e6, human(p*4), human(p*2), human(p*0.5),
human(float64(kv)), human(float64(kv)*float64(m.Context)), 2*p)
}
fmt.Println("\nFLOPs per token ≈ 2 × parameters: every weight takes part in one multiply and one add.")
fmt.Println("That is why decoding is limited by how fast memory can be read, not by arithmetic.")
}
Remember this#
- Go’s place in AI is the system around the model: clients, gateways, serving, pipelines, agents, and the platform.
- Three ways to run a model from Go: call a server (default), bind a native runtime, or pure Go (small models and learning).
- An LLM API is JSON over HTTP with a streamed response; the standard library handles it.
- Know the cost arithmetic: parameters × bytes, KV cache per token, ~2 FLOPs per parameter per token.
Try it#
- Run
modelcost.go. Add a model you use; compare the parameter estimate with its published size. Which components did the simple formula leave out? - List the components of an AI product you know and mark each as “numerical core” or “systems around it”. Which language is each written in today?
- Pick one library from the table and find its latest release date and open-issue count. Would you depend on it?
Check yourself#
- Why is Go a poor fit for training and a good fit for serving?
- What are the three ways a Go program can run a model?
- Roughly how many floating-point operations does one generated token cost?