The idea in one minute#
Between your program and the GPU’s transistors sit several layers: frameworks, libraries of ready-made kernels, the CUDA runtime, the driver, and — in containers — a small piece of glue that makes the GPU visible. When something breaks (“CUDA version mismatch”, “no GPU found”), the fix depends on knowing which layer is complaining.
A picture#
flowchart TB
APP["Your application<br/>Go service, Python script"]
SRV["Inference server or framework<br/>vLLM, Triton, PyTorch, llama.cpp"]
LIB["Kernel libraries<br/>cuBLAS, cuDNN, TensorRT, NCCL"]
RT["CUDA runtime<br/>libcudart"]
DRV["NVIDIA driver<br/>kernel module + libcuda, libnvidia-ml"]
HW[("GPU hardware")]
APP --> SRV --> LIB --> RT --> DRV --> HW
NVML["NVML<br/>management: go-nvml, nvidia-smi"] --> DRV
class APP,SRV compute
class LIB,RT io
class DRV,NVML neutral
class HW memoryHow it really works#
The layers, bottom up#
| Layer | What it is | Lives |
|---|---|---|
| Driver | Kernel module that owns the hardware, plus user-space libcuda and libnvidia-ml | On the host; one version per machine |
| CUDA runtime / toolkit | libcudart, the nvcc compiler, headers | With the application (often inside the container) |
| Kernel libraries | Hand-tuned kernels: cuBLAS (linear algebra), cuDNN (neural-net layers), cuFFT, NCCL (multi-GPU communication), TensorRT (optimizing compiler) | With the application |
| Framework / server | Chooses and sequences kernels: PyTorch, ONNX Runtime, vLLM, llama.cpp, Triton | With the application |
| Application | Your code | — |
The one compatibility rule#
The driver must be at least as new as the CUDA runtime the application was built for. Drivers are backward compatible; a new driver runs old CUDA applications. An old driver cannot run a newer CUDA runtime.
So: keep host drivers current, and let each application bring its own CUDA runtime and libraries. This is exactly how containers are used.
GPUs in containers#
A container shares the host’s kernel, so it must use the host’s driver — but should carry its own CUDA libraries. The NVIDIA Container Toolkit handles the split: when a container starts, it injects the host’s driver libraries and GPU device files into the container. The image supplies everything above the driver.
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
If that prints a GPU table, the stack is healthy up to the driver.
Two APIs, two purposes#
- CUDA — compute. Allocate memory, launch kernels.
- NVML (NVIDIA Management Library) — management. Read temperature, utilization, memory
use, running processes, errors.
nvidia-smiis a command-line front end to NVML, andgo-nvmlis NVIDIA’s official Go binding. Kubernetes’ GPU components are Go programs built on it.
Beyond NVIDIA#
| Vendor | Compute API | Notes |
|---|---|---|
| NVIDIA | CUDA | Largest library ecosystem; the default in AI |
| AMD | ROCm / HIP | CUDA-like; supported by the major frameworks |
| Intel | oneAPI / SYCL | CPUs, GPUs and accelerators |
| Apple | Metal / MLX | Unified memory on M-series chips |
| Cross-vendor | Vulkan compute, WebGPU, OpenCL | Portable, fewer AI libraries |
The concepts from this module — kernels, device memory, queues, synchronization — carry to all of them with different names.
Code#
Read GPU state from Go. This is the program every GPU-aware Go tool starts from. It uses
nvidia-smi’s machine-readable output, so it needs no cgo and no extra modules; module IV
shows the go-nvml version.
// gpus.go — list GPUs from Go using nvidia-smi's CSV output.
package main
import (
"encoding/csv"
"fmt"
"os/exec"
"strings"
)
type GPU struct {
Index, Name, Driver string
MemUsedMiB, MemTotalMiB, UtilPct string
TempC, PowerW string
}
func List() ([]GPU, error) {
out, err := exec.Command("nvidia-smi",
"--query-gpu=index,name,driver_version,memory.used,memory.total,utilization.gpu,temperature.gpu,power.draw",
"--format=csv,noheader,nounits").Output()
if err != nil {
return nil, fmt.Errorf("nvidia-smi not available: %w", err)
}
rows, err := csv.NewReader(strings.NewReader(string(out))).ReadAll()
if err != nil {
return nil, err
}
var gpus []GPU
for _, r := range rows {
for i := range r {
r[i] = strings.TrimSpace(r[i])
}
gpus = append(gpus, GPU{r[0], r[1], r[2], r[3], r[4], r[5], r[6], r[7]})
}
return gpus, nil
}
func main() {
gpus, err := List()
if err != nil {
fmt.Println(err)
fmt.Println("(no NVIDIA GPU here — that is fine; read the code and move on)")
return
}
for _, g := range gpus {
fmt.Printf("GPU %s %-24s driver %s %s/%s MiB util %s%% %s C %s W\n",
g.Index, g.Name, g.Driver, g.MemUsedMiB, g.MemTotalMiB, g.UtilPct, g.TempC, g.PowerW)
}
}
Remember this#
- Driver (host) → CUDA runtime → libraries → framework → application.
- The driver must be at least as new as the application’s CUDA runtime.
- In containers, the driver comes from the host and everything else from the image.
- CUDA computes; NVML manages. Go tooling mostly uses NVML.
Try it#
- Run
gpus.go. On a machine without a GPU, confirm it fails cleanly. - Extend it into a loop that prints one line per second, like a tiny
nvidia-smi dmon. - An application container was built for CUDA 12.4 and fails on an old host with a driver from 2021. Which layer do you upgrade, and why not the other one?
Check yourself#
- State the driver/runtime compatibility rule.
- What does the NVIDIA Container Toolkit inject into a container?
- What is the difference between CUDA and NVML?