Learn to see inside running systems: what to measure, how the measurements are collected and stored, how to turn them into decisions — and what changes when the system is a rack of GPUs generating tokens.
This course starts at “what is a metric?” and ends at “design the telemetry for a multi-tenant inference platform”. It assumes you can program in Go. It does not assume you have run Prometheus, read a trace, or been on call.
It leans on two other paths for the systems being observed — Inference Engineering and GPU Engineering — and on Go Engineering if you want the language itself in depth. You can take modules I–IV without any of them.
How this course works#
Every lesson follows the same shape:
- The idea in one minute — the whole lesson in a few sentences.
- An analogy — something from everyday life with the same shape.
- A picture — a hand-drawn diagram of the mechanism.
- How it really works — the precise version, with real names and numbers.
- Code — a small Go program you can run. Almost all of them use only the standard library.
- Remember this — the three or four facts worth keeping.
- Try it and Check yourself — exercises and questions.
Why build the tools before using them?#
You will not use a hand-written metrics library in production; you will use Prometheus client libraries and OpenTelemetry. But a counter is twenty lines of Go, a histogram is forty, and a trace context is one HTTP header. Building each one once removes the mystery, and the mystery is what makes people instrument badly: labels with unbounded values, averages of percentiles, alerts on causes instead of symptoms. Each lesson builds the mechanism first and then shows the production tool that does the same job.
flowchart LR
APP["Your service<br/>counters, spans, logs"] --> COL["Collection<br/>scrape or push"]
COL --> STORE[("Storage<br/>time series, columns, objects")]
STORE --> Q["Query<br/>PromQL, SQL, trace search"]
Q --> DASH["Dashboards"]
Q --> ALERT["Alerts"]
ALERT --> HUMAN["A person<br/>or an agent"]
HUMAN -->|"asks a new question"| Q
class APP compute
class COL io
class STORE memory
class Q,ALERT queue
class DASH,HUMAN neutralThe path#
flowchart LR I["I Why observability"] --> II["II Metrics"] II --> III["III Logs, traces, profiles"] III --> IV["IV Running it"] IV --> V["V AI infrastructure"] V --> VI["VI Tools and frontier"] class I neutral class II,III compute class IV queue class V memory class VI io
| Module | You will be able to | Level |
|---|---|---|
| I — Why Observability | Say what to measure and why; define an SLO; read a percentile correctly | Foundations |
| II — Metrics | Instrument a Go service, write PromQL, choose histogram buckets, keep cardinality bounded | Basic |
| III — Logs, Traces and Profiles | Emit structured events, propagate a trace, use OpenTelemetry, read a profile | Basic |
| IV — Running It | Design the pipeline, build dashboards that answer questions, alert on SLO burn, control cost, work an incident | Intermediate |
| V — AI Infrastructure | Observe GPUs, inference engines, LLM requests, a serving fleet, cost and power, and output quality | Advanced |
| VI — Tools and Frontier | Map the tool landscape, explain how the field is changing, and know what to watch | Expert |
| Projects | Build five Go programs that make the ideas stick | All |
What you need#
- Go 1.22 or newer and a terminal.
- Nothing else for modules I–III: every program runs offline.
- Docker (or Podman) for the optional exercises that start Prometheus, Grafana or an OpenTelemetry Collector.
- No GPU required. Module V explains GPU and engine telemetry with simulators and real metric names; the exercises that read a real GPU say so and give a no-GPU alternative.
Getting started in fifteen minutes#
- Read I.01 and run its program.
- Read II.01 and
II.02;
curlyour own/metricsendpoint. - If you came for AI infrastructure and already know Prometheus, jump to V.01 and come back to module II’s histograms and cardinality lessons when a percentile or a bill surprises you.
A promise about names and versions#
Metric names, tool versions and project statuses in this course were checked on 3 October 2026. Mechanisms — what a counter is, why a percentile cannot be averaged, why a GPU-utilization gauge misleads — do not change. Names do: a lesson says so wherever a name comes from a specification that is still marked unstable, and VI.03 lists where to check what has moved.