Below the API

Learning path

Observability Engineering

Seeing inside running systems — metrics, logs, traces and profiles, then the GPU, token and cost telemetry that AI infrastructure adds.

5 levels · 7 modules · 30 topics · ~24h 25m completed ·
Start learning →

The path

  1. Foundations

    Build the mental model.

    4 topics · ~2h 40m

    Why Observability

    Before any tool: what are you trying to learn about a running system, and how do you state "working" precisely enough that a machine can check it? Module overview →

    1. Monitoring vs Observability30 min
    2. The Signals40 min
    3. SLIs, SLOs and Error Budgets45 min
    4. Percentiles and Tails45 min
  2. Basic

    Understand the core mechanisms.

    9 topics · ~7h 30m

    Metrics

    Metrics are the cheapest signal and the one every alert is built on. This module builds a metrics library by hand, then uses the real one, then learns to query it without being fooled. Module overview →

    1. Metric Types and the Data Model45 min
    2. Instrumenting a Go Service50 min
    3. PromQL Essentials55 min
    4. Histograms55 min
    5. Cardinality40 min
    Logs, Traces and Profiles

    Metrics tell you something is wrong. These three signals keep the detail needed to say what. Module overview →

    1. Structured Logs and Wide Events40 min
    2. Distributed Tracing55 min
    3. OpenTelemetry1h
    4. Profiling and eBPF50 min
  3. Intermediate

    Learn the optimization techniques.

    5 topics · ~4h 10m

    Running It

    Emitting telemetry is the easy half. This module is about the system that receives it, the screens and alerts built on it, what it costs, and how to use it when something is on fire. Module overview →

    1. The Telemetry Pipeline55 min
    2. Dashboards40 min
    3. Alerting on SLOs55 min
    4. Sampling and Cost50 min
    5. Debugging an Incident50 min
  4. Advanced

    Study systems at production scale.

    7 topics · ~6h 30m

    AI Infrastructure

    Everything so far applies to any service. This module is about what changes when the service is a model on a GPU: new resources to watch, new units of work, new failure modes — and a product … Module overview →

    1. What Is Different45 min
    2. GPU Telemetry1h
    3. Inference Engine Metrics1h 10m
    4. Tracing LLM Requests55 min
    5. Platform and Fleet1h
    6. Cost, Power and Efficiency50 min
    7. Quality and Evals50 min
  5. Expert

    Design platforms and read the frontier.

    4 topics · ~3h 35m

    Tools and Frontier

    The mechanisms in modules I–V are stable. The products that implement them are not: this module maps what exists, explains the direction the field is moving, and gives you a method for … Module overview →

    1. The Tool Landscape50 min
    2. How It Is Changing1h
    3. What to Watch45 min
    4. Designing Observability for an Inference Platform1h

Build

Reference

About this path

Learn to see inside running systems: what to measure, how the measurements are collected and stored, how to turn them into decisions — and what changes when the system is a rack of GPUs generating tokens.

This course starts at “what is a metric?” and ends at “design the telemetry for a multi-tenant inference platform”. It assumes you can program in Go. It does not assume you have run Prometheus, read a trace, or been on call.

It leans on two other paths for the systems being observed — Inference Engineering and GPU Engineering — and on Go Engineering if you want the language itself in depth. You can take modules I–IV without any of them.


How this course works#

Every lesson follows the same shape:

  1. The idea in one minute — the whole lesson in a few sentences.
  2. An analogy — something from everyday life with the same shape.
  3. A picture — a hand-drawn diagram of the mechanism.
  4. How it really works — the precise version, with real names and numbers.
  5. Code — a small Go program you can run. Almost all of them use only the standard library.
  6. Remember this — the three or four facts worth keeping.
  7. Try it and Check yourself — exercises and questions.

Why build the tools before using them?#

You will not use a hand-written metrics library in production; you will use Prometheus client libraries and OpenTelemetry. But a counter is twenty lines of Go, a histogram is forty, and a trace context is one HTTP header. Building each one once removes the mystery, and the mystery is what makes people instrument badly: labels with unbounded values, averages of percentiles, alerts on causes instead of symptoms. Each lesson builds the mechanism first and then shows the production tool that does the same job.

flowchart LR
  APP["Your service<br/>counters, spans, logs"] --> COL["Collection<br/>scrape or push"]
  COL --> STORE[("Storage<br/>time series, columns, objects")]
  STORE --> Q["Query<br/>PromQL, SQL, trace search"]
  Q --> DASH["Dashboards"]
  Q --> ALERT["Alerts"]
  ALERT --> HUMAN["A person<br/>or an agent"]
  HUMAN -->|"asks a new question"| Q
  class APP compute
  class COL io
  class STORE memory
  class Q,ALERT queue
  class DASH,HUMAN neutral

The path#

flowchart LR
  I["I Why observability"] --> II["II Metrics"]
  II --> III["III Logs, traces, profiles"]
  III --> IV["IV Running it"]
  IV --> V["V AI infrastructure"]
  V --> VI["VI Tools and frontier"]
  class I neutral
  class II,III compute
  class IV queue
  class V memory
  class VI io
ModuleYou will be able toLevel
I — Why ObservabilitySay what to measure and why; define an SLO; read a percentile correctlyFoundations
II — MetricsInstrument a Go service, write PromQL, choose histogram buckets, keep cardinality boundedBasic
III — Logs, Traces and ProfilesEmit structured events, propagate a trace, use OpenTelemetry, read a profileBasic
IV — Running ItDesign the pipeline, build dashboards that answer questions, alert on SLO burn, control cost, work an incidentIntermediate
V — AI InfrastructureObserve GPUs, inference engines, LLM requests, a serving fleet, cost and power, and output qualityAdvanced
VI — Tools and FrontierMap the tool landscape, explain how the field is changing, and know what to watchExpert
ProjectsBuild five Go programs that make the ideas stickAll

What you need#

  • Go 1.22 or newer and a terminal.
  • Nothing else for modules I–III: every program runs offline.
  • Docker (or Podman) for the optional exercises that start Prometheus, Grafana or an OpenTelemetry Collector.
  • No GPU required. Module V explains GPU and engine telemetry with simulators and real metric names; the exercises that read a real GPU say so and give a no-GPU alternative.

Getting started in fifteen minutes#

  1. Read I.01 and run its program.
  2. Read II.01 and II.02; curl your own /metrics endpoint.
  3. If you came for AI infrastructure and already know Prometheus, jump to V.01 and come back to module II’s histograms and cardinality lessons when a percentile or a bill surprises you.

A promise about names and versions#

Metric names, tool versions and project statuses in this course were checked on 3 October 2026. Mechanisms — what a counter is, why a percentile cannot be averaged, why a GPU-utilization gauge misleads — do not change. Names do: a lesson says so wherever a name comes from a specification that is still marked unstable, and VI.03 lists where to check what has moved.

↑↓ navigate ↵ open