The idea in one minute#
Monitoring answers questions you thought of in advance: is the error rate above 1%? Observability is the property of a system that lets you answer questions you did not think of in advance: why are requests from one customer, on one model, slow only when the prompt is long? — without shipping new code to find out.
Monitoring tells you that something is wrong. Observability lets you find out why. You need both, and the second is built from richer data, not from more dashboards.
An analogy#
A car dashboard is monitoring: a few fixed gauges and a warning light. A mechanic’s diagnostic port is observability: it lets them ask the engine arbitrary questions about what it was doing when the light came on.
The light is useful. It is not enough to fix the car.
A picture#
flowchart TB
SYS["Running system"] -->|"emits telemetry"| TEL[("Telemetry<br/>metrics, logs, traces, profiles")]
TEL --> KNOWN["Known questions<br/>dashboards and alerts"]
TEL --> UNKNOWN["New questions<br/>slice by any field"]
KNOWN -->|"something is wrong"| P["A person"]
P -->|"why?"| UNKNOWN
UNKNOWN -->|"which requests, which hosts, which change"| FIX["Cause found"]
class SYS compute
class TEL memory
class KNOWN,UNKNOWN queue
class P,FIX neutralHow it really works#
Known unknowns and unknown unknowns#
| Monitoring | Observability | |
|---|---|---|
| Question | Decided before the incident | Invented during the incident |
| Data | A few aggregated numbers | Detailed events with many fields |
| Typical tool | Dashboard, alert rule | Ad-hoc query, trace search, profile diff |
| Fails when | The failure is new | The data needed was never recorded |
A system is observable to the degree that its outputs let you reconstruct its internal state. The term comes from control theory; in software it reduces to one practical test:
Can I explain why this particular request was slow, using only what the system already emitted?
What makes a system observable#
- It emits events with context. Not
request failed, but which tenant, which model, which replica, which version, how many tokens, how long each phase took. - The context is consistent across components. The same request ID or trace ID appears in the gateway, the server and the log line, so you can follow one request through all of them.
- You can slice by any of those fields afterwards. The fields you need are rarely the ones you predicted.
What observability is not#
- Not three pillars you buy. Metrics, logs and traces are data types (lesson 02). Having all three in separate tools that cannot be joined is still poor observability.
- Not more dashboards. Fifty dashboards answer fifty predicted questions.
- Not free. Every event costs CPU, network and storage. Module IV is largely about deciding what to keep.
The loop you are building#
detect → an alert on a user-facing symptom fires
triage → how bad, who is affected, since when
localize → which component, version, tenant, host
explain → what that component was doing (trace, profile, logs)
fix → roll back, scale, patch
learn → add the missing signal so next time is fasterEach module of this course strengthens one step. Detection is modules I, II and IV; localizing and explaining is module III; module V repeats the whole loop for GPUs and LLM serving.
Code#
A service that only counts errors can tell you that 3% of requests fail. The same service recording one event per request can tell you which 3%.
// why.go — the same failures seen through a counter and through events.
package main
import (
"fmt"
"math/rand"
"sort"
)
type Event struct {
Tenant, Model, Replica string
PromptTokens int
Failed bool
}
func main() {
rng := rand.New(rand.NewSource(1))
tenants := []string{"acme", "globex", "initech"}
models := []string{"small", "large"}
replicas := []string{"r0", "r1", "r2", "r3"}
var events []Event
for i := 0; i < 20000; i++ {
e := Event{
Tenant: tenants[rng.Intn(len(tenants))],
Model: models[rng.Intn(len(models))],
Replica: replicas[rng.Intn(len(replicas))],
PromptTokens: 100 + rng.Intn(8000),
}
// The hidden bug: replica r2 runs out of memory on long prompts for the large model.
if e.Replica == "r2" && e.Model == "large" && e.PromptTokens > 6000 {
e.Failed = rng.Float64() < 0.9
}
events = append(events, e)
}
// Monitoring view: one number.
failed := 0
for _, e := range events {
if e.Failed {
failed++
}
}
fmt.Printf("monitoring: error rate = %.2f%% (something is wrong; no idea what)\n\n",
100*float64(failed)/float64(len(events)))
// Observability view: group failures by each field and see where they concentrate.
for _, dim := range []struct {
name string
key func(Event) string
}{
{"tenant", func(e Event) string { return e.Tenant }},
{"model", func(e Event) string { return e.Model }},
{"replica", func(e Event) string { return e.Replica }},
{"long prompt", func(e Event) string { return fmt.Sprint(e.PromptTokens > 6000) }},
} {
total, bad := map[string]int{}, map[string]int{}
for _, e := range events {
k := dim.key(e)
total[k]++
if e.Failed {
bad[k]++
}
}
keys := make([]string, 0, len(total))
for k := range total {
keys = append(keys, k)
}
sort.Strings(keys)
fmt.Printf("by %-12s", dim.name)
for _, k := range keys {
fmt.Printf(" %s=%.1f%%", k, 100*float64(bad[k])/float64(total[k]))
}
fmt.Println()
}
}The counter says roughly 3%. The grouped view shows the failures are evenly spread across tenants, and concentrated on one replica, one model and long prompts. That is the difference this course is about.
Remember this#
- Monitoring answers predicted questions; observability lets you answer new ones.
- Observability comes from events with rich, consistent context — not from more dashboards.
- The practical test: can you explain one specific slow request from what was already emitted?
- Telemetry has a cost; deciding what to keep is part of the engineering.
Try it#
- Run
why.go. Change the hidden bug to depend onTenantinstead ofReplica. Does the grouped view still find it? - Add a field the bug depends on but that you do not record (say, a kernel version). What does the grouped view show now? What does that teach you about choosing fields?
- For a service you know, write down one question you could not answer during its last incident. Which field was missing?
Check yourself#
- State the difference between monitoring and observability in one sentence.
- Why can a system with metrics, logs and traces still be hard to debug?
- What are the six steps of the incident loop?