The idea in one minute#
Debugging with telemetry is a search. Start from the symptom the alert reported, and narrow by repeatedly asking one question: what is different about the bad requests compared with the good ones? Each signal narrows a different way: metrics tell you when and how much, grouping tells you which, traces tell you where in the request, profiles and logs tell you why.
Two habits make people fast: check what changed first — a deploy, a config, a traffic shift explains most incidents — and mitigate before you understand.
An analogy#
A doctor does not begin with an MRI. Vital signs, then “where does it hurt?”, then a focused test, then treatment — and if the patient is bleeding, stop the bleeding before finishing the diagnosis.
A picture#
flowchart TB
A["Alert: error budget burning"] --> S["Scope<br/>since when, how bad, who"]
S --> CH{"What changed<br/>at that time?"}
CH -->|"a deploy or config"| RB["Roll back first<br/>investigate afterwards"]
CH -->|"nothing obvious"| SL["Slice the symptom<br/>by version, route, tenant, model, zone, host"]
SL --> ONE["One dimension stands out"]
ONE --> TR["Open a trace from that slice<br/>via an exemplar"]
TR --> SPAN["Which span is slow or failing?"]
SPAN --> WHY["Logs and profile of that component"]
WHY --> FIX["Mitigate, then fix"]
RB --> FIX
FIX --> LEARN["Add the missing signal"]
class A warn
class S,CH,SL queue
class ONE,TR,SPAN,WHY compute
class RB,FIX,LEARN neutralHow it really works#
Step 1 — scope#
Three numbers, in under a minute:
- Since when? Widen the graph until you see the start. Sudden step, or gradual ramp?
- How bad? Error ratio or latency against the SLO; how much budget is gone.
- Who? All users, or one tenant, one region, one model?
A step change points to an event (deploy, failover, a dependency). A ramp points to something accumulating (a leak, a queue, a filling disk, growing traffic).
Step 2 — what changed#
Look at annotations for the moment the graph bent: deploys, config pushes, feature flags, autoscaling events, node replacements, a certificate rotation, a traffic spike. If something changed, undo it and see whether the symptom goes with it. Understanding can wait; users cannot.
Step 3 — slice#
Group the symptom by each dimension and look for the one where “bad” concentrates:
sum by (version) (rate(http_requests_total{code=~"5.."}[5m]))
sum by (instance) (rate(http_requests_total{code=~"5.."}[5m]))
sum by (route) (rate(http_requests_total{code=~"5.."}[5m]))
With wide events this is one query per field, and some tools automate it: they compare the events inside the bad period with a baseline and rank the fields that differ most. It is the same idea as the program in I.01.
Typical findings: one version (bad deploy), one instance (bad host), one zone (infrastructure), one tenant (their traffic changed), one route or model (a specific code path).
Step 4 — trace#
Use an exemplar or a trace search (service = X, status = error, duration > 2 s) to open a
few bad requests and a few good ones. Compare waterfalls:
| What you see | Usually means |
|---|---|
| One child span much longer than usual | That dependency is slow |
| A long gap before the first child | Waiting: a queue, a lock, a connection pool |
| Many identical sequential spans | An N+1 loop or a retry storm |
| The trace ends abruptly at a hop | Timeout, crash, or broken propagation |
| Everything slower in proportion | Resource starvation on that host |
Step 5 — explain#
Now you know which component and roughly what it was doing. Read its logs filtered by the trace ID. Compare its CPU or heap profile with yesterday’s. Check its USE metrics: is it saturated, throttled, out of memory, restarting?
Step 6 — mitigate, fix, learn#
Mitigations in rough order of safety: roll back, shift traffic away, scale out, shed load, restart, disable a feature. Then write down the timeline while it is fresh, and ask the question that improves the system: what signal would have made this ten minutes instead of an hour? Add it.
Traps#
- Correlation. Two graphs that move together may share a cause. A cause should precede its effect and survive slicing.
- Looking where the light is. The dashboard you have is not necessarily where the problem is.
- Averages. A healthy mean with a broken tail; a healthy fleet with one broken instance.
- Trusting a flat line. Zero errors because there are no errors, or because the metric stopped being reported?
- Changing several things at once. You will not know which one worked.
A worked example#
09:14 Page: API error budget fast burn (14x).
09:15 Scope: error ratio 3%, started 09:02 as a step. All regions.
09:16 Changed at 09:02? A gateway config push. Roll it back.
09:19 Error ratio unchanged. Not the config. Restore it.
09:21 Slice: by route — only /v1/chat. By model — only "large". By instance — only 2 of 12.
09:23 Those two instances: restarted at 09:01 (node replacement), same node pool.
09:25 Trace from an exemplar: the request reaches the server and fails in 40 ms, "CUDA out of memory".
09:27 Compare instances: the two bad ones run on a GPU type with half the memory.
09:29 Mitigate: cordon that node pool for this model; errors stop.
09:45 Fix: add a GPU-memory requirement to the deployment.
Later Learn: add GPU model as a resource attribute; alert when a replica's KV-cache
capacity differs from its peers.Notice step 09:16: the first hypothesis was wrong, and it cost three minutes because it was tested quickly and reversibly.
Code#
The slicing step, automated: rank the fields by how unevenly failures are distributed across their values.
// slice.go — given request events, find which field best explains the failures.
package main
import (
"fmt"
"math/rand"
"sort"
)
type Event map[string]string
func main() {
rng := rand.New(rand.NewSource(21))
pick := func(xs ...string) string { return xs[rng.Intn(len(xs))] }
var events []Event
for i := 0; i < 30000; i++ {
e := Event{
"version": pick("1.8.1", "1.8.2"),
"route": pick("/v1/chat", "/v1/embeddings", "/v1/models"),
"model": pick("small", "large"),
"zone": pick("a", "b", "c"),
"gpu": pick("gpu-80gb", "gpu-80gb", "gpu-80gb", "gpu-40gb"),
"tenant": pick("acme", "globex", "initech", "umbrella"),
"failed": "false",
}
p := 0.002
if e["gpu"] == "gpu-40gb" && e["model"] == "large" && e["route"] == "/v1/chat" {
p = 0.7
}
if rng.Float64() < p {
e["failed"] = "true"
}
events = append(events, e)
}
type score struct {
field, worst string
lift float64
}
var scores []score
overallBad := 0
for _, e := range events {
if e["failed"] == "true" {
overallBad++
}
}
overall := float64(overallBad) / float64(len(events))
for _, f := range []string{"version", "route", "model", "zone", "gpu", "tenant"} {
total, bad := map[string]int{}, map[string]int{}
for _, e := range events {
total[e[f]]++
if e["failed"] == "true" {
bad[e[f]]++
}
}
best := score{field: f}
for v := range total {
lift := float64(bad[v]) / float64(total[v]) / overall
if lift > best.lift {
best.lift, best.worst = lift, v
}
}
scores = append(scores, best)
}
sort.Slice(scores, func(i, j int) bool { return scores[i].lift > scores[j].lift })
fmt.Printf("overall failure ratio: %.2f%%\n\n", overall*100)
fmt.Println("field worst value failure ratio vs overall")
for _, s := range scores {
fmt.Printf("%-8s %-15s %5.1fx\n", s.field, s.worst, s.lift)
}
fmt.Println("\nFields near 1.0x do not explain the failures. Start with the top of the list,")
fmt.Println("then slice again inside that value.")
}
Remember this#
- Scope, check what changed, slice, trace, explain, mitigate.
- The core question: how do bad requests differ from good ones?
- Mitigate before you understand; test hypotheses quickly and reversibly.
- End every incident by adding the signal that would have shortened it.
Try it#
- Run
slice.go. Extend it to slice a second time inside the top field’s worst value. Does it find the full three-field cause? - Rewrite the worked example with one missing signal (no GPU attribute). Where does the investigation stall?
- Take a past incident from your own experience and label each step with the stage above. How long did each take?
Check yourself#
- What does a step change suggest, compared with a ramp?
- Why roll back before understanding?
- What does a long gap before a span’s first child usually mean?