The idea in one minute#
An alert should mean: a human must act now, and here is what users are experiencing. That rules out most alerts people write. “CPU above 80%” is a cause that may or may not hurt anyone. “1% of requests are failing” is a symptom.
The dependable way to alert on symptoms is the error budget burn rate: page when the budget is being spent fast enough to run out soon, checked over a long window (so it is significant) and a short window (so it stops firing when the problem stops).
An analogy#
A smoke detector that sounds whenever the oven is on gets its battery removed. One that sounds only for real smoke gets obeyed. Every false page teaches people to ignore the next one; alert quality is a safety property.
A picture#
flowchart TB
SLI["Error ratio<br/>from recording rules"] --> W1["Long window: 1 h<br/>burn rate above 14.4?"]
SLI --> W2["Short window: 5 min<br/>burn rate above 14.4?"]
W1 --> AND{"Both true?"}
W2 --> AND
AND -->|"yes"| PAGE["Page<br/>2% of the monthly budget gone in an hour"]
AND -->|"no"| NEXT["Check the slower pair<br/>6 h and 30 min at 6x"]
NEXT -->|"yes"| PAGE
NEXT -->|"no"| TICKET["3 d and 6 h at 1x: open a ticket"]
class SLI memory
class W1,W2,AND,NEXT queue
class PAGE warn
class TICKET neutralHow it really works#
Symptoms, not causes#
| Alert on | Not on |
|---|---|
| Error ratio above budget burn | A single 500 |
| p99 latency SLI burning budget | CPU, memory or GPU utilization |
| Queue depth growing while goodput falls | Queue depth above a fixed number |
Disk will be full in four hours (predict_linear) | Disk at 80% |
| A scrape target has been down for minutes | One failed scrape |
Causes belong on dashboards. The exception is an imminent, certain failure — a disk filling, a certificate expiring — where the “cause” is the symptom arriving on a schedule.
Why a simple threshold fails#
error_ratio > 0.001 for 5m on a 99.9% SLO either fires for harmless blips (short for) or
takes an hour to notice a total outage (long for). The trouble is that one threshold cannot
express both “a lot of budget, quickly” and “a steady leak”.
Multi-window, multi-burn-rate#
Burn rate = error ratio ÷ error budget (I.03). For a 30-day SLO:
| Severity | Budget consumed | Long window | Short window | Burn rate |
|---|---|---|---|---|
| Page | 2% | 1 h | 5 min | 14.4 |
| Page | 5% | 6 h | 30 min | 6 |
| Ticket | 10% | 3 d | 6 h | 1 |
groups:
- name: slo-api
rules:
- record: job:slo_errors:ratio_rate5m
expr: |
sum(rate(http_requests_total{job="api",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="api"}[5m]))
- record: job:slo_errors:ratio_rate1h
expr: |
sum(rate(http_requests_total{job="api",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{job="api"}[1h]))
- alert: ErrorBudgetFastBurn
expr: |
job:slo_errors:ratio_rate1h > (14.4 * 0.001)
and job:slo_errors:ratio_rate5m > (14.4 * 0.001)
labels: { severity: page }
annotations:
summary: "API is burning its 30-day error budget 14x too fast"
runbook: https://runbooks.example.internal/api/error-budgetThe long window makes it significant; the short window makes it reset quickly once the problem is fixed. Tools such as Sloth and Pyrra generate these rules from a short SLO definition, and OpenSLO is a vendor-neutral format for writing the definition.
Low traffic#
With ten requests an hour, one failure is a 10% error ratio. Options: lengthen the windows, require a minimum request count, combine related services into one SLO, or add synthetic probes so there is always traffic to measure.
Routing and hygiene#
- Severity:
page(now, wakes someone),ticket(this week),info(dashboard only). - Group: one notification for “40 pods of service X are failing”, not forty.
- Inhibit: if the cluster is down, suppress alerts for everything inside it.
- Every page has a runbook link and names the user-visible impact.
- Review: for each page, was action required? If not, fix or delete the alert. Track pages per on-call shift; more than a couple is a problem to engineer away.
Alerts worth having beyond SLOs#
| Alert | Why |
|---|---|
| Dead man’s switch | The alerting pipeline itself is alive (IV.01) |
up == 0 for N minutes | A target vanished; its SLO alert would be silent |
| Absent metric | Instrumentation was removed or renamed |
| Certificate / quota / disk exhaustion predicted | Certain, scheduled failures |
| Telemetry volume anomaly | A cardinality or log explosion before the bill |
Code#
Simulate a month with three incidents and see which alerting rule catches what.
// burn.go — threshold alert vs multi-window burn-rate alert on a simulated error ratio.
package main
import "fmt"
const (
slo = 0.999
budget = 1 - slo
)
// window returns the mean error ratio over the last n minutes ending at t.
func window(errs []float64, t, n int) float64 {
lo := t - n + 1
if lo < 0 {
lo = 0
}
s := 0.0
for i := lo; i <= t; i++ {
s += errs[i]
}
return s / float64(t-lo+1)
}
func main() {
const minutes = 30 * 24 * 60
errs := make([]float64, minutes)
type incident struct {
name string
start, len int
ratio float64
}
incidents := []incident{
{"blip: 5% errors for 6 min", 5000, 6, 0.05},
{"outage: 30% errors for 40 min", 15000, 40, 0.30},
{"slow leak: 0.4% errors for 2 days", 25000, 2 * 24 * 60, 0.004},
}
for _, in := range incidents {
for i := in.start; i < in.start+in.len; i++ {
errs[i] = in.ratio
}
}
firstFire := func(rule func(t int) bool, from, to int) int {
for t := from; t < to; t++ {
if rule(t) {
return t - from
}
}
return -1
}
rules := []struct {
name string
f func(t int) bool
}{
{"threshold >0.1% for 5m", func(t int) bool {
for i := t - 4; i <= t; i++ {
if i < 0 || errs[i] <= budget {
return false
}
}
return true
}},
{"fast burn 14.4x (1h & 5m)", func(t int) bool {
return window(errs, t, 60) > 14.4*budget && window(errs, t, 5) > 14.4*budget
}},
{"slow burn 6x (6h & 30m)", func(t int) bool {
return window(errs, t, 360) > 6*budget && window(errs, t, 30) > 6*budget
}},
{"ticket 1x (3d & 6h)", func(t int) bool {
return window(errs, t, 4320) > budget && window(errs, t, 360) > budget
}},
}
fmt.Printf("%-36s", "incident")
for _, r := range rules {
fmt.Printf(" %-26s", r.name)
}
fmt.Println()
for _, in := range incidents {
fmt.Printf("%-36s", in.name)
for _, r := range rules {
d := firstFire(r.f, in.start, in.start+in.len+4500)
if d < 0 {
fmt.Printf(" %-26s", "silent")
} else {
fmt.Printf(" %-26s", fmt.Sprintf("fires after %d min", d))
}
}
used := in.ratio * float64(in.len) / float64(minutes) / budget
fmt.Printf(" budget used: %.0f%%\n", used*100)
}
}Read the table by row. The blip uses under 1% of the budget, yet the fixed threshold pages for it; the burn-rate rules stay silent. The outage is caught by the fast-burn rule within a couple of minutes. The slow leak consumes over a quarter of the budget: the fixed threshold pages for it at once — at 3 a.m., for something that can wait until morning — while the burn-rate rules turn it into a ticket.
One property worth knowing: a total outage trips the fast-burn rule in under a minute, because 100% errors for one minute is already 1.7% over the hour. That is the intended behaviour, not noise.
Remember this#
- Page on symptoms users feel, at a rate that threatens the SLO.
- Burn rate over a long and a short window: significant, and quick to reset.
- Every page: actionable, urgent, with a runbook. Review and delete the rest.
- Also alert on the pipeline itself and on absence.
Try it#
- Run
burn.go. Change the SLO to 99.5% and see which rules change behaviour. - Write the three burn-rate rules for a latency SLI (“99% of requests under 500 ms”).
- Take five alerts from a real system and classify each as symptom or cause. Which would you delete?
Check yourself#
- Why does one fixed threshold fail for both blips and slow leaks?
- What does each of the two windows contribute?
- What three things must every page have?