The idea in one minute#
An SLI (service level indicator) is a measurement of what users experience, written as a ratio: good events ÷ all events. An SLO (service level objective) is a target for that ratio over a window: 99.5% of requests succeed, over 30 days. The error budget is what is left: the 0.5% you are allowed to fail.
This turns “is the service good?” into arithmetic, and gives you a principled answer to “should we page someone?”: page when the budget is being spent too fast.
An analogy#
A monthly mobile data plan. The plan is the SLO. Usage is the SLI. The remaining gigabytes are the error budget. You do not panic at every megabyte; you worry when the rate of use means you will run out before the month ends.
A picture#
flowchart TB
EV["Every request"] --> G{"Good?<br/>succeeded and fast enough"}
G -->|"yes"| GOOD["good count"]
G -->|"no"| BAD["bad count"]
GOOD --> SLI["SLI = good / total"]
BAD --> SLI
SLI --> CMP{"SLI vs SLO<br/>over 30 days"}
CMP --> BUD["Error budget left"]
BUD -->|"plenty"| SHIP["Ship features, take risk"]
BUD -->|"burning fast"| PAGE["Page someone"]
BUD -->|"gone"| FREEZE["Slow down, fix reliability"]
class EV neutral
class G,CMP queue
class GOOD,BAD,SLI,BUD memory
class PAGE,FREEZE warn
class SHIP computeHow it really works#
Writing an SLI#
A good SLI is a ratio of events, measured as close to the user as you can get.
| Kind | Good event | Example |
|---|---|---|
| Availability | Request returned a non-5xx response | good = status < 500 |
| Latency | Request finished within a threshold | good = duration < 300 ms |
| Quality | Response was complete and correct | good = not truncated, no fallback used |
| Freshness | Data newer than a threshold | good = age < 60 s |
Two rules:
- Count events, do not average measurements. “99% of requests under 300 ms” is an SLI. “Average latency under 300 ms” hides the slow tail (lesson 04).
- Exclude what the user caused. A 400 for a malformed request is not your failure. A 429 because you ran out of capacity is.
Choosing the target#
| SLO | Allowed bad time in 30 days |
|---|---|
| 99% | 7 h 12 min |
| 99.5% | 3 h 36 min |
| 99.9% | 43 min |
| 99.95% | 21.6 min |
| 99.99% | 4.3 min |
Each extra nine costs roughly ten times more engineering. Choose the lowest target your users would not notice, not the highest you can imagine. 100% is never the right answer: it forbids all change.
An SLA is a contract with penalties. Keep the internal SLO stricter than the SLA so you learn about trouble before you owe money.
Error budget and burn rate#
error budget = 1 − SLO (99.9% → 0.1% of requests)
burn rate = observed error ratio ÷ error budgetBurn rate 1 means you will spend exactly the budget over the window. Burn rate 14.4 means a 30-day budget is gone in 50 hours — or 2% of it in one hour. Alerting on burn rate (IV.03) is how you get paged for real problems and left alone for blips.
SLOs for LLM serving, briefly#
Streaming generation needs more than one latency SLI, because the user experiences two different waits:
| SLI | Good event |
|---|---|
| Time to first token (TTFT) | First token within, say, 500 ms |
| Time per output token (TPOT) | Each token within, say, 50 ms on average for that request |
| Availability | Stream completed without a server error |
A request is good only if it met all of them. The fraction of requests that did is called goodput when expressed as a rate, and it is the number a serving fleet is sized against. V.03 builds on this.
Code#
// budget.go — SLI, error budget and burn rate from raw counts.
package main
import "fmt"
type Window struct {
Name string
Total, Good float64
Hours float64
}
func main() {
const slo = 0.999 // 99.9% over 30 days
const windowHours = 30 * 24.0
budget := 1 - slo
fmt.Printf("SLO %.2f%% → error budget %.2f%% of requests, or %.0f min of total outage per 30 days\n\n",
slo*100, budget*100, budget*windowHours*60)
windows := []Window{
{"quiet hour", 360000, 359900, 1},
{"bad deploy, 1 h", 360000, 354600, 1},
{"slow leak, 6 h", 2160000, 2153500, 6},
}
fmt.Println("window SLI error ratio burn rate budget used exhausts in")
for _, w := range windows {
sli := w.Good / w.Total
errRatio := 1 - sli
burn := errRatio / budget
used := burn * w.Hours / windowHours
exhaust := "never at this rate"
if burn > 1 {
exhaust = fmt.Sprintf("%.1f days", windowHours/burn/24)
}
fmt.Printf("%-18s %.4f%% %9.3f%% %8.1fx %10.1f%% %s\n",
w.Name, sli*100, errRatio*100, burn, used*100, exhaust)
}
fmt.Println("\nPage when burn rate is high over a short AND a longer window (lesson IV.03).")
}
Remember this#
- SLI = good events ÷ total events, measured near the user.
- SLO = a target for the SLI over a window. Error budget = 1 − SLO.
- Burn rate says how fast the budget is being spent; it is what you alert on.
- LLM serving needs several SLIs at once (TTFT, per-token speed, completion); a request is good only if it meets all of them.
Try it#
- Run
budget.go. Change the SLO to 99.5%. Which windows still look alarming? - Write an SLI for a service you use, as a precise good/total definition. Which requests did you exclude, and why?
- A team says “our SLO is 100%”. Write two sentences explaining what that would forbid.
Check yourself#
- Why is an SLI a ratio of events and not an average?
- How many minutes of total outage does 99.9% allow in 30 days?
- What does a burn rate of 6 mean?