The idea in one minute#
A dashboard is an answer to a question, laid out so that a tired person can read it in ten seconds. Good dashboards are layered: one overview that says whether users are happy, then one screen per service showing the RED metrics, then resource detail showing USE. You start at the top and descend only where something looks wrong.
Most bad dashboards share one fault: they show everything that can be measured rather than what someone needs to decide.
An analogy#
A hospital monitor shows heart rate, blood pressure and oxygen — big, at the top. It does not show the full blood panel; that is a separate report you request when the vital signs are off.
A picture#
flowchart TB L1["Level 1: Service health<br/>SLO status, error budget, are users affected"] --> L2["Level 2: One service, RED<br/>rate, errors, duration by route or model"] L2 --> L3["Level 3: Resources, USE<br/>CPU, memory, queue depth, GPU"] L3 --> L4["Level 4: Drill-down<br/>traces, logs, profiles for one instance"] class L1 queue class L2 compute class L3 memory class L4 neutral
How it really works#
Layout rules#
- Symptoms at the top, causes below. The top row is what users feel.
- Read left to right as a sentence: traffic → errors → latency → saturation.
- One question per panel, stated in the title: “Error ratio by route”, not “Errors”.
- Units and thresholds on the graph. Draw the SLO line.
- Same time range and same colours everywhere. Errors are always the same colour.
- Mark deploys and config changes as annotations. Most incidents begin at one.
- Variables, not copies: one dashboard with a
serviceormodeldrop-down, not twenty near-identical ones.
The service dashboard#
| Row | Panels | Query shape |
|---|---|---|
| Traffic | Requests/s by route | sum by (route) (rate(requests_total[$__rate_interval])) |
| Errors | Error ratio, with the SLO line | errors ÷ total |
| Duration | p50, p95, p99; a heatmap beside it | histogram_quantile |
| Saturation | In-flight, queue depth, waiting | gauges |
| Dependencies | The same RED row for each downstream call | client-side metrics |
| Resources | CPU, memory, restarts, throttling | USE |
Choosing a visualization#
| You want to show | Use |
|---|---|
| A trend | Time series line |
| A distribution over time | Heatmap |
| A current value against a limit | Stat or gauge with thresholds |
| Ranking | Bar chart or table sorted by value |
| Many instances at once | A table, or a status grid — never forty overlapping lines |
| Parts of a whole over time | Stacked area, only when the parts genuinely add up |
Avoid pie charts for anything that changes, dual y-axes, and averaged latency.
Traps#
- Averages across instances hide the one that is broken. Show max, or p99 across instances, or a top-5 table.
- A graph that is flat at zero: is the service healthy, or is the metric missing? Show
upand scrape health. - Stacking latencies or percentiles. They do not add.
- Auto-scaled axes make noise look like an incident. Fix the y-axis from zero for ratios.
- Too long a rate window smooths away a short spike; too short makes noise. Use the dashboard’s interval-aware variable.
Dashboards as code#
Keep dashboards in version control and generate them — Grafana’s JSON model, Jsonnet/Grafonnet, Terraform, or the Grafana Foundation SDK (which has a Go builder). Reasons: review, reuse of a standard service row across teams, and no more “who changed this panel?”.
The mixin pattern packages dashboards, recording rules and alerts for one piece of software together. Before building a dashboard for Kubernetes, a database or an inference engine, look for an existing mixin or the vendor’s published dashboard and adapt it.
When a dashboard is the wrong tool#
Dashboards answer questions you anticipated. During a novel incident you need ad-hoc queries over events and traces (I.01). If every incident ends with “let me build a dashboard for that”, you are accumulating screens nobody will open again. Build a dashboard when a question recurs.
Code#
A dashboard is data. This program generates a standard RED row for any list of services, which is all “dashboards as code” means.
// dash.go — generate a RED dashboard definition from a list of services.
package main
import (
"encoding/json"
"fmt"
"os"
)
type Panel struct {
Title string `json:"title"`
Type string `json:"type"`
Unit string `json:"unit"`
Expr string `json:"expr"`
Row int `json:"row"`
}
func redRow(row int, svc string) []Panel {
sel := fmt.Sprintf(`{service=%q}`, svc)
return []Panel{
{svc + ": requests/s by route", "timeseries", "reqps",
fmt.Sprintf(`sum by (route) (rate(http_requests_total%s[$__rate_interval]))`, sel), row},
{svc + ": error ratio (SLO line at 0.1%)", "timeseries", "percentunit",
fmt.Sprintf(`sum(rate(http_requests_total{service=%q,code=~"5.."}[$__rate_interval])) / sum(rate(http_requests_total%s[$__rate_interval]))`, svc, sel), row},
{svc + ": p99 duration", "timeseries", "s",
fmt.Sprintf(`histogram_quantile(0.99, sum by (route) (rate(http_request_duration_seconds%s[$__rate_interval])))`, sel), row},
{svc + ": duration heatmap", "heatmap", "s",
fmt.Sprintf(`sum(rate(http_request_duration_seconds%s[$__rate_interval]))`, sel), row},
{svc + ": in flight (saturation)", "timeseries", "short",
fmt.Sprintf(`sum(http_in_flight_requests%s)`, sel), row},
}
}
func main() {
services := []string{"gateway", "inference", "retrieval"}
var panels []Panel
for i, s := range services {
panels = append(panels, redRow(i, s)...)
}
dash := map[string]any{
"title": "Service health (generated)",
"annotations": []string{"deploys", "config changes"},
"variables": []string{"cluster", "namespace"},
"panels": panels,
}
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.Encode(dash)
fmt.Fprintf(os.Stderr, "%d panels for %d services, all identical in shape\n", len(panels), len(services))
}Remember this#
- Layer dashboards: user-facing health → service RED → resource USE → drill-down.
- One question per panel; symptoms above causes; mark deploys.
- Never average across instances or stack percentiles.
- Keep dashboards in version control and generate the repetitive parts.
Try it#
- Run
dash.goand add amodelvariable to each query. - Open a dashboard you use. Delete (on paper) every panel you have never acted on. What is left?
- Sketch the Level 1 dashboard for a product with three services. It may have at most six panels.
Check yourself#
- What goes on the top row of a service dashboard?
- Why is averaging a metric across instances dangerous?
- When should you not build a new dashboard?