The idea in one minute#
Telemetry is a cost centre that grows faster than traffic unless you manage it. Each signal has one cost driver: metrics cost per active series, logs per gigabyte, traces per span. So each has one main lever: bound cardinality, reduce volume, and sample.
Sampling well means keeping the events that are interesting — errors, slow requests, rare customers — and a small, known fraction of the boring ones, while recording the sampling rate so that counts can still be reconstructed.
An analogy#
A quality inspector cannot open every box on the conveyor. Opening every hundredth box catches systematic faults. Opening every box that rattles catches the rest. Doing both, and writing down which rule selected each box, lets you still estimate the true defect rate.
A picture#
flowchart LR
ALL["All traces<br/>100%"] --> H{"Head sampling<br/>decide at the start"}
H -->|"keep 10%"| BUF["Buffer whole traces<br/>in the collector"]
H -->|"drop 90%"| GONE1["Gone, including<br/>90% of the errors"]
ALL -.->|"or skip head sampling"| BUF
BUF --> T{"Tail sampling<br/>decide when the trace ends"}
T -->|"error? keep"| KEEP[("Stored")]
T -->|"slow? keep"| KEEP
T -->|"otherwise keep 1%"| KEEP
T -->|"drop"| GONE2["Dropped"]
class ALL neutral
class H,T queue
class BUF,KEEP memory
class GONE1 warn
class GONE2 neutralHow it really works#
What drives the bill#
| Signal | Billed or sized by | Grows with | Main lever |
|---|---|---|---|
| Metrics | Active series (and samples/s) | Label combinations | Cardinality (II.05), scrape interval |
| Logs | GB ingested, GB retained | Requests × bytes per line | Drop, sample, shorten retention |
| Traces | Spans ingested | Requests × spans per request | Sampling |
| Profiles | Processes × sample rate | Fleet size | Sampling frequency |
A useful rule of thumb when planning: telemetry that costs more than a few percent of the infrastructure it observes deserves a review. For GPU fleets the ratio is usually tiny — one GPU-hour buys a great deal of telemetry — which is an argument for more detail there, not less (V.06).
Sampling traces#
Head sampling — the SDK decides when the root span starts.
parentbased_traceidratioat 0.1 keeps a consistent 10% of traces, honouring the parent’s decision so traces are complete.- Cheap: dropped traces cost almost nothing.
- Blind: it decides before knowing whether the request will fail.
Tail sampling — a collector buffers every span until the trace is complete, then applies policies:
processors:
tail_sampling:
decision_wait: 10s
policies:
- { name: errors, type: status_code, status_code: { status_codes: [ERROR] } }
- { name: slow, type: latency, latency: { threshold_ms: 2000 } }
- { name: vip, type: string_attribute, string_attribute: { key: tenant.tier, values: [enterprise] } }
- { name: baseline, type: probabilistic, probabilistic: { sampling_percentage: 2 } }
- Keeps what you will actually look at.
- Costs memory, and every span of a trace must reach the same collector: put a
trace-ID-aware load balancer (the
loadbalancingexporter) in front of the sampling tier.
Keep your metrics exact. Compute RED metrics from all spans before sampling (the
Collector’s spanmetrics connector does this), or from the service’s own counters. Never
derive rates from sampled traces without multiplying by the inverse of the sampling
probability.
Reducing logs#
| Technique | Typical saving |
|---|---|
| Drop DEBUG and health-check lines at the agent | Large, free |
| One wide event instead of ten lines per request | Several-fold |
| Sample successes, keep all errors | 5–20× |
| Deduplicate repeated messages ("… repeated 4,000 times") | Spiky but large |
| Route by value: hot store for a week, object storage for a year | Cost per GB drops 10× or more |
| Convert a log pattern into a metric and drop the lines | Removes whole categories |
Reducing metrics#
Everything in II.05, plus: lengthen the scrape interval for slow-moving metrics (60 s for capacity, 15 s for serving), drop unused metrics at ingestion (most exporters expose hundreds you never query — list which metrics no dashboard or rule references), and aggregate away per-pod labels with recording rules where only the service total matters.
Aggregation and downsampling#
Old data rarely needs full resolution. Keep 15-second samples for two weeks and 5-minute and 1-hour aggregates for a year. Downsample by keeping min, max, sum and count for each interval — never an average alone, and never a percentile; keep the histogram.
Budgeting#
metrics: active_series × price_per_series
logs: requests/s × lines/request × bytes/line × 86,400 × retention_days × price_per_GB
traces: requests/s × spans/request × bytes/span × sample_rate × 86,400 × price_per_GBAttribute the bill to teams with a label, show each team its own number, and give each a budget. Cost becomes controllable only when the people who add the label see the price.
Code#
Head versus tail sampling on the same traffic: what each keeps, and what each costs.
// sampling.go — head vs tail sampling: storage cost and how many interesting traces survive.
package main
import (
"fmt"
"math/rand"
)
type Trace struct {
Err bool
Seconds float64
Spans int
}
func main() {
rng := rand.New(rand.NewSource(9))
traces := make([]Trace, 200000)
for i := range traces {
t := Trace{Seconds: 0.2 + rng.ExpFloat64()*0.3, Spans: 8 + rng.Intn(10)}
if rng.Float64() < 0.004 {
t.Err = true
}
if rng.Float64() < 0.01 {
t.Seconds += 3 + rng.Float64()*5
}
traces[i] = t
}
interesting := func(t Trace) bool { return t.Err || t.Seconds > 2 }
type result struct{ kept, keptSpans, keptInteresting, buffered int }
run := func(head float64, tail bool, baseline float64) result {
var r result
for _, t := range traces {
if rng.Float64() >= head {
continue // dropped at the source: costs nothing downstream
}
r.buffered += t.Spans
keep := true
if tail {
keep = interesting(t) || rng.Float64() < baseline
}
if keep {
r.kept++
r.keptSpans += t.Spans
if interesting(t) {
r.keptInteresting++
}
}
}
return r
}
total, totalSpans := 0, 0
for _, t := range traces {
totalSpans += t.Spans
if interesting(t) {
total++
}
}
fmt.Printf("%d traces, %d spans, %d interesting (errors or > 2 s)\n\n", len(traces), totalSpans, total)
fmt.Println("strategy spans stored interesting kept spans buffered in collector")
for _, s := range []struct {
name string
head float64
tail bool
baseline float64
}{
{"keep everything", 1, false, 0},
{"head 5%", 0.05, false, 0},
{"tail: errors+slow+2% baseline", 1, true, 0.02},
{"head 50% then tail", 0.5, true, 0.02},
} {
r := run(s.head, s.tail, s.baseline)
fmt.Printf("%-30s %6.1f%% %6.1f%% %6.1f%%\n", s.name,
100*float64(r.keptSpans)/float64(totalSpans),
100*float64(r.keptInteresting)/float64(total),
100*float64(r.buffered)/float64(totalSpans))
}
}
Remember this#
- Metrics cost per series, logs per byte, traces per span. Pull the matching lever.
- Head sampling is cheap and blind; tail sampling is selective and needs buffering.
- Compute metrics before sampling, and record sampling rates.
- Downsample with min, max, sum and count — and make each team see its own bill.
Try it#
- Run
sampling.go. Find the cheapest configuration that keeps at least 99% of interesting traces. - Estimate the daily log volume of a service you know, then apply three techniques from the table and recompute.
- List the metrics your service exposes. How many appear in a dashboard or an alert?
Check yourself#
- Why can head sampling not keep all errors?
- What infrastructure requirement does tail sampling add?
- Why should RED metrics be computed before traces are sampled?