Below the API

Cardinality

Basic Intermediate 40 min Difficulty 2/5

Prerequisites 01, 04

The idea in one minute#

Cardinality is the number of distinct time series. It is the product of the number of values of every label on a metric — and products grow fast. Each series costs memory in the database whether or not anyone ever queries it, so cardinality is what a metrics system’s cost and stability depend on, far more than request volume.

The rule: labels are for bounded sets you aggregate over. Anything unbounded — user IDs, request IDs, prompts, raw URLs — belongs in logs and traces, not in a label.

An analogy#

A filing cabinet with one folder per (customer, product, branch, day). Add “per salesperson” and the number of folders multiplies by the number of salespeople, most of them holding one sheet of paper. The paper is cheap. The folders are what fill the room.

A picture#

flowchart TB
  M["request_duration_seconds"] --> L1["route<br/>40 values"]
  L1 --> L2["code<br/>8 values"]
  L2 --> L3["instance<br/>30 values"]
  L3 --> L4["le buckets<br/>12 + 3"]
  L4 --> N["40 x 8 x 30 x 15<br/>= 144,000 series"]
  N --> ADD["add tenant: x 2,000"]
  ADD --> BOOM["288,000,000 series<br/>the database falls over"]
  class M neutral
  class L1,L2,L3,L4 memory
  class N queue
  class ADD,BOOM warn

How it really works#

The arithmetic#

series for one metric = product of (distinct values of each label)
                        × (buckets + 3, for a classic histogram)

In practice not every combination occurs, so the real number is lower than the product — but the product is the ceiling you must plan for.

What a series costs: a few kilobytes of memory in Prometheus while it is active, an index entry, and a line on the invoice if your vendor bills per active series. A healthy single Prometheus handles millions of active series; trouble begins when one deploy adds tens of millions.

Churn#

A series that stops receiving samples is still indexed for the retention period. Labels whose values change with every deploy — pod name, container ID, a version hash — create churn: the active count looks fine while the total climbs. In Kubernetes this is normal and manageable; it becomes a problem when multiplied by another high-cardinality label.

Finding the offenders#

topk(10, count by (__name__) ({__name__=~".+"}))     # which metrics have the most series
count(count by (tenant) (http_requests_total))       # how many values does `tenant` have?

Prometheus’s status page (TSDB status) lists the highest-cardinality metrics and labels.

Fixes, in order of preference#

FixHow
Do not add the labelPut the detail in a trace or a structured log instead
Bucket the valueprompt_tokens → size="small|medium|large"; status code → class
NormalizeRoute template instead of path; strip IDs
Bound itKeep the top N tenants by name, the rest as tenant="other"
Drop at ingestionmetric_relabel_configs in Prometheus; a processor in the OTel Collector
Use a native histogramRemoves the bucket multiplier (lesson 04)
Aggregate earlyA recording rule without the noisy label, then drop the raw series

When you genuinely need per-entity detail#

Per-tenant or per-user questions are legitimate. The honest options:

  1. Exemplars: keep the metric low-cardinality and attach a trace ID to it.
  2. Wide events in a column store: one row per request, any field, aggregated at query time (III.01, VI.02). Cost scales with requests, not with distinct values — the opposite trade-off from metrics.
  3. A separate, deliberately bounded metric: per-tenant token counters for the 200 tenants you bill, with nothing else on it.

In AI infrastructure#

The tempting labels are worse than usual:

Tempting labelValuesDo instead
user_id, session_id, conversation_idUnboundedSpan attribute
prompt or a prompt hashUnboundedNever a label; opt-in content on spans
modelTens — fineKeep
lora_adapterCan be thousandsTop-N plus other
gpu UUID × pod × containerBounded but large on a big fleet, high churnFine for hardware metrics; do not multiply with request labels
finish_reason, error_typeA handful — fineKeep

Code#

// cardinality.go — how many series will this metric create?
package main

import "fmt"

type Label struct {
	Name   string
	Values int
}

func series(labels []Label, bucketSeries int) int {
	n := 1
	for _, l := range labels {
		n *= l.Values
	}
	return n * bucketSeries
}

func main() {
	const bytesPerSeries = 4096.0 // rough in-memory cost of an active series
	base := []Label{{"route", 40}, {"code", 8}, {"instance", 30}}

	cases := []struct {
		name    string
		labels  []Label
		buckets int
	}{
		{"counter", base, 1},
		{"classic histogram, 12 buckets", base, 15},
		{"native histogram", base, 1},
		{"classic + model (20)", append(append([]Label{}, base...), Label{"model", 20}), 15},
		{"classic + tenant (2,000)", append(append([]Label{}, base...), Label{"tenant", 2000}), 15},
		{"classic + top-50 tenants", append(append([]Label{}, base...), Label{"tenant", 51}), 15},
		{"native + top-50 tenants", append(append([]Label{}, base...), Label{"tenant", 51}), 1},
	}
	fmt.Println("metric shape                         max series        memory")
	for _, c := range cases {
		n := series(c.labels, c.buckets)
		fmt.Printf("%-34s %12d   %9.1f GB\n", c.name, n, float64(n)*bytesPerSeries/1e9)
	}
	fmt.Println("\nThe ceiling is a product. Remove a factor, or shrink one, before it ships.")
}

Remember this#

  • Cardinality = product of label value counts (× buckets). Cost follows series, not traffic.
  • Labels are for small, bounded sets. Unbounded detail goes in traces, logs or wide events.
  • Bucket, normalize, cap at top-N, or drop at ingestion.
  • Native histograms remove the bucket multiplier.

Try it#

  1. Run cardinality.go with your own service’s labels. What is the ceiling?
  2. Design a per-tenant token metric for 5,000 tenants that stays under 20,000 series.
  3. For a metric you own, decide which one label you would remove if you had to halve its cost.

Check yourself#

  1. How do you compute the worst-case series count of a metric?
  2. What is churn, and which labels cause it?
  3. Where should a request ID go, if not in a label?

↑↓ navigate ↵ open