The idea in one minute#
There are hundreds of observability products and a small number of jobs: instrument, collect, store, query, alert — and, for AI, read the accelerator, read the engine, trace the model call, score the output. Every tool does one or more of these jobs. Place a tool on that map and you know what it replaces and what it still needs beside it.
Choose in this order: open standards for instrumentation first (so nothing else is permanent), then storage by the shape of your questions, then the interface people will actually use at 3 a.m.
An analogy#
A kitchen. There are countless brands, and only so many jobs: cut, heat, cool, store, serve. You do not need to know every brand. You need to know which job each appliance does, and not to buy three ovens and no fridge.
A picture#
flowchart TB
subgraph INS["Instrument"]
I1["OpenTelemetry SDKs"]
I2["Prometheus client libraries"]
I3["eBPF: OBI, Beyla, Pixie"]
I4["Exporters: node, kube-state, DCGM"]
end
subgraph COLL["Collect and process"]
C1["OTel Collector, Grafana Alloy"]
C2["Prometheus scrape, vmagent"]
C3["Fluent Bit, Vector"]
end
subgraph STO["Store and query"]
S1["Metrics: Prometheus, Mimir, Thanos, VictoriaMetrics"]
S2["Logs: Loki, OpenSearch, ClickHouse"]
S3["Traces: Tempo, Jaeger"]
S4["Profiles: Pyroscope, Parca"]
end
subgraph USE["Use"]
U1["Grafana, vendor UIs"]
U2["Alertmanager, on-call tools"]
U3["SLO tools: Sloth, Pyrra"]
end
subgraph AI["AI-specific"]
A1["GPU: DCGM, vendor exporters"]
A2["Engine metrics: vLLM, SGLang, Dynamo"]
A3["LLM tracing and evals: Langfuse, Phoenix, LangSmith, ..."]
A4["Load and benchmark: GuideLLM, AIPerf, inference-perf"]
end
INS --> COLL --> STO --> USE
AI --> COLL
class I1,I2,I3,I4 compute
class C1,C2,C3 io
class S1,S2,S3,S4 memory
class U1,U2,U3 queue
class A1,A2,A3,A4 neutralHow it really works#
General-purpose stack, by job#
| Job | Open source | Notes |
|---|---|---|
| Instrument code | OpenTelemetry SDKs; Prometheus client libraries | OTel for traces and logs; either for metrics |
| Instrument without code | OBI (OpenTelemetry eBPF Instrumentation), Grafana Beyla, Pixie, Coroot | Breadth, protocol-level detail |
| Export system state | node_exporter, cAdvisor, kube-state-metrics, blackbox exporter | The standard Kubernetes set |
| Collect and process | OpenTelemetry Collector; Grafana Alloy (a Collector distribution); Fluent Bit; Vector | The Collector is the default hub |
| Metrics storage | Prometheus; Grafana Mimir, Thanos, Cortex; VictoriaMetrics | Single node → horizontally scaled with object storage |
| Log storage | Grafana Loki; OpenSearch / Elasticsearch; ClickHouse | Label-indexed vs full-text vs columnar |
| Trace storage | Grafana Tempo; Jaeger (v2 is built on the OTel Collector); ClickHouse | |
| Profile storage | Grafana Pyroscope; Parca | |
| All signals in one | SigNoz, OpenObserve, Uptrace, ClickStack (ClickHouse + HyperDX) | OTel-native, columnar back ends |
| Dashboards | Grafana; Perses | |
| Alert routing | Alertmanager; Grafana Alerting | Plus an on-call scheduler |
| SLOs | Sloth, Pyrra; OpenSLO as the spec | Generate burn-rate rules |
| Kubernetes packaging | kube-prometheus-stack, the Prometheus Operator, the OpenTelemetry Operator | Start here rather than assembling by hand |
| Energy | Kepler | Per-pod energy estimates |
Commercial platforms covering most of the table in one product include Datadog, Dynatrace, New Relic, Splunk, Honeycomb, Chronosphere and Grafana Cloud, and each major cloud has its own (CloudWatch, Azure Monitor, Google Cloud Observability, including managed Prometheus services).
AI infrastructure, by layer#
| Layer (V.01) | Tools | What you get |
|---|---|---|
| Accelerator | NVIDIA DCGM + dcgm-exporter (installed by the GPU Operator); go-nvml; nvidia-smi, nvtop, nvitop for ad-hoc use; AMD’s device metrics exporter; cloud-provider accelerator metrics | Power, energy, temperature, memory, errors, links |
| Node health | Node Problem Detector; NVIDIA NVSentinel; DRA device taints | Detect, cordon, drain, remediate |
| Device profiling | NVIDIA Nsight Systems and Nsight Compute; PyTorch profiler; eBPF-based GPU profilers | Kernel-level time on the device |
| Engine | Built-in Prometheus endpoints of vLLM, SGLang, NVIDIA Dynamo and Triton, TensorRT-LLM, llama.cpp server, KServe, Ray Serve | Queue, TTFT, tokens, KV cache |
| Router / gateway | Gateway API Inference Extension endpoint picker and llm-d; Envoy AI Gateway; LiteLLM; Kong and other API gateways with AI plugins | Per-tenant tokens, routing decisions |
| Platform | Kubernetes metrics; KEDA; Kueue and other batch schedulers’ metrics | Scaling, pending, quotas |
| LLM application tracing | OpenTelemetry GenAI instrumentations; OpenLLMetry (Traceloop); OpenInference (Arize) | Standard spans for model, tool and agent calls |
| LLM observability and evals | Langfuse; Arize Phoenix / Arize AX; LangSmith; Braintrust; Opik (Comet); W&B Weave; MLflow Tracing; Helicone; LLM modules of Datadog, Grafana, New Relic, Dynatrace, Honeycomb | Trace views built for prompts, token cost, datasets, evaluators |
| Load testing and benchmarks | GuideLLM; NVIDIA AIPerf (successor to GenAI-Perf); inference-perf (Kubernetes SIG); llm-d-benchmark; public references InferenceX and MLPerf Inference | Capacity and goodput numbers to compare your metrics against |
LLM-observability tools: how they differ#
They all show a trace of model calls with tokens and cost. The differences that matter:
| Question | Why it matters |
|---|---|
| Does it ingest plain OTLP with the GenAI conventions? | Otherwise you are locked to its SDK |
| Open source and self-hostable? Under which licence? | Prompts are sensitive; many teams must keep them in-house. Langfuse’s core is MIT; Phoenix uses Elastic License 2.0; others are hosted-only |
| Is evaluation built in (datasets, judges, human review queues)? | Tracing without scoring is half the job (V.07) |
| Does it join to infrastructure telemetry? | Otherwise an LLM trace and a GPU metric live in different worlds |
| Is it tied to one framework? | LangSmith is deepest with LangChain/LangGraph; that is a strength only if you use them |
| What happens at volume? | Many were built for development-time tracing, not millions of requests per hour |
The category is consolidating into the general platforms: ClickHouse acquired Langfuse in January 2026 (the open-source licence and self-hosting were kept), and every large observability vendor now ships an LLM module. Expect “LLM observability” to become a feature of your observability platform rather than a separate product — which is one more reason to emit standard OTel.
How to choose#
- Instrument with OpenTelemetry and Prometheus exposition. This decision outlasts all the others.
- Start from what your platform installs. On Kubernetes with GPUs:
kube-prometheus-stack, the GPU Operator’s dcgm-exporter, your engine’s/metrics. That is a working system in an afternoon. - Pick storage by question shape. Known questions and alerts → a Prometheus-compatible TSDB. Unknown questions sliced by many fields → a column store with wide events.
- Decide where prompts may live before choosing an LLM tracing tool.
- Count operating cost honestly. Self-hosting is free until it pages you.
- Run a real incident drill on the candidate before committing. The test is how fast a tired engineer gets from the page to the cause.
A sensible default, small to large#
| Scale | Stack |
|---|---|
| One team, one cluster | Prometheus + Grafana + Alertmanager; Loki; dcgm-exporter; engine metrics; an OTel Collector; one LLM tracing tool if you build applications |
| Several clusters | Add remote write to Mimir / Thanos / VictoriaMetrics or a managed Prometheus; Tempo with tail sampling; continuous profiling |
| Large fleet | Collector gateways with per-tenant limits; a column store for wide events; SLO tooling; automated node remediation; showback |
Code#
A tool is a set of jobs. This program checks a proposed stack for gaps and overlaps — the question to ask of any architecture diagram.
// stackcheck.go — does this stack cover every job, and where does it overlap?
package main
import (
"fmt"
"sort"
"strings"
)
func main() {
jobs := []string{
"instrument", "collect", "metrics store", "log store", "trace store", "profiles",
"dashboards", "alert routing", "gpu telemetry", "engine metrics", "llm tracing", "evals", "load test",
}
stack := map[string][]string{
"OpenTelemetry SDK + Collector": {"instrument", "collect"},
"Prometheus": {"collect", "metrics store"},
"Grafana": {"dashboards"},
"Alertmanager": {"alert routing"},
"Loki": {"log store"},
"Tempo": {"trace store"},
"dcgm-exporter": {"gpu telemetry"},
"vLLM /metrics": {"engine metrics"},
"Langfuse": {"llm tracing", "evals", "trace store"},
}
covered := map[string][]string{}
for tool, js := range stack {
for _, j := range js {
covered[j] = append(covered[j], tool)
}
}
fmt.Println("job covered by")
var gaps []string
for _, j := range jobs {
tools := covered[j]
sort.Strings(tools)
note := ""
switch {
case len(tools) == 0:
note = "← GAP"
gaps = append(gaps, j)
case len(tools) > 1:
note = "← overlap: decide which is the source of truth"
}
fmt.Printf("%-15s %-48s %s\n", j, strings.Join(tools, ", "), note)
}
fmt.Printf("\n%d gaps: %s\n", len(gaps), strings.Join(gaps, ", "))
}Remember this#
- Tools are implementations of a few jobs. Map a tool to its jobs before comparing it.
- Open instrumentation first; it keeps every later choice reversible.
- For AI: accelerator exporter + engine metrics + gateway accounting + standard GenAI traces + an evaluator. LLM observability is becoming a feature of general platforms.
- Decide where prompts are allowed to be stored before you pick a tracing tool.
Try it#
- Run
stackcheck.gowith the stack you actually use. What are the gaps? - For one LLM observability tool, answer the six questions in the table from its documentation.
- Design the smallest stack that covers every job for a single-node inference server.
Check yourself#
- Which decision in a stack is the hardest to reverse, and how do you make it reversible?
- What distinguishes a column store from a time-series database as a home for telemetry?
- Name one tool per AI layer.