The idea in one minute#
Telemetry comes in a small number of shapes, called signals. Metrics are numbers sampled over time: cheap, good for trends and alerts. Logs are timestamped records of individual things that happened. Traces follow one request across components and show where its time went. Profiles show which lines of code consumed CPU or memory.
They differ in cost per request and in how much detail survives. Metrics throw away the individual request to stay cheap; the other three keep it and pay for that.
An analogy#
A shop can keep three kinds of records. The till total at closing time (a metric): tiny, tells you how the day went, cannot tell you what Mrs. Shah bought. The receipts (logs): one per sale, detailed, a shoebox full by Friday. A customer’s path through the shop on the security camera (a trace): shows she waited four minutes at the till, which no receipt records. And a time-and-motion study of the staff (a profile): where the working hours actually go.
A picture#
flowchart TB REQ["One request"] --> M["Metric<br/>+1 to a counter, one bucket of a histogram"] REQ --> L["Log<br/>a record with fields"] REQ --> T["Trace<br/>a tree of timed spans"] CPU["The process itself"] --> P["Profile<br/>stack samples, 100 per second"] M --> MQ["Is the service healthy?<br/>Is it getting worse?"] L --> LQ["What exactly happened<br/>to this request?"] T --> TQ["Where did the time go,<br/>across services?"] P --> PQ["Which code is<br/>burning the CPU?"] class REQ,CPU neutral class M,L,T,P memory class MQ,LQ,TQ,PQ queue
How it really works#
The four signals side by side#
| Signal | One data point is | Cost grows with | Keeps the individual request? | Best at |
|---|---|---|---|---|
| Metric | A number for a (name, labels) pair at a time | Number of distinct label combinations | No | Alerting, trends, capacity |
| Log | A timestamped record | Requests × bytes per record | Yes | Detail about one event; audit |
| Trace | A tree of spans sharing a trace ID | Requests × spans (reduced by sampling) | Yes | Latency breakdown across components |
| Profile | Aggregated stack samples | Processes × sampling rate | No (it is about code, not requests) | Finding expensive code |
The key property of a metric: its cost does not grow with traffic. A counter incremented a million times per second is still one number. That is why metrics are the foundation of alerting, and also why they cannot tell you about any single request.
Events: the shape underneath#
Logs and spans are both events: a set of key-value fields describing one thing that happened. A span is an event with a duration, a parent and a trace ID. A “wide event” is a single record per request carrying every field you know about it — dozens to hundreds. Metrics can be derived from events (count them, bucket their durations); the reverse is impossible. This is why a growing part of the industry stores wide events in column databases and computes metrics on query (VI.02).
Linking the signals#
Signals are far more useful joined than separate:
| Link | Mechanism | What it buys |
|---|---|---|
| Log ↔ trace | Put trace_id and span_id in every log record | From an error log, open the whole request |
| Metric → trace | Exemplars: a histogram bucket remembers one trace ID that landed in it | From a latency spike, open a slow request |
| Trace → profile | Attach span IDs to profile samples | From a slow span, see the code that ran |
| All ↔ all | Shared resource attributes: service.name, version, host, pod, GPU | Slice everything by the same dimensions |
OpenTelemetry (III.03) exists largely to make these links standard.
Other things you will hear called signals#
- Events in the Kubernetes sense: state changes such as a pod being evicted.
- Health checks and synthetic probes: requests you send yourself, to measure from outside.
- Real-user monitoring: timings recorded in the client.
- For AI systems, evaluation results — scores for output quality — behave like a signal too (V.07).
Choosing#
Ask what question you need to answer and how often:
"Is it broken?" every 30 s, forever → metric
"What happened to request X?" occasionally → log or trace
"Where does time go between services?" → trace
"Why is this process using 8 cores?" → profileStart with metrics for the symptoms users feel, add traces at the boundaries between components, keep logs structured, and turn profiling on continuously once you operate anything expensive — a GPU server qualifies.
Code#
One simulated request, emitted as each of the three request-scoped signals, so you can compare the bytes each costs.
// signals.go — the same request as a metric update, a log record and a trace.
package main
import (
"encoding/json"
"fmt"
"time"
)
type Span struct {
TraceID string `json:"trace_id"`
SpanID string `json:"span_id"`
ParentID string `json:"parent_id,omitempty"`
Name string `json:"name"`
Millis float64 `json:"duration_ms"`
}
func main() {
start := time.Date(2026, 10, 3, 9, 0, 0, 0, time.UTC)
queue, prefill, decode := 12.0, 85.0, 940.0
total := queue + prefill + decode
// 1. Metric: nothing about this request survives except +1 in one bucket.
fmt.Println("METRIC (what changes in memory)")
fmt.Println(` requests_total{route="/v1/chat",code="200"} += 1`)
fmt.Println(` request_seconds_bucket{route="/v1/chat",le="2.5"} += 1`)
fmt.Println(" bytes sent per request: 0 (scraped as two numbers, whatever the traffic)")
// 2. Log: one record, all context, no structure across components.
logRec := map[string]any{
"ts": start.Format(time.RFC3339Nano), "level": "info", "msg": "request done",
"route": "/v1/chat", "tenant": "acme", "model": "large", "status": 200,
"prompt_tokens": 1830, "output_tokens": 212, "duration_ms": total,
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
}
lb, _ := json.Marshal(logRec)
fmt.Printf("\nLOG (%d bytes)\n %s\n", len(lb), lb)
// 3. Trace: the same request as a tree of timed spans.
tid := "4bf92f3577b34da6a3ce929d0e0e4736"
spans := []Span{
{tid, "a1", "", "POST /v1/chat", total},
{tid, "b1", "a1", "queue", queue},
{tid, "b2", "a1", "prefill", prefill},
{tid, "b3", "a1", "decode", decode},
}
tb, _ := json.Marshal(spans)
fmt.Printf("\nTRACE (%d bytes, %d spans)\n", len(tb), len(spans))
for _, s := range spans {
indent := ""
if s.ParentID != "" {
indent = " "
}
fmt.Printf(" %s%-16s %7.1f ms\n", indent, s.Name, s.Millis)
}
fmt.Println("\nAt 1,000 requests/s: the metric still costs two numbers per scrape;")
fmt.Printf("logs cost ~%d kB/s and traces ~%d kB/s before sampling.\n", len(lb), len(tb))
}
Remember this#
- Metrics are cheap because they forget the individual request. Logs and traces are detailed because they do not.
- Logs and spans are both events; metrics can be derived from events, never the reverse.
- The value is in the links: trace IDs in logs, exemplars on histograms, shared resource attributes.
- Pick the signal from the question, and the question’s frequency.
Try it#
- Run
signals.go. Add ten more fields to the log record. How does cost at 1,000 requests/s change? What would 1% sampling do, and what would you lose? - From the four spans, compute what a histogram of
decodeduration would need to store. Which questions can the histogram no longer answer? - Take one log line from a real system you use. Rewrite it as a structured event with at least eight fields.
Check yourself#
- Why does a metric’s cost not grow with traffic?
- What is an exemplar and which two signals does it connect?
- Which signal would you reach for to learn why a process uses twice the CPU after a deploy?