Below the API

Glossary

Short definitions of every term the course uses. The lesson in brackets is where it is explained.

TermMeaning
AlertmanagerPrometheus component that deduplicates, groups, silences and routes alerts. (IV.01)
BaggageKey-value pairs propagated with a request alongside the trace context. (III.02)
Burn rateObserved error ratio divided by the error budget; how fast the budget is being spent. (I.03, IV.03)
CardinalityThe number of distinct time series; the product of label value counts. (II.05)
CollectorOpenTelemetry’s standalone pipeline process: receivers, processors, exporters. (III.03)
Context propagationPassing trace and span IDs from one component to the next. (III.02)
Coordinated omissionA load-testing error where slow responses delay later requests, hiding stalls. (I.04)
CounterA metric that only increases; queried with rate. (II.01)
DCGMNVIDIA’s Data Center GPU Manager; source of fleet GPU telemetry. (V.02)
Dead man’s switchAn always-firing alert whose absence means the alerting pipeline is broken. (IV.01)
DRADynamic Resource Allocation: Kubernetes API for requesting devices by attribute. (V.05)
eBPFVerified programs run inside the Linux kernel; used for profiling and zero-code instrumentation. (III.04)
Error budget1 − SLO: the fraction of events allowed to be bad. (I.03)
ExemplarA trace ID stored with a metric sample, linking a spike to a request. (I.02, III.02)
Finish reasonWhy generation stopped: stop, length, tool_calls, … (V.04, V.07)
Flame graphA profile drawn so that width is share of samples. (III.04)
GaugeA metric that can go up and down. (II.01)
GenAI semantic conventionsOpenTelemetry’s gen_ai.* names for model, tool and agent telemetry. (V.04)
GoodputThe rate of requests (or tokens) that met every SLO. (I.03, V.03)
Head samplingDeciding whether to keep a trace when it starts. (III.02, IV.04)
HistogramBucketed counts of observations, from which percentiles are estimated. (II.04)
ITLInter-token latency: the gap between consecutive output tokens. (V.03)
KV-cache usageFraction of an engine’s key-value cache blocks in use; the real memory pressure. (V.03)
LabelA key-value pair identifying one series of a metric. (II.01)
MCPModel Context Protocol: a standard for exposing tools and data to models. (V.04, VI.02)
Native histogramPrometheus’s exponential-bucket histogram stored as one series. (II.04)
NVMLNVIDIA Management Library; what nvidia-smi reads. (V.02)
OBIOpenTelemetry eBPF Instrumentation. (III.04)
OccupancyTokens served divided by tokens a replica could serve at its SLO. (V.06)
OTLPThe OpenTelemetry protocol for all signals. (III.03)
PercentileThe value below which a given share of measurements fall. (I.04)
PreemptionAn engine evicting a running request to free KV-cache space. (V.03)
Prefix-cache hit rateCached prompt tokens divided by prompt tokens. (V.03)
ProfileAggregated stack samples showing where a resource is spent. (III.04)
PromQLThe Prometheus query language. (II.03)
PUEPower usage effectiveness: facility power divided by IT power. (V.06)
Recording ruleA query evaluated on a schedule and stored as a new series. (II.03)
REDRate, Errors, Duration: the metrics for a service. (II.02)
Resource attributesAttributes describing the source of telemetry (service.name, pod, GPU). (III.03)
SaturationWork waiting for a resource. (II.02)
SLI / SLO / SLAIndicator (a good/total ratio) / objective (a target for it) / agreement (a contract). (I.03)
SpanOne timed operation in a trace. (III.02)
Tail samplingDeciding whether to keep a trace after it has finished. (IV.04)
TemporalityWhether a metric is reported as a running total (cumulative) or a change (delta). (III.03)
Time seriesTimestamped values identified by a metric name and labels. (II.01)
TPOTTime per output token, averaged over a request. (V.03)
TraceThe tree of spans for one request. (III.02)
traceparentThe W3C header carrying trace ID, parent span ID and flags. (III.02)
TTFTTime to first token. (V.03)
USEUtilization, Saturation, Errors: the metrics for a resource. (II.02)
Wide eventOne record per unit of work carrying every known field. (III.01)
XidA numbered NVIDIA driver error report in the kernel log. (V.02)

↑↓ navigate ↵ open