Short definitions of every term the course uses. The lesson in brackets is where it is explained.
| Term | Meaning |
|---|---|
| Alertmanager | Prometheus component that deduplicates, groups, silences and routes alerts. (IV.01) |
| Baggage | Key-value pairs propagated with a request alongside the trace context. (III.02) |
| Burn rate | Observed error ratio divided by the error budget; how fast the budget is being spent. (I.03, IV.03) |
| Cardinality | The number of distinct time series; the product of label value counts. (II.05) |
| Collector | OpenTelemetry’s standalone pipeline process: receivers, processors, exporters. (III.03) |
| Context propagation | Passing trace and span IDs from one component to the next. (III.02) |
| Coordinated omission | A load-testing error where slow responses delay later requests, hiding stalls. (I.04) |
| Counter | A metric that only increases; queried with rate. (II.01) |
| DCGM | NVIDIA’s Data Center GPU Manager; source of fleet GPU telemetry. (V.02) |
| Dead man’s switch | An always-firing alert whose absence means the alerting pipeline is broken. (IV.01) |
| DRA | Dynamic Resource Allocation: Kubernetes API for requesting devices by attribute. (V.05) |
| eBPF | Verified programs run inside the Linux kernel; used for profiling and zero-code instrumentation. (III.04) |
| Error budget | 1 − SLO: the fraction of events allowed to be bad. (I.03) |
| Exemplar | A trace ID stored with a metric sample, linking a spike to a request. (I.02, III.02) |
| Finish reason | Why generation stopped: stop, length, tool_calls, … (V.04, V.07) |
| Flame graph | A profile drawn so that width is share of samples. (III.04) |
| Gauge | A metric that can go up and down. (II.01) |
| GenAI semantic conventions | OpenTelemetry’s gen_ai.* names for model, tool and agent telemetry. (V.04) |
| Goodput | The rate of requests (or tokens) that met every SLO. (I.03, V.03) |
| Head sampling | Deciding whether to keep a trace when it starts. (III.02, IV.04) |
| Histogram | Bucketed counts of observations, from which percentiles are estimated. (II.04) |
| ITL | Inter-token latency: the gap between consecutive output tokens. (V.03) |
| KV-cache usage | Fraction of an engine’s key-value cache blocks in use; the real memory pressure. (V.03) |
| Label | A key-value pair identifying one series of a metric. (II.01) |
| MCP | Model Context Protocol: a standard for exposing tools and data to models. (V.04, VI.02) |
| Native histogram | Prometheus’s exponential-bucket histogram stored as one series. (II.04) |
| NVML | NVIDIA Management Library; what nvidia-smi reads. (V.02) |
| OBI | OpenTelemetry eBPF Instrumentation. (III.04) |
| Occupancy | Tokens served divided by tokens a replica could serve at its SLO. (V.06) |
| OTLP | The OpenTelemetry protocol for all signals. (III.03) |
| Percentile | The value below which a given share of measurements fall. (I.04) |
| Preemption | An engine evicting a running request to free KV-cache space. (V.03) |
| Prefix-cache hit rate | Cached prompt tokens divided by prompt tokens. (V.03) |
| Profile | Aggregated stack samples showing where a resource is spent. (III.04) |
| PromQL | The Prometheus query language. (II.03) |
| PUE | Power usage effectiveness: facility power divided by IT power. (V.06) |
| Recording rule | A query evaluated on a schedule and stored as a new series. (II.03) |
| RED | Rate, Errors, Duration: the metrics for a service. (II.02) |
| Resource attributes | Attributes describing the source of telemetry (service.name, pod, GPU). (III.03) |
| Saturation | Work waiting for a resource. (II.02) |
| SLI / SLO / SLA | Indicator (a good/total ratio) / objective (a target for it) / agreement (a contract). (I.03) |
| Span | One timed operation in a trace. (III.02) |
| Tail sampling | Deciding whether to keep a trace after it has finished. (IV.04) |
| Temporality | Whether a metric is reported as a running total (cumulative) or a change (delta). (III.03) |
| Time series | Timestamped values identified by a metric name and labels. (II.01) |
| TPOT | Time per output token, averaged over a request. (V.03) |
| Trace | The tree of spans for one request. (III.02) |
traceparent | The W3C header carrying trace ID, parent span ID and flags. (III.02) |
| TTFT | Time to first token. (V.03) |
| USE | Utilization, Saturation, Errors: the metrics for a resource. (II.02) |
| Wide event | One record per unit of work carrying every known field. (III.01) |
| Xid | A numbered NVIDIA driver error report in the kernel log. (V.02) |