PidokuInfra

OpenTelemetry

Basic Intermediate 1h Difficulty 3/5 Topic 03 of 04

Prerequisites II.01, 01, 02

The idea in one minute#

OpenTelemetry (OTel) is the vendor-neutral standard for producing and moving telemetry. It is four things: an API and SDK per language for creating metrics, logs, traces and profiles; a wire protocol, OTLP; semantic conventions, which fix the names (http.request.method, service.name) so data from different sources lines up; and the Collector, a standalone process that receives, transforms and forwards telemetry.

The point is separation: your code is instrumented once, against a standard, and where the data goes is a configuration decision you can change without touching the code.

An analogy#

Shipping containers. Before them, every port and ship handled cargo its own way. A standard box did not make ships faster; it meant any ship, crane and truck could handle any cargo. OTLP is the box, the Collector is the port, and backends are the destinations.

A picture#

flowchart TB
  subgraph APP["Your process"]
    API["OTel API<br/>what your code calls"] --> SDK["OTel SDK<br/>sampling, batching, export"]
    AUTO["Auto-instrumentation<br/>HTTP, gRPC, SQL libraries"] --> API
  end
  SDK -->|"OTLP"| AG["Collector as agent<br/>on every node"]
  PROM["Prometheus endpoints<br/>/metrics"] -->|"scrape"| AG
  EBPF["eBPF instrumentation<br/>no code changes"] -->|"OTLP"| AG
  AG -->|"OTLP"| GWC["Collector as gateway<br/>tail sampling, redaction, routing"]
  GWC --> B1[("Metrics backend")]
  GWC --> B2[("Trace backend")]
  GWC --> B3[("Log backend")]
  class API,SDK,AUTO compute
  class AG,GWC io
  class PROM,EBPF neutral
  class B1,B2,B3 memory

How it really works#

The pieces#

PieceWhat it isWhy it is separate
APIInterfaces: Tracer, Meter, LoggerLibraries can instrument themselves without forcing an implementation on you; with no SDK installed, calls are no-ops
SDKThe implementation: sampling, batching, exportersThe application owner configures it
Instrumentation librariesReady-made spans and metrics for common frameworksYou should not hand-write HTTP server spans
OTLPThe protocol (gRPC on 4317, HTTP on 4318)One format for every signal, accepted by nearly every backend
Semantic conventionsStandard attribute and metric namesDashboards and queries work across services and vendors
ResourceAttributes describing the source: service.name, service.version, k8s.pod.name, host.nameAttached to everything a process emits; this is what joins the signals
CollectorA pipeline process: receivers → processors → exportersKeeps vendor logic, credentials and heavy processing out of the app

The Collector#

A pipeline is declared in YAML:

YAML
receivers:
  otlp:
    protocols: { grpc: {}, http: {} }
  prometheus:
    config:
      scrape_configs:
        - job_name: vllm
          static_configs: [{ targets: ["vllm:8000"] }]
processors:
  memory_limiter: { check_interval: 1s, limit_percentage: 80 }
  k8sattributes: {}            # add pod, namespace, node to every record
  batch: {}
exporters:
  otlphttp/traces: { endpoint: https://traces.example.internal }
  prometheusremotewrite: { endpoint: https://metrics.example.internal/api/v1/write }
service:
  pipelines:
    traces:  { receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [otlphttp/traces] }
    metrics: { receivers: [otlp, prometheus], processors: [memory_limiter, k8sattributes, batch], exporters: [prometheusremotewrite] }

Two deployment roles, usually combined:

  • Agent — one per node (a DaemonSet) or per pod (a sidecar). Close to the app: cheap export, adds local metadata, scrapes local endpoints.
  • Gateway — a central, horizontally-scaled pool. Does what needs a wider view or credentials: tail sampling, redaction, routing to several backends.

Useful processors to know by name: batch, memory_limiter, k8sattributes, resourcedetection, filter, transform (rewrite with the OTTL language), tail_sampling, redaction.

In Go#

Go
// One-time setup
exp, _ := otlptracegrpc.New(ctx)
tp := sdktrace.NewTracerProvider(
    sdktrace.WithBatcher(exp),
    sdktrace.WithResource(resource.NewSchemaless(
        semconv.ServiceName("gateway"), semconv.ServiceVersion("1.8.2"))),
)
otel.SetTracerProvider(tp)
otel.SetTextMapPropagator(propagation.TraceContext{})

// Automatic server and client spans
handler := otelhttp.NewHandler(mux, "gateway")
client := &http.Client{Transport: otelhttp.NewTransport(http.DefaultTransport)}

// A manual span where it matters
ctx, span := otel.Tracer("gateway").Start(ctx, "pick_replica")
span.SetAttributes(attribute.String("model", model))
defer span.End()

The SDK reads standard environment variables, so most configuration needs no code: OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_TRACES_SAMPLER, OTEL_RESOURCE_ATTRIBUTES.

Where each part stands (checked 3 October 2026)#

AreaStatusWhat it means for you
Traces, metrics, logs: API, SDK, OTLPStableSafe to depend on
CollectorCore components stable; many contrib components are alpha or betaCheck each component’s stated stability
Declarative configuration (one YAML file for the SDK)Declared stable in 2026Replaces a pile of environment variables
Profiles signalPublic alpha since March 2026The fourth signal; usable, the format may still change (lesson 04)
eBPF instrumentation (OBI)Beta, announced at KubeCon EU in April 2026; 1.0 is the project’s stated 2026 goalZero-code traces and RED metrics for HTTP, gRPC and SQL
HTTP, database, messaging semantic conventionsStable for the core HTTP set; others at varying stagesNames are settled where marked stable
GenAI semantic conventionsEvery attribute still marked Development; moved to their own repository in June 2026Use them, and expect renames (V.04)

Stability is tracked per signal and per component, not for “OpenTelemetry” as a whole. Before depending on anything, read its stability marker.

OTel and Prometheus together#

They are no longer alternatives:

  • Prometheus 3 has a native OTLP receiver (/api/v1/otlp/v1/metrics) and accepts the dotted, UTF-8 names OTel uses.
  • OTLP exponential histograms become Prometheus native histograms (II.04).
  • The Collector scrapes Prometheus endpoints and writes Prometheus remote-write.
  • One real difference to know: temporality. Prometheus counters are cumulative (total since start); OTLP can also send delta (change since last export). Prometheus wants cumulative; configure the SDK or convert in the Collector.

A common, sensible shape: Prometheus exposition for infrastructure and anything that already has an exporter, OTel SDKs for application traces and logs, and a Collector in the middle joining both with the same resource attributes.

Code#

OTLP is just structured data. This program builds a minimal OTLP/JSON trace payload by hand so you can see exactly what an SDK sends.

Go
// otlp.go — what an OTLP trace export looks like on the wire (JSON encoding).
package main

import (
	"encoding/json"
	"fmt"
	"os"
	"time"
)

type KV struct {
	Key   string         `json:"key"`
	Value map[string]any `json:"value"`
}

func str(k, v string) KV         { return KV{k, map[string]any{"stringValue": v}} }
func integer(k string, v int) KV { return KV{k, map[string]any{"intValue": fmt.Sprint(v)}} }

type Span struct {
	TraceID      string `json:"traceId"`
	SpanID       string `json:"spanId"`
	ParentSpanID string `json:"parentSpanId,omitempty"`
	Name         string `json:"name"`
	Kind         int    `json:"kind"` // 2 = SERVER, 1 = INTERNAL
	Start        string `json:"startTimeUnixNano"`
	End          string `json:"endTimeUnixNano"`
	Attributes   []KV   `json:"attributes"`
}

func main() {
	t0 := time.Date(2026, 10, 3, 9, 0, 0, 0, time.UTC)
	ns := func(d time.Duration) string { return fmt.Sprint(t0.Add(d).UnixNano()) }
	tid := "4bf92f3577b34da6a3ce929d0e0e4736"

	payload := map[string]any{
		"resourceSpans": []any{map[string]any{
			// The resource: attached once, applies to every span below.
			"resource": map[string]any{"attributes": []KV{
				str("service.name", "inference-gateway"),
				str("service.version", "1.8.2"),
				str("k8s.pod.name", "gateway-7c9f-x2k"),
			}},
			"scopeSpans": []any{map[string]any{
				"scope": map[string]any{"name": "gateway"},
				"spans": []Span{
					{tid, "a1b2c3d4e5f60718", "", "POST /v1/chat/completions", 2, ns(0), ns(1042 * time.Millisecond),
						[]KV{str("http.request.method", "POST"), integer("http.response.status_code", 200)}},
					{tid, "00f067aa0ba902b7", "a1b2c3d4e5f60718", "pick_replica", 1, ns(2 * time.Millisecond), ns(3 * time.Millisecond),
						[]KV{str("replica", "vllm-3")}},
				},
			}},
		}},
	}
	enc := json.NewEncoder(os.Stdout)
	enc.SetIndent("", "  ")
	enc.Encode(payload)
	fmt.Fprintln(os.Stderr, "\nPOST this to http://localhost:4318/v1/traces on any OTel Collector.")
}

Remember this#

  • OTel = API/SDK + OTLP + semantic conventions + Collector.
  • Instrument against the API; choose destinations in configuration.
  • Resource attributes are what join metrics, logs, traces and profiles.
  • Stability is per signal and per component. Check the marker.
  • Prometheus and OTel interoperate; mind cumulative versus delta temporality.

Try it#

  1. Run otlp.go. Start a Collector with the debug exporter and POST the output to it.
  2. Write a Collector config that receives OTLP, drops spans from health-check routes, and exports to two backends.
  3. Instrument lesson II.02’s server with otelhttp and compare the spans with your hand-built ones from lesson 02.

Check yourself#

  1. Why is the API separate from the SDK?
  2. What are the two Collector deployment roles, and what belongs in each?
  3. Name two OTel areas that were not yet stable in October 2026.

↑↓ navigate↵ openesc close