PidokuInfra

Observability for LLM Services

Advanced Intermediate 1h Difficulty 3/5 Topic 08 of 11

Prerequisites X.10, 01

Section X.10 covered the metrics catalogue and dashboards. This file covers what’s specific to LLM services and what a generic APM setup will miss.


1. What generic observability misses#

A standard APM setup gives you:
  ✓ request rate, error rate, latency percentiles
  ✓ CPU, memory, network
  ✓ traces across services

It does NOT give you:
  ✗ KV cache utilization        ← your actual capacity limit
  ✗ running batch size          ← your actual efficiency
  ✗ queue wait vs prefill time  ← the TTFT decomposition
  ✗ prefix cache hit rate
  ✗ preemption rate
  ✗ tokens generated after abort
  ✗ per-request token counts and cost
  ✗ output quality signals
  ✗ GPU-specific health (Xid, ECC, throttling)

Every item in the second list is essential and none of it comes for free. Instrumenting them is a deliberate project.


2. The three pillars, applied#

METRICS   aggregate, cheap, alertable
          → "is the service healthy right now?"
          → the catalogue in Section X.10

TRACES    per-request, sampled, detailed
          → "why was THIS request slow?"
          → span the gateway → scheduler → engine → response

LOGS      per-request, structured, queryable
          → "what is the distribution of X across all requests?"
          → the source for capacity planning and benchmark design

For LLM services, logs do the heaviest lifting, because the questions you need to answer are distributional (length distributions, per-tenant patterns, cost attribution) and metrics can’t carry that cardinality.


3. The per-request log record#

JSON
{
  "ts": "2026-03-14T10:23:45.123Z",
  "event": "inference_complete",
  "request_id": "req_abc123",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "tenant": "team-search",
  "api_key_id": "key_xyz",
  "model": "llama-3-70b",
  "model_version": "v1.5.0-fp8-tp8",
  "replica": "llama-70b-7f9d4-x2k1p",
  "region": "us-east-1",

  "prompt_tokens": 1843,
  "cached_prompt_tokens": 1536,
  "output_tokens": 200,
  "max_tokens_requested": 2048,

  "queue_wait_ms": 12.4,
  "tokenize_ms": 1.2,
  "prefill_ms": 187.3,
  "ttft_ms": 234.1,
  "itl_p50_ms": 21.4,
  "itl_p95_ms": 38.2,
  "itl_max_ms": 112.0,
  "detokenize_ms": 8.1,
  "e2e_ms": 4512.0,

  "batch_size_at_admission": 47,
  "kv_blocks_peak": 128,
  "kv_block_seconds": 578.0,
  "preempted": false,
  "preemption_count": 0,

  "finish_reason": "stop",
  "temperature": 0.7,
  "top_p": 0.9,
  "guided_format": null,
  "n": 1,

  "status": 200,
  "client_disconnected": false,
  "tokens_after_disconnect": 0
}

Fields most often missing and most valuable:

  • cached_prompt_tokens — measures prefix caching’s actual benefit
  • kv_block_seconds — the missing term in cost attribution (Section XI.03)
  • batch_size_at_admission — efficiency signal per request
  • tokens_after_disconnect — pure waste, per request
  • replica — enables the per-replica imbalance analysis

4. Tracing an LLM request properly#

Go
// go.opentelemetry.io/otel — one span per stage of the request.
var tracer = otel.Tracer("inference")

func handle(ctx context.Context, req *Request) error {
	ctx, root := tracer.Start(ctx, "inference.request")
	defer root.End()
	// OpenTelemetry GenAI semantic conventions
	root.SetAttributes(
		attribute.String("gen_ai.system", "vllm"),
		attribute.String("gen_ai.request.model", req.Model),
		attribute.Int("gen_ai.request.max_tokens", req.MaxTokens),
		attribute.Float64("gen_ai.request.temperature", req.Temperature),
	)

	// span runs one stage and records its duration and error.
	span := func(name string, fn func(trace.Span) error) error {
		_, s := tracer.Start(ctx, name)
		defer s.End()
		return fn(s)
	}

	span("auth", func(trace.Span) error { return authenticate(ctx, req) })
	span("rate_limit", func(trace.Span) error { return checkQuota(ctx, req) })
	span("tokenize", func(s trace.Span) error {
		req.IDs = tokenize(req)
		s.SetAttributes(attribute.Int("gen_ai.usage.input_tokens", len(req.IDs)))
		return nil
	})
	span("route", func(s trace.Span) error {
		req.Replica = router.Choose(req)
		s.SetAttributes(attribute.String("inference.replica", req.Replica.ID),
			attribute.Bool("inference.prefix_cache_hit", req.Replica.LikelyHit))
		return nil
	})
	span("queue", func(s trace.Span) error { // ← the interesting span
		<-req.Admitted
		s.SetAttributes(attribute.Int("inference.batch_size", engine.RunningCount()),
			attribute.Float64("inference.kv_usage", engine.KVUsage()))
		return nil
	})
	span("prefill", func(s trace.Span) error {
		s.SetAttributes(attribute.Int("inference.cached_tokens", req.CachedTokens))
		return nil
	})
	return span("decode", func(s trace.Span) error {
		s.SetAttributes(attribute.Int("gen_ai.usage.output_tokens", req.OutputTokens),
			attribute.String("gen_ai.response.finish_reason", req.FinishReason))
		return nil
	})
}

The queue span with the batch size and KV usage as attributes is what makes a trace diagnostic rather than decorative: when you look at a slow request, you immediately see the system state it encountered.

Sampling: 100% of errors, 100% of requests exceeding the SLO, 0.1-1% of the rest. Head-based sampling misses slow requests; use tail-based sampling if your tracing backend supports it.


5. Quality observability#

The dimension generic tooling has no concept of.

CHEAP, CONTINUOUS SIGNALS (compute for every request)
  output token count
  finish_reason distribution
  structured output validity (if applicable)
  presence of refusal patterns (regex on a sample)
  repetition detection (n-gram repetition rate)
  language of the response vs the request

SAMPLED SIGNALS (1% of traffic)
  LLM-as-judge score against a rubric
  embedding-based similarity to the reference model's output
  toxicity / safety classifier score

USER SIGNALS (all traffic, sparse)
  thumbs up/down
  regeneration rate               ← strong negative signal
  copy/paste rate                 ← positive signal if measurable
  conversation abandonment
  turns to resolution

PERIODIC (daily)
  full evaluation suite against production traffic samples
  drift detection: has the input distribution changed?

Regeneration rate is the best single user-side quality signal available in most products: it requires no explicit feedback, correlates well with dissatisfaction, and moves quickly.


6. The alerting philosophy#

PAGE (wake someone up)
  SLO violation sustained > 10 minutes
  error rate > 5%
  a region or model completely unavailable
  security event

TICKET (fix during business hours)
  efficiency degradation (low batch, low DRAM_ACTIVE)
  cost per token drift > 20%
  preemption rate elevated
  quality signals outside bounds
  hardware degradation (ECC errors accumulating)

DASHBOARD ONLY (no alert)
  normal variation
  informational metrics

NEVER ALERT ON
  GPU utilization
  raw queue depth (use wait time)
  individual request latency
  anything that fires more than once a week without action

The last line is the discipline that keeps alerting useful. An alert that fires weekly and is always acknowledged without action should be deleted or converted to a ticket.


7. The on-call runbook#

For each alert, the runbook entry should contain:

ALERT: TTFTSLOViolation

WHAT IT MEANS
  p95 TTFT for prompts ≤ 2k tokens has exceeded 800 ms for 10 minutes.

FIRST CHECKS (in order)
  1. Grafana → Service Health → is queue_wait_p95 elevated?
       YES → capacity problem. Go to CAPACITY below.
       NO  → prefill got slower. Go to PREFILL below.
  2. Is this one replica or all? (per-replica heatmap)
       one → check that replica's health; consider removing it
  3. Did anything deploy in the last hour? (deployment annotations)
       yes → consider rollback

CAPACITY
  - check avg_running_batch vs max_num_seqs
  - check kv_cache_usage
  - check arrival rate vs the last 24 hours
  - if arrival rate is up: scale (link to the runbook)
  - if arrival rate is normal but batch is low: routing problem

PREFILL
  - check the prompt length distribution (has it shifted?)
  - check prefix_cache_hit_rate (has it dropped? routing change?)
  - check for GPU throttling

ESCALATION
  if unresolved in 20 minutes, page <team>

RELATED
  dashboards, past incidents, the capacity plan

A runbook that names specific dashboards and specific next queries is worth ten times one that says “investigate the cause.”


8. Production implications#

  • Budget real engineering time for observability. The LLM-specific signals don’t come from any product.
  • Structured per-request logs are the foundation. Retain 30+ days.
  • Use OpenTelemetry GenAI semantic conventions so your traces are portable and tool-readable.
  • Instrument quality, not just performance. It’s the failure mode that matters.
  • Tail-based trace sampling so you capture the slow requests.
  • Write runbooks with specific next steps, and update them after every incident.
  • Prune alerts quarterly. Delete anything that fires without action.
  • Emit model_version and replica everywhere — correlation depends on it.

9. Common mistakes#

Relying on generic APM. Misses everything LLM-specific.

No quality observability. You learn about regressions from support tickets.

Head-based trace sampling. Systematically misses slow requests.

Metrics without per-request logs. Can’t do distributional analysis or capacity planning.

Alerting on GPU utilization.

Runbooks that say “investigate.”

Not emitting model_version. Can’t correlate regressions with deployments.

Alert fatigue. Alerts that fire without action train people to ignore them.


10. Hands-on exercise#

A. Emit the full record. Implement the per-request log from section 3, including the four “often missing” fields. Verify each is populated correctly.

B. Trace properly. Implement the tracing from section 4, with the queue span carrying system state. Trace a slow request and confirm you can see why it was slow from the trace alone.

C. Quality signals. Implement the cheap continuous quality signals from section 5. Establish baselines over a week. Then deploy a degraded model and see which signals move first and by how much.

D. Write a runbook. For your top three alerts, write the runbook entries in the format of section 7, with specific dashboards and queries.

E. Audit alerts. For an existing service, list every alert and when it last fired. How many fired without resulting in action? Delete or downgrade those.

F. Distributional analysis. Using your per-request logs, answer: what’s the output-length distribution? Which tenant has the highest cost per request? What’s the prefix cache benefit per tenant? These queries are impossible without the logs.


11. Interview questions#

  1. What does generic APM miss for an LLM service?
  2. What per-request fields would you log, and what does each enable?
  3. Why do logs matter more than metrics for LLM capacity planning?
  4. How would you observe output quality continuously?
  5. Why is tail-based trace sampling important here?
  6. What makes a good runbook entry?
  7. What would you never alert on, and why?

12. Further reading#

↑↓ navigate↵ openesc close