Below the API

AI Infrastructure

AdvancedModule7 topics~6h 30m
completed

Topics, in order

About this module

Everything so far applies to any service. This module is about what changes when the service is a model on a GPU: new resources to watch, new units of work, new failure modes — and a product whose output can be wrong while every infrastructure metric is green.

#LessonThe question it answers
01What Is DifferentWhich assumptions from web-service observability break?
02GPU TelemetryWhat does a GPU report, and which numbers lie?
03Inference Engine MetricsWhat should an LLM server expose, and how do I read it?
04Tracing LLM RequestsHow do I trace a model call, a tool call, an agent?
05Platform and FleetWhat do I watch across a cluster, and what do I scale on?
06Cost, Power and EfficiencyWhat does a token cost in dollars and joules?
07Quality and EvalsHow do I observe whether the answers are any good?

When you finish you can list the metrics that matter at each layer of an inference platform, explain why GPU utilization is not one of them, and compute cost per million tokens from counters.

This module refers often to Inference Engineering (prefill, decode, KV cache, continuous batching) and GPU Engineering (memory bandwidth, power). Each lesson links the specific background it needs.

↑↓ navigate ↵ open