1. What this covers#
The operational management of a GPU fleet: keeping nodes healthy, keeping capacity allocated sensibly, and knowing when to buy more.
DAILY node health, cordoning bad hardware, watching utilization
WEEKLY rebalancing placement, reviewing tenant growth
MONTHLY capacity review against the plan, cost review
QUARTERLY procurement, hardware refresh planning2. Node lifecycle#
PROVISIONED → VALIDATED → READY → SERVING → DRAINING → CORDONED → REPAIRED
↑ ↓
└──────────────────────┘Validation (before a node ever serves traffic)#
#!/bin/bash
# node-validation.sh — run on every new or repaired node
set -e
# 1. All GPUs present and healthy
[ "$(nvidia-smi -L | wc -l)" -eq 8 ] || fail "expected 8 GPUs"
nvidia-smi --query-gpu=ecc.errors.uncorrected.aggregate.total \
--format=csv,noheader | grep -qv '[1-9]' || fail "ECC errors"
# 2. PCIe links at full width and generation
nvidia-smi -q | grep -A4 "GPU Link Info" | grep -q "x16" || fail "degraded PCIe"
# 3. NVLink topology as expected
nvidia-smi topo -m | grep -q "SYS" && fail "unexpected SYS link between GPUs"
# 4. Achieved bandwidth within tolerance
./bandwidthTest --memory=pinned | grep -q "PASS" || fail "H2D bandwidth"
./p2pBandwidthLatencyTest > /tmp/p2p.txt
python check_p2p.py /tmp/p2p.txt --min-gbps 300 || fail "NVLink bandwidth"
# 5. Collectives
./all_reduce_perf -b 8M -e 8M -g 8 | python check_nccl.py --min-busbw 120
# 6. Host: CPU count, memory, NUMA, storage
nproc | grep -q "^[0-9]\{3\}" || fail "CPU count"
numactl --hardware | grep -q "available: 2 nodes" || warn "NUMA layout"
# 7. Model cache volume present and writable
[ -w /var/cache/models ] || fail "model cache not writable"
echo "VALIDATION PASSED"Run this on every node before it joins the pool. A node with a degraded PCIe link or a missing NVLink connection will silently underperform and you will spend days finding it (Section X, case 5).
3. Health monitoring and automatic cordoning#
CONTINUOUS CHECKS (DCGM + node-problem-detector)
Xid errors → classify: software (restart) vs hardware (cordon)
ECC uncorrectable → cordon immediately
ECC correctable, rising → schedule replacement
thermal / power throttling → investigate cooling; may be environmental
clock reduction sustained → investigate
GPU not responding → cordon
NVLink errors → cordon (usually a cable or connector)
PCIe replay errors rising → cordon
AUTOMATIC RESPONSE
taint the node: inference.platform/unhealthy=true:NoSchedule
drain running pods gracefully (respecting terminationGracePeriodSeconds)
create a ticket with the diagnostic output
alert if > 2% of nodes are cordonedAutomate the cordon. Manual response means a bad node keeps taking traffic and causing incidents until someone notices the pattern.
4. The capacity dashboard#
The one the platform team looks at daily:
FLEET
total GPUs, by type
allocated / free / cordoned
allocation by model
allocation by tenant
UTILIZATION
cluster-wide DRAM_ACTIVE (the real utilization — Section X.04)
average running batch across all replicas
effective utilization (Section X.04's composite)
utilization by model — which models are under-utilized?
DEMAND
requests/sec and tokens/sec, by model and tenant
trend: 7-day and 30-day
peak-to-average ratio
CAPACITY HEALTH
headroom: (capacity − peak demand) / capacity
time spent over 80% utilization (last 7 days)
rejection rate
queue wait p95
COST
$/hour, $/M tokens
cost by model and tenant
cost trend vs demand trend (are we getting more or less efficient?)“Cost trend vs demand trend” is the summary metric for the platform team. If demand grew 40% and cost grew 40%, you’re standing still. If cost grew 15%, you’re improving.
5. Fragmentation and consolidation#
SYMPTOM: free GPUs exist but a new deployment can't be placed.
WEEKLY REVIEW
1. Compute the placement that the current demand would produce
from scratch (the "ideal" placement).
2. Compare to the actual placement.
3. If the ideal uses ≥ 10% fewer nodes, schedule a consolidation.
CONSOLIDATION PROCEDURE
during a low-traffic window:
1. drain the target replicas one at a time
2. reschedule onto the consolidated placement
3. verify readiness before draining the next
→ never move more than one replica of a model at a time
→ abort if latency degradesConsolidation is disruptive and should be rare. Prevent fragmentation by standardizing GPU counts (Section XII.04) rather than fixing it repeatedly.
6. Growth and procurement#
MONTHLY CAPACITY REVIEW
actual demand vs the plan's projection
→ if actual > projected by > 15%, re-plan
headroom remaining at peak
→ if < 20%, start the procurement process
THE PROCUREMENT TIMELINE
decision T
approval T + 2-4 weeks
order placed T + 4 weeks
delivery T + 8-20 weeks ← the long pole
rack, cable, validate T + 10-22 weeks
in service T + 11-23 weeks
→ you must decide roughly 3-6 months before you need the capacity.
→ therefore: your headroom must cover 3-6 months of growth.This is why the growth headroom factor in Section XI.02 is 1.2-1.5 and not 1.05. The lead time forces you to carry capacity you don’t yet need.
ALTERNATIVES TO BUYING
cloud burst capacity (expensive per hour, no lead time)
external API provider for overflow (Section XII.09)
efficiency work (quantization, prefix caching, consolidation)
→ often cheaper and faster than procurementAlways evaluate efficiency work against procurement. A 40% efficiency improvement delivered in 6 weeks beats hardware delivered in 20 weeks, and it’s permanent.
7. Hardware heterogeneity#
Real fleets accumulate generations: A100s, H100s, H200s, L40Ss.
MANAGING IT
label nodes by GPU type and capability
nvidia.com/gpu.product, nvidia.com/gpu.memory
models declare requirements and preferences in the registry
gpu_type: [H100-80GB, H200-141GB] # acceptable
gpu_preference: H200-141GB # preferred
the scheduler scores accordingly (Section XII.04)
WORKLOAD-TO-HARDWARE MATCHING
decode-heavy, latency-sensitive → H200 (bandwidth)
prefill-heavy → H100 (FLOPs/$)
small models, moderate traffic → L40S / A10G (cost)
batch / offline → oldest hardware, or spot
→ measure $/M tokens per hardware type per workload (Section XI.03)
and route accordinglyHeterogeneity is an opportunity, not just a burden: matching workloads to the hardware they’re actually bound by can improve fleet-wide cost 20-30%.
8. Production implications#
- Validate every node before it serves. The script in section 2.
- Automate cordoning on hardware faults.
- Track effective utilization, not
nvidia-smi. (Section X.04.) - Review capacity monthly against the plan.
- Start procurement 3-6 months ahead. Lead times force it.
- Evaluate efficiency work against procurement. Often faster and cheaper.
- Standardize GPU counts per model to prevent fragmentation.
- Label and match hardware to workload.
- Track cost trend vs demand trend as the platform team’s KPI.
9. Common mistakes#
No node validation. Degraded hardware silently underperforms.
Manual response to hardware faults. Repeated incidents from the same node.
Tracking nvidia-smi utilization. Says 100% always (Section X.04).
Procurement started when headroom runs out. Too late by 3-6 months.
Not considering efficiency work as an alternative to buying.
Treating all GPUs as fungible. Wastes the bandwidth-rich ones on prefill-heavy work.
Frequent consolidation. Disruptive; prevent fragmentation instead.
No capacity review cadence. You discover problems from incidents.
10. Hands-on exercise#
A. Write the validation script. Adapt section 2’s script to your hardware. Run it on every node you have access to. Did any fail? What did you find?
B. Build the capacity dashboard. Implement the panels from section 4. Which numbers surprised you?
C. Compute the efficiency trend. For the last 3 months, plot demand (tokens/day) and cost ($/day). Is the ratio improving? By how much?
D. Ideal vs actual placement. Compute what placement the current demand would produce from scratch. Compare to actual. How much fragmentation is there?
E. Hardware matching. For each hardware type you have, measure $/M tokens for a decode-heavy and a prefill-heavy workload. Build the matching table. Is your current allocation optimal?
F. Procurement model. Given your growth rate and a 16-week lead time, compute when you must decide to order for a given headroom target.
11. Interview questions#
- What would you validate on a new GPU node before it serves traffic?
- How do you respond to a hardware Xid error automatically?
- What’s the right utilization metric for a GPU fleet, and why not
nvidia-smi? - How far ahead must you start procurement, and why?
- When is efficiency work a better answer than buying hardware?
- How do you manage a heterogeneous GPU fleet?
- What single metric summarizes whether a platform team is doing well?
12. Further reading#
- [REFERENCE] NVIDIA DCGM and GPU Operator documentation
- [REFERENCE] Kubernetes node-problem-detector
- [FUNDAMENTAL] Google SRE Book, capacity planning and demand forecasting
- Next: 11 — Building the platform: a roadmap