PidokuInfra

Cluster and Capacity Management

Expert Advanced 1h Difficulty 3/5 Topic 10 of 11

Prerequisites XI.02, 04


1. What this covers#

The operational management of a GPU fleet: keeping nodes healthy, keeping capacity allocated sensibly, and knowing when to buy more.

DAILY       node health, cordoning bad hardware, watching utilization
WEEKLY      rebalancing placement, reviewing tenant growth
MONTHLY     capacity review against the plan, cost review
QUARTERLY   procurement, hardware refresh planning

2. Node lifecycle#

PROVISIONED → VALIDATED → READY → SERVING → DRAINING → CORDONED → REPAIRED
                                     ↑                      ↓
                                     └──────────────────────┘

Validation (before a node ever serves traffic)#

Shell
#!/bin/bash
# node-validation.sh — run on every new or repaired node
set -e

# 1. All GPUs present and healthy
[ "$(nvidia-smi -L | wc -l)" -eq 8 ] || fail "expected 8 GPUs"
nvidia-smi --query-gpu=ecc.errors.uncorrected.aggregate.total \
           --format=csv,noheader | grep -qv '[1-9]' || fail "ECC errors"

# 2. PCIe links at full width and generation
nvidia-smi -q | grep -A4 "GPU Link Info" | grep -q "x16" || fail "degraded PCIe"

# 3. NVLink topology as expected
nvidia-smi topo -m | grep -q "SYS" && fail "unexpected SYS link between GPUs"

# 4. Achieved bandwidth within tolerance
./bandwidthTest --memory=pinned | grep -q "PASS" || fail "H2D bandwidth"
./p2pBandwidthLatencyTest > /tmp/p2p.txt
python check_p2p.py /tmp/p2p.txt --min-gbps 300 || fail "NVLink bandwidth"

# 5. Collectives
./all_reduce_perf -b 8M -e 8M -g 8 | python check_nccl.py --min-busbw 120

# 6. Host: CPU count, memory, NUMA, storage
nproc | grep -q "^[0-9]\{3\}" || fail "CPU count"
numactl --hardware | grep -q "available: 2 nodes" || warn "NUMA layout"

# 7. Model cache volume present and writable
[ -w /var/cache/models ] || fail "model cache not writable"

echo "VALIDATION PASSED"

Run this on every node before it joins the pool. A node with a degraded PCIe link or a missing NVLink connection will silently underperform and you will spend days finding it (Section X, case 5).


3. Health monitoring and automatic cordoning#

CONTINUOUS CHECKS (DCGM + node-problem-detector)
  Xid errors                    → classify: software (restart) vs hardware (cordon)
  ECC uncorrectable             → cordon immediately
  ECC correctable, rising       → schedule replacement
  thermal / power throttling    → investigate cooling; may be environmental
  clock reduction sustained     → investigate
  GPU not responding            → cordon
  NVLink errors                 → cordon (usually a cable or connector)
  PCIe replay errors rising     → cordon

AUTOMATIC RESPONSE
  taint the node: inference.platform/unhealthy=true:NoSchedule
  drain running pods gracefully (respecting terminationGracePeriodSeconds)
  create a ticket with the diagnostic output
  alert if > 2% of nodes are cordoned

Automate the cordon. Manual response means a bad node keeps taking traffic and causing incidents until someone notices the pattern.


4. The capacity dashboard#

The one the platform team looks at daily:

FLEET
  total GPUs, by type
  allocated / free / cordoned
  allocation by model
  allocation by tenant

UTILIZATION
  cluster-wide DRAM_ACTIVE (the real utilization — Section X.04)
  average running batch across all replicas
  effective utilization (Section X.04's composite)
  utilization by model — which models are under-utilized?

DEMAND
  requests/sec and tokens/sec, by model and tenant
  trend: 7-day and 30-day
  peak-to-average ratio

CAPACITY HEALTH
  headroom: (capacity − peak demand) / capacity
  time spent over 80% utilization (last 7 days)
  rejection rate
  queue wait p95

COST
  $/hour, $/M tokens
  cost by model and tenant
  cost trend vs demand trend (are we getting more or less efficient?)

“Cost trend vs demand trend” is the summary metric for the platform team. If demand grew 40% and cost grew 40%, you’re standing still. If cost grew 15%, you’re improving.


5. Fragmentation and consolidation#

SYMPTOM: free GPUs exist but a new deployment can't be placed.

WEEKLY REVIEW
  1. Compute the placement that the current demand would produce
     from scratch (the "ideal" placement).
  2. Compare to the actual placement.
  3. If the ideal uses ≥ 10% fewer nodes, schedule a consolidation.

CONSOLIDATION PROCEDURE
  during a low-traffic window:
    1. drain the target replicas one at a time
    2. reschedule onto the consolidated placement
    3. verify readiness before draining the next
  → never move more than one replica of a model at a time
  → abort if latency degrades

Consolidation is disruptive and should be rare. Prevent fragmentation by standardizing GPU counts (Section XII.04) rather than fixing it repeatedly.


6. Growth and procurement#

MONTHLY CAPACITY REVIEW
  actual demand vs the plan's projection
  → if actual > projected by > 15%, re-plan
  headroom remaining at peak
  → if < 20%, start the procurement process

THE PROCUREMENT TIMELINE
  decision                 T
  approval                 T + 2-4 weeks
  order placed             T + 4 weeks
  delivery                 T + 8-20 weeks       ← the long pole
  rack, cable, validate    T + 10-22 weeks
  in service               T + 11-23 weeks

→ you must decide roughly 3-6 months before you need the capacity.
→ therefore: your headroom must cover 3-6 months of growth.

This is why the growth headroom factor in Section XI.02 is 1.2-1.5 and not 1.05. The lead time forces you to carry capacity you don’t yet need.

ALTERNATIVES TO BUYING
  cloud burst capacity (expensive per hour, no lead time)
  external API provider for overflow (Section XII.09)
  efficiency work (quantization, prefix caching, consolidation)
    → often cheaper and faster than procurement

Always evaluate efficiency work against procurement. A 40% efficiency improvement delivered in 6 weeks beats hardware delivered in 20 weeks, and it’s permanent.


7. Hardware heterogeneity#

Real fleets accumulate generations: A100s, H100s, H200s, L40Ss.

MANAGING IT
  label nodes by GPU type and capability
    nvidia.com/gpu.product, nvidia.com/gpu.memory
  models declare requirements and preferences in the registry
    gpu_type: [H100-80GB, H200-141GB]        # acceptable
    gpu_preference: H200-141GB               # preferred
  the scheduler scores accordingly (Section XII.04)

WORKLOAD-TO-HARDWARE MATCHING
  decode-heavy, latency-sensitive   → H200 (bandwidth)
  prefill-heavy                     → H100 (FLOPs/$)
  small models, moderate traffic    → L40S / A10G (cost)
  batch / offline                   → oldest hardware, or spot

  → measure $/M tokens per hardware type per workload (Section XI.03)
    and route accordingly

Heterogeneity is an opportunity, not just a burden: matching workloads to the hardware they’re actually bound by can improve fleet-wide cost 20-30%.


8. Production implications#

  • Validate every node before it serves. The script in section 2.
  • Automate cordoning on hardware faults.
  • Track effective utilization, not nvidia-smi. (Section X.04.)
  • Review capacity monthly against the plan.
  • Start procurement 3-6 months ahead. Lead times force it.
  • Evaluate efficiency work against procurement. Often faster and cheaper.
  • Standardize GPU counts per model to prevent fragmentation.
  • Label and match hardware to workload.
  • Track cost trend vs demand trend as the platform team’s KPI.

9. Common mistakes#

No node validation. Degraded hardware silently underperforms.

Manual response to hardware faults. Repeated incidents from the same node.

Tracking nvidia-smi utilization. Says 100% always (Section X.04).

Procurement started when headroom runs out. Too late by 3-6 months.

Not considering efficiency work as an alternative to buying.

Treating all GPUs as fungible. Wastes the bandwidth-rich ones on prefill-heavy work.

Frequent consolidation. Disruptive; prevent fragmentation instead.

No capacity review cadence. You discover problems from incidents.


10. Hands-on exercise#

A. Write the validation script. Adapt section 2’s script to your hardware. Run it on every node you have access to. Did any fail? What did you find?

B. Build the capacity dashboard. Implement the panels from section 4. Which numbers surprised you?

C. Compute the efficiency trend. For the last 3 months, plot demand (tokens/day) and cost ($/day). Is the ratio improving? By how much?

D. Ideal vs actual placement. Compute what placement the current demand would produce from scratch. Compare to actual. How much fragmentation is there?

E. Hardware matching. For each hardware type you have, measure $/M tokens for a decode-heavy and a prefill-heavy workload. Build the matching table. Is your current allocation optimal?

F. Procurement model. Given your growth rate and a 16-week lead time, compute when you must decide to order for a given headroom target.


11. Interview questions#

  1. What would you validate on a new GPU node before it serves traffic?
  2. How do you respond to a hardware Xid error automatically?
  3. What’s the right utilization metric for a GPU fleet, and why not nvidia-smi?
  4. How far ahead must you start procurement, and why?
  5. When is efficiency work a better answer than buying hardware?
  6. How do you manage a heterogeneous GPU fleet?
  7. What single metric summarizes whether a platform team is doing well?

12. Further reading#

  • [REFERENCE] NVIDIA DCGM and GPU Operator documentation
  • [REFERENCE] Kubernetes node-problem-detector
  • [FUNDAMENTAL] Google SRE Book, capacity planning and demand forecasting
  • Next: 11 — Building the platform: a roadmap

↑↓ navigate↵ openesc close