1. The question#
How many GPUs do I need? And its harder cousin: how many will I need in six months, and when do I have to order them?
Section V.15 did the per-node arithmetic. This file adds the production factors: traffic modeling, headroom, redundancy, growth, and lead times.
2. The formula#
GPUs = ceil(
peak_demand
÷ per_node_capacity_at_SLO
÷ target_utilization
× redundancy_factor
× growth_headroom
) × GPUs_per_node
where:
peak_demand tokens/sec at your busiest 5-minute window
per_node_capacity measured, at your SLO (not peak throughput!)
target_utilization 0.60-0.75 for latency-sensitive (the queueing wall)
redundancy_factor 1 + (1/nodes) for N+1, or 1.5 for N+50%
growth_headroom 1.2-1.5 depending on your planning horizonTwo terms people omit and shouldn’t: target_utilization (Section I.05 — latency explodes
near 100%) and growth_headroom (GPU lead times are months).
3. Step 1 — model the traffic#
This is the input everything else depends on, and it’s usually the weakest part of a capacity plan.
FROM PRODUCT / BUSINESS
daily active users
interactions per user per day
peak-to-average ratio
growth rate
FROM MEASUREMENT (structured logs, Section X.10)
prompt length distribution (p50, p90, p95, p99)
output length distribution
request rate by hour of day and day of week
the shape of the peak (sharp spike or broad plateau?)
DERIVED
peak_requests_per_second
peak_prefill_tokens_per_second
peak_decode_tokens_per_secondWORKED:
50,000 DAU × 8 sessions/day × 6 turns = 2.4M requests/day
Peak hour = 14% of daily volume (measured, typical for consumer)
= 336,000 requests in the peak hour = 93 req/s
Peak 5-minute burst = 1.4× the peak hour rate = 130 req/s
p50 prompt 480 tokens, p95 3,200; mean 890 (the tail matters!)
p50 output 190 tokens, p95 1,100; mean 320
peak prefill: 130 × 890 = 115,700 tok/s
peak decode: 130 × 320 = 41,600 tok/sUse the MEAN for capacity, not the median. The mean is what determines total work; the tail pulls it well above the median. Using p50 will underprovision by 40-80% for a heavy-tailed distribution.
4. Step 2 — measure per-node capacity at the SLO#
Not peak throughput. Capacity at the SLO.
Run the benchmark from Section X.07, open-loop, at increasing arrival rates.
Find the rate at which goodput starts falling — that's your capacity.
arrival rate p95 TTFT p95 ITL goodput
4 req/s 280 ms 22 ms 100%
8 340 ms 26 ms 100%
12 480 ms 31 ms 99.4%
14 690 ms 38 ms 98.1%
16 1,240 ms 44 ms 91.3% ← knee
18 3,100 ms 52 ms 71.0%
→ capacity at SLO (goodput ≥ 97%) = 14 req/s per node
→ peak throughput would say 18+, and would be wrongThe gap between “peak throughput” and “capacity at SLO” is typically 25-40%. Planning with the former guarantees SLO violations at peak.
5. Step 3 — assemble the plan#
peak demand: 130 req/s
capacity per node at SLO: 14 req/s
──────────
nodes for peak: 9.3 → 10
target utilization 0.70: 10 / 0.70 = 14.3 → 15 nodes
(running 10 nodes' worth of work on 15 nodes'
capacity keeps you off the queueing wall)
N+1 redundancy: 16 nodes
growth headroom 1.3
(6-month horizon,
30% expected growth): 16 × 1.3 = 20.8 → 21 nodes
ANSWER: 21 nodes × 8 GPUs = 168 GPUsSanity-check it against the token arithmetic:
21 nodes × 14 req/s × 320 output tokens = 94,080 output tok/s of capacity
Peak demand: 41,600 output tok/s
Ratio: 2.26x
Is 2.26x too much? Decompose:
1/0.70 (utilization) × 1.07 (N+1) × 1.3 (growth) = 1.99
plus the rounding: 2.26
→ Consistent. The headroom is deliberate, not accidental.Always do this cross-check. If the ratio is much larger than the product of your explicit factors, you’ve double-counted something.
6. The factors, justified#
Target utilization (the biggest and most-argued factor)#
From queueing theory (Section I.05): W_queue ≈ W_service × ρ/(1-ρ)
ρ = 0.60 → queue adds 1.5× service time
ρ = 0.70 → 2.3×
ρ = 0.80 → 4×
ρ = 0.90 → 9×
ρ = 0.95 → 19×
For a latency SLO with 300 ms of queueing budget and 200 ms service:
queue budget / service = 1.5 → ρ ≈ 0.60
For batch workloads with no latency SLO: ρ = 0.95 is fine.This factor is where the money is, and it’s where the argument happens. “Why are we running at 70%?” is answered by the queueing formula and your latency budget, not by preference.
Redundancy#
N+1: survive one node failure. redundancy = (N+1)/N
N+2: survive two, or one failure during a deployment
2N: survive an AZ loss (Section XI.06)
For GPU fleets: N+1 within an AZ, and enough capacity in a second AZ to
serve degraded traffic if the first is lost.Growth headroom and lead time#
GPU procurement lead times: weeks to quarters, depending on the part and
your relationship with the supplier.
Plan: headroom = expected_growth_over(lead_time + safety_margin)
lead time 12 weeks, growth 8%/month → 12 weeks = 2.8 months
→ 1.08^2.8 = 1.24
→ plus safety: 1.3Order before you need it. Capacity planning that produces an order date after the need date is a plan to have an outage.
7. What changes the plan#
Re-run capacity planning when any of these change materially:
INPUT EFFECT
prompt length distribution linear-to-quadratic on prefill
output length distribution linear on decode
peak-to-average ratio linear on peak sizing
model size or architecture large, nonlinear
precision (quantization) 1.5-3x
prefix cache hit rate large for prefill-heavy workloads
SLO thresholds large (via target utilization)
context length limits large (via concurrency)
new product features unpredictable — model them explicitlyPrompt length distribution drift is the most common surprise. A product change that adds retrieved context to every request can double your prefill load overnight without any change in request rate.
8. Producing the artifact#
A capacity plan is a document, not a number. It should contain:
1. TRAFFIC MODEL assumptions, sources, measured distributions
2. UNIT CAPACITY measured, with the benchmark methodology
3. THE CALCULATION every factor, with justification
4. THE ANSWER GPU count, instance types, timeline
5. COST monthly, and cost per million tokens
6. SENSITIVITY what if traffic is 2x? what if prompts double?
7. TRIGGERS what measurements would invalidate this plan
8. REVIEW DATESection 6 (sensitivity) is what makes it useful. A plan with a single number is fragile; a plan with “if peak exceeds 160 req/s we need 26 nodes, order by March” is actionable.
9. Production implications#
- Base the traffic model on measurement. Structured per-request logs (Section X.10) are the source. Product forecasts are an input, not the input.
- Measure unit capacity at the SLO, with a realistic benchmark.
- Justify the utilization target from the queueing formula and your latency budget.
- Track actual vs planned monthly. A plan you don’t check is a guess.
- Set triggers: “if p95 prompt length exceeds 4,000 tokens, re-plan.”
- Account for lead time. Order date = need date − lead time − safety.
- Include the efficiency levers in the plan: “this assumes FP8 and 60% prefix cache hit rate; without them, add 8 nodes.”
10. Common mistakes#
Planning with peak throughput instead of capacity at SLO. 25-40% underprovisioned.
Using median instead of mean lengths. 40-80% underprovisioned for heavy-tailed traffic.
Planning for 90%+ utilization with a latency SLO. The queueing wall.
Forgetting growth and lead time. The plan is correct and arrives late.
No sensitivity analysis. One assumption changes and the plan is void.
Not accounting for the efficiency levers. Planning at FP16 when you’ll deploy FP8 means 2x overprovisioning.
Planning once. Traffic shape drifts continuously.
11. Hands-on exercise#
A. Build the traffic model. From real (or realistic synthetic) request logs, extract: the length distributions, the hourly rate profile, the peak-to-average ratio, and the peak 5-minute burst. Plot each.
B. Measure unit capacity at the SLO. Using your Section X.07 harness, run open-loop at increasing rates. Find the goodput knee. Compare to peak throughput — how big is the gap?
C. Produce the plan. Write the full capacity plan document from section 8 for a service you know. Include the sensitivity analysis.
D. Sensitivity. Vary each input ±50% and compute the effect on the GPU count. Rank the inputs by impact. Which most deserves careful measurement?
E. The efficiency lever. Compute the plan with and without: FP8, prefix caching, chunked prefill. How many GPUs does each lever save? Express each as annual dollars.
F. Cross-check. For your plan, verify the token arithmetic matches the request arithmetic (section 5’s sanity check).
12. Interview questions#
- Walk me through capacity planning for an LLM service.
- Why plan with capacity-at-SLO rather than peak throughput?
- Why target 70% utilization? Justify it quantitatively.
- Why use the mean rather than the median request length?
- What would invalidate a capacity plan, and how would you detect it?
- How does prefix caching change your capacity requirement?
- Your plan says 21 nodes. How do you sanity-check that?
13. Further reading#
- [FUNDAMENTAL] Google SRE Book, capacity planning and demand forecasting chapters
- [FUNDAMENTAL] Little’s Law and basic queueing theory
- [ESTABLISHED] Section V.15 of this curriculum — the per-node arithmetic
- Next: 03 — Cost per token