The capstone of Section V. Work every example with a calculator. This is the skill that gets you hired and that you’ll use weekly.
1. The method#
Every capacity question follows the same six steps:
1. MEMORY BUDGET what fits?
weights + KV + activations + overhead ≤ GPU memory
2. CONCURRENCY how many at once?
max_seqs = KV_budget / (KV_per_token × avg_context)
3. STEP TIME how fast per iteration?
T_step = max(bytes/BW, flops/FLOPS) / efficiency
4. THROUGHPUT how many tokens/sec?
output_tok/s = batch / T_step
plus prefill capacity
5. LATENCY CHECK does it meet the SLO?
ITL = T_step; TTFT = queue + prefill
6. FLEET SIZE how many nodes?
demand / per-node capacity × redundancy × headroomConstants to keep at hand:
H100 SXM: 80 GB, 3.35 TB/s, ~990 TFLOP/s BF16 dense
H200: 141 GB, 4.80 TB/s, ~990 TFLOP/s
A100 80G: 80 GB, 2.04 TB/s, ~312 TFLOP/s
L40S: 48 GB, 0.86 TB/s, ~362 TFLOP/s
Realistic efficiency: 60-75% of theoreticalDiagram — The capacity method as a pipeline#
flowchart TB M["Model config + precision"] --> W["Weight bytes"] G["GPU: VRAM and bandwidth"] --> F["Free for KV<br/>= VRAM x utilization - weights - overhead"] W --> F F --> CC["Max concurrent sequences<br/>= free / KV bytes per token x context"] G --> ST["Decode step time<br/>~ bytes read per step / bandwidth"] W --> ST CC --> TP["Throughput = concurrency / step time"] ST --> TP TP --> N["Nodes = demand / per-node capacity<br/>x redundancy x headroom"] class W,F,CC memory class ST,TP compute class N queue class M,G neutral
2. Example 1 — Can I serve this model on this GPU?#
Question: Llama-3-8B on one L40S (48 GB), 8k context. How many users?
1. MEMORY
weights BF16 = 8.03e9 × 2 = 16.1 GB
framework overhead ≈ 2 GB
activation headroom ≈ 2 GB
KV budget = 48 − 16.1 − 4 = 27.9 GB
2. CONCURRENCY
KV_per_token = 2 × 32 × 8 × 128 × 2 = 131,072 B = 128 KiB
per sequence at 8k = 1.074 GB
max_seqs = 27.9 / 1.074 = 25 sequences
3. STEP TIME (at batch 25, context 8k avg)
bytes = 16.1 GB + 25 × 8192 × 128 KiB = 16.1 + 26.8 = 42.9 GB
T_theoretical = 42.9 / 860 GB/s = 49.9 ms
T_real (70% eff) = 71 ms
4. THROUGHPUT
25 / 0.071 = 352 output tokens/sec
5. LATENCY
ITL = 71 ms — acceptable for batch work, poor for interactive chat
(users perceive < 50 ms as smooth)
ANSWER: 25 concurrent users at 8k context, 352 tok/s, ITL 71 ms.
The L40S's low bandwidth (860 GB/s) is the binding constraint.
An H100 would give ~4x the throughput at ~3x the price — better value
for this workload.Lesson: cheap GPUs with weak memory bandwidth are poor for decode. Compare $/GB/s, not $/GB or $/TFLOP.
3. Example 2 — Sizing a chat product#
Question: 50,000 daily active users, chat product, Llama-3-70B. What do I need?
0. TRAFFIC MODEL (do this first — everything depends on it)
50,000 DAU × 8 conversations/day × 6 turns = 2.4M requests/day
Peak hour = 15% of daily (typical for a consumer product)
= 360,000 requests/hour = 100 requests/sec at peak
Average: 800 input tokens (with history), 250 output tokens
1. TOKEN DEMAND AT PEAK
output: 100 req/s × 250 tok = 25,000 output tok/s
input: 100 req/s × 800 tok = 80,000 input tok/s (prefill)
2. PER-NODE CAPACITY (8×H100, TP=8)
weights: 141 GB / 8 = 17.6 GB per GPU
KV_per_token = 2 × 80 × 8 × 128 × 2 = 320 KiB total, 40 KiB per GPU
KV budget per GPU = 80 − 17.6 − 4 = 58.4 GB
avg context ≈ 800 + 125 = 925 tokens (mid-generation)
→ generous: assume 2k for safety
per sequence per GPU at 2k = 40 KiB × 2048 = 82 MB
max_seqs = 58.4 GB / 0.082 GB = 712 ← memory allows a lot
But latency limits us. At batch 256, context 2k:
bytes/GPU = 17.6 + 256 × 2048 × 40 KiB = 17.6 + 21.0 = 38.6 GB
T = 38.6/3350 = 11.5 ms theoretical
+ TP AllReduce overhead ~25% + 70% efficiency → ~21 ms
ITL = 21 ms ✓ good
output tok/s per node = 256 / 0.021 = 12,190
3. PREFILL CAPACITY
prefill throughput ≈ 8 × 990 TFLOP/s × 0.5 / (2 × 70e9) per token
≈ 28,000 prefill tok/s per node (compute bound)
Need 80,000 → 2.9 nodes for prefill alone
4. NODES NEEDED
decode: 25,000 / 12,190 = 2.1 nodes
prefill: 80,000 / 28,000 = 2.9 nodes
Since they share the same GPUs and interleave:
total GPU time fraction = 2.1 + 2.9 = 5.0 node-equivalents
→ 5 nodes at 100% utilization
5. HEADROOM AND REDUNDANCY
target 70% utilization (latency wall, Section I.05): 5/0.7 = 7.1
N+1 redundancy: 8
growth headroom 30%: 10 nodes
ANSWER: 10 nodes × 8 H100 = 80 GPUs.
At $2.50/GPU-hour: $4,800/day = $1.75M/year.
Cost per request: $4800 / 2.4M = $0.002Note how prefill dominated. With 800-token prompts and 250-token outputs, prefill is 58% of the GPU time. Prefix caching would be enormously valuable here — in a chat product, turn 6 re-prefills turns 1-5 every time. With caching, prefill drops to ~15% and you need ~6 nodes instead of 10. That’s a $700k/year optimization from one config flag.
4. Example 3 — The long-context question#
Question: A customer wants 128k context on Llama-3-8B. What’s the impact?
Baseline (8k context, one H100):
KV per seq = 1.07 GB
max_seqs = 59.9 / 1.07 = 55
step bytes = 16.1 + 55 × 8192 × 128 KiB = 16.1 + 59.0 = 75.1 GB
T = 22.4 ms real; throughput = 55/0.0224 = 2,455 tok/s
128k context:
KV per seq = 17.2 GB
max_seqs = 59.9 / 17.2 = 3
step bytes = 16.1 + 3 × 131072 × 128 KiB = 16.1 + 51.5 = 67.6 GB
T = 20.2 ms real; throughput = 3/0.0202 = 149 tok/s
Prefill: 128k tokens
FLOPs = 2×8e9×131072 + 2×32×131072²×4096 = 6,601 TFLOP
at 600 TFLOP/s achieved = 11 seconds ← TTFT
RATIO:
throughput: 2455 / 149 = 16.5x worse
cost per output token: 16.5x higher
TTFT: 11 s vs 0.15 s = 73x worse
concurrency: 18x worse
ANSWER: 128k context costs ~16x more per output token and gives 11-second TTFT.
Recommendations:
- Price it at 15-20x
- Route to a dedicated pool (don't poison the 8k pool)
- Enable chunked prefill (mandatory — an 11 s prefill would freeze everyone)
- Enable prefix caching (if the long document recurs, this is a 100x win)
- Evaluate RAG instead: retrieving 4k of relevant context costs 1/16th5. Example 4 — Quantization decision#
Question: Should we quantize Llama-3-70B to FP8? Workload: 500 in / 500 out, batch-heavy.
BF16 on 8×H100:
weights/GPU: 17.6 GB, KV budget 58.4 GB
at 1k avg context: KV/seq/GPU = 40 KiB × 1024 = 41 MB
max_seqs = 1,424 (memory) — latency-limited instead
at batch 256: bytes = 17.6 + 10.5 = 28.1 GB → T = 8.4 ms theo, 15 ms real
throughput = 256/0.015 = 17,067 tok/s
FP8 on 8×H100:
weights/GPU: 8.8 GB, KV budget 67.2 GB (+15%)
at batch 256: bytes = 8.8 + 10.5 = 19.3 GB → T = 5.8 ms theo, 10.3 ms real
throughput = 256/0.0103 = 24,854 tok/s → 1.46x
AND prefill gets ~1.8x from FP8 tensor cores
Could also raise batch: at batch 512,
bytes = 8.8 + 21.0 = 29.8 GB → T = 15.9 ms real
throughput = 512/0.0159 = 32,201 tok/s → 1.89x, ITL still acceptable
DECISION FACTORS:
✓ 1.5-1.9x throughput = 1.5-1.9x cost reduction
✓ FP8 quality loss typically < 0.5% on standard benchmarks
? Must validate on OUR task with long generations (Section IV.12)
✓ Hopper has native FP8 — this is [ESTABLISHED], not experimental
ANSWER: Yes, subject to quality validation. Expected saving ~$800k/year on a
$1.7M fleet. Run the Level 1-4 validation from Section IV.12 first.Contrast: INT4 weight-only would give ~2.5x on decode but ~1.0x on prefill. With a
500:500 workload where prefill is ~30% of GPU time, the blended gain is
1/(0.7/2.5 + 0.3/1.0) = 1.73x — comparable to FP8 but with more quality risk. FP8 wins here.
6. Example 5 — Diagnosing a slow system#
Question: “Our p95 TTFT is 4 seconds. We expected 400 ms.” What’s wrong?
Work through the possibilities in order of likelihood:
Measured: p95 TTFT 4,000 ms, p95 ITL 45 ms (fine), throughput 1,800 tok/s
Hypothesis 1: QUEUEING
Check: queue_wait metric. If it's 3,500 ms of the 4,000 → confirmed.
Cause: arrival rate exceeds capacity, or admission is too conservative.
Verify with Little's Law: L = λW. If λ=20 req/s and W=8 s, L=160 in system.
If max_num_seqs=64, then 96 are queueing. → CAPACITY PROBLEM.
Fix: add capacity, or reduce per-request cost, or shed load.
Hypothesis 2: LONG PREFILLS BLOCKING
Check: prompt length distribution. Is p99 prompt 50k tokens while p50 is 500?
Check: ITL variance. Spikes correlated with long-prompt arrivals?
Fix: chunked prefill. Also route long prompts to a separate pool.
Hypothesis 3: PREFILL IS GENUINELY SLOW
Check: prefill_ms metric alone. If prefill is 3,500 ms for a 2,000-token
prompt, that's 570 tok/s — far below the ~28,000 tok/s expected.
Cause: wrong attention kernel, unfused ops, FP32, no tensor cores.
Fix: profile it.
Hypothesis 4: LOW BATCH / POOR UTILIZATION
Check: average running batch size. If it's 4 when memory allows 60,
requests are arriving but not being admitted.
Cause: max_num_seqs set too low, or preemption thrashing.
Hypothesis 5: NETWORK / INGRESS
Check: client-measured TTFT vs server-measured. A 3.5 s gap = buffering proxy.
Hypothesis 6: TOKENIZATION
Check: for very long prompts with a slow Python tokenizer, this can be seconds.
METHOD: instrument each phase separately. queue_wait, tokenize_ms, prefill_ms,
first_token_ms, network. The breakdown identifies the phase in one look.The general principle: always decompose the metric before hypothesizing. A single TTFT number supports six explanations; a breakdown supports one.
7. Example 6 — Comparing hardware#
Question: H100 vs H200 vs A100 for a decode-heavy 70B workload. Assume $2.50, $3.20, $1.60 per GPU-hour.
Per-GPU with TP=8, weights 17.6 GB (BF16), batch 128, context 4k:
HBM BW KV budget bytes/step T_real tok/s/node $/node/hr
A100 80G 2.04 TB/s 58.4 GB 17.6+21=38.6 27 ms 4,741 $12.80
H100 80G 3.35 TB/s 58.4 GB 38.6 16.5 ms 7,758 $20.00
H200 141G 4.80 TB/s 119.4 GB 38.6 11.5 ms 11,130 $25.60
Cost per million output tokens:
A100: 12.80 / (4741 × 3600) × 1e6 = $0.75
H100: 20.00 / (7758 × 3600) × 1e6 = $0.72
H200: 25.60 / (11130 × 3600) × 1e6 = $0.64
But H200's larger memory allows a bigger batch:
KV budget 119.4 GB → batch 512 at 4k context
bytes = 17.6 + 84 = 101.6 GB → T = 30 ms real
tok/s/node = 512/0.030 = 17,067
cost = 25.60 / (17067 × 3600) × 1e6 = $0.42 ← best by far
ANSWER: H200, if available. Its bandwidth AND capacity advantages compound:
more bandwidth per step, and more memory means a bigger batch.
A100 is competitive on cost but gives worse latency (27 ms vs 11.5 ms ITL).
Choose by: latency SLO first, then $/M tokens.Lesson: compare on $/M tokens at your actual operating point, not on spec-sheet ratios.
8. Common estimation errors#
✗ Forgetting activation headroom → OOM on the first long prompt
✗ Using theoretical bandwidth → 30-40% optimistic
✗ Ignoring TP communication overhead → 20-30% optimistic
✗ Assuming 100% utilization → 40% optimistic
✗ Forgetting prefill in the time budget → 30-70% optimistic for RAG
✗ Using average context instead of p95 → OOM under load
✗ Forgetting redundancy and headroom → no capacity for failures
✗ Ignoring the length distribution's tail → a few requests dominateRule of thumb: take your theoretical number and multiply by 0.5-0.6 for a planning estimate. Then measure, and calibrate your personal fudge factor.
9. Hands-on exercise#
A. Build the calculator. Write a script implementing all six steps, taking model config, GPU spec, workload shape, and SLOs as input. Output: max concurrency, throughput, ITL, TTFT, nodes needed, cost per million tokens. This becomes your standard tool.
B. Validate it. Run a real workload on real hardware and compare every predicted number to the measured one. Compute your fudge factor per quantity. Which prediction is furthest off?
C. Do all six examples yourself with a calculator, without looking at the answers. Then compare.
D. New scenarios. Size a fleet for:
- Code completion: 2,000 in / 30 out, p95 TTFT < 200 ms, 500 req/s
- Document summarization: 30,000 in / 800 out, batch job, no latency SLO, 10,000 docs/hour
- Agentic tool loop: 6,000 in / 150 out, 12 steps per task, p95 step latency < 2 s For each: which optimization matters most, and why?
E. Sensitivity. For example 2, vary each assumption ±50% and rank by impact on fleet size. Which assumption most deserves careful measurement?
10. Interview questions#
- Size a cluster for 50,000 DAU on a 70B model. Walk me through your reasoning.
- A customer wants 128k context. Quantify the impact on capacity and cost.
- Should we quantize to FP8? What’s the expected gain and what would you check first?
- Our p95 TTFT is 4 seconds. Walk me through your diagnosis.
- H100 or H200 for a decode-heavy workload? Show the arithmetic.
- What’s the most commonly underestimated factor in capacity planning?
- Why did prefill dominate in the chat example, and what would you do about it?
11. Further reading#
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022)
- [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023) — evaluation methodology
- [REFERENCE] vLLM
benchmarks/benchmark_serving.py— how to measure what you predicted - Next: Section VI — GPU Computing