★ A wrong benchmark is worse than no benchmark, because it produces confident wrong decisions.
1. The requirements for a valid benchmark#
1. REALISTIC WORKLOAD SHAPE
Real prompt and output length distributions, not fixed lengths.
2. REALISTIC ARRIVAL PATTERN
Poisson or replayed traces, not "fire N requests at once."
3. STEADY STATE
Warm up. Discard the first N seconds. Measure long enough.
4. CLIENT-SIDE MEASUREMENT
Through the real ingress path.
5. THE RIGHT METRICS
TTFT and ITL percentiles separately, bucketed by prompt length.
Throughput AND cost. Goodput, not just throughput.
6. CONTROLLED ENVIRONMENT
Locked clocks, persistence mode, no other tenants, documented config.
7. REPRODUCIBLE
Fixed seeds where possible, recorded configuration, versioned.
8. HONEST REPORTING
All the numbers, including the unfavorable ones. Error bars.Most published inference benchmarks violate 1, 2, and 5. Most internal ones violate 3 and 6.
2. The workload shape (requirement 1)#
WRONG:
All prompts 1,024 tokens, all outputs 128 tokens.
→ tells you about one point in a 2-D space, and not one your traffic occupies.
RIGHT:
Sample from your measured distribution:
prompt lengths: lognormal or empirical, with a heavy tail
output lengths: lognormal, correlated with prompt length and task
Or better: REPLAY REAL TRACES (anonymized lengths, not content).// A reasonable synthetic distribution for chat
type Workload struct{ PromptTokens, OutputTokens int }
func sampleWorkload(rng *rand.Rand, n int) []Workload {
clip := func(x, lo, hi float64) int { return int(math.Max(lo, math.Min(hi, x))) }
out := make([]Workload, n)
for i := range out {
// prompt: heavy-tailed, mean ~600, p99 ~8000
prompt := math.Exp(5.8 + 1.1*rng.NormFloat64())
// output: correlated with prompt (longer questions → longer answers)
output := math.Exp(4.6 + 0.15*math.Log(prompt/600) + 0.9*rng.NormFloat64())
out[i] = Workload{clip(prompt, 10, 32768), clip(output, 1, 4096)}
}
return out
}
Record and publish the distribution you used. A benchmark without its length distribution is uninterpretable.
3. The arrival pattern (requirement 2)#
CLOSED LOOP (N concurrent clients, each sending the next request
immediately after the previous completes)
→ measures maximum throughput at a fixed concurrency
→ does NOT measure latency under realistic load
→ the system never queues, because clients self-throttle
→ USE FOR: capacity ceiling
OPEN LOOP (requests arrive at rate λ, independent of completions)
→ measures what users experience
→ reveals queueing behavior and the latency knee
→ can overload the system (that's the point)
→ USE FOR: SLO validation, finding the operating limit
TRACE REPLAY (recorded arrival timestamps)
→ most realistic; captures burstiness and diurnal patterns
→ USE FOR: final validationReport which you used. A “1,000 tok/s” number from a closed-loop test at concurrency 256 and one from an open-loop test at the SLO limit are different quantities.
The standard mistake: running closed-loop, reporting the throughput, and implying the latency is achievable at that throughput in production. It usually isn’t, because open-loop arrival creates queueing that closed-loop hides.
4. Steady state (requirement 3)#
Sources of non-steady-state behavior:
- CUDA kernel JIT loading (first use of each kernel)
- cuBLAS autotuning (first call per shape)
- allocator pool growth
- CUDA graph capture
- torch.compile
- prefix cache warming
- GPU clock ramp-up (boost clocks decay under sustained load)
- thermal equilibrium
WARMUP: at least 60 seconds of representative load, or 500 requests.
MEASURE: at least 5 minutes, or 2,000 requests, whichever is longer.The clock ramp is real and often missed. A 30-second benchmark runs at boost clocks; a 2-hour production workload runs at sustained clocks, 10-20% lower. Short benchmarks systematically overstate performance.
5. Client-side measurement (requirement 4)#
SERVER-SIDE TTFT: what the engine measured
CLIENT-SIDE TTFT: what the user experiences
The difference includes: network RTT, TLS, proxy buffering, LB overhead,
client-side parsing.
A buffering proxy makes server-side TTFT look perfect while client-side
TTFT is 20x worse (Section II.08).Always measure both and report both. If they differ by more than the network RTT, investigate.
// Correct client-side measurement: timestamps taken as each chunk ARRIVES.
type Result struct {
TTFT time.Duration
ITLs []time.Duration
PromptTokens int
}
func measure(ctx context.Context, url string, payload []byte) (Result, error) {
req, _ := http.NewRequestWithContext(ctx, "POST", url, bytes.NewReader(payload))
req.Header.Set("Content-Type", "application/json")
t0 := time.Now()
resp, err := http.DefaultClient.Do(req)
if err != nil {
return Result{}, err
}
defer resp.Body.Close()
var res Result
var prev time.Time
sc := bufio.NewScanner(resp.Body)
for sc.Scan() {
data, ok := strings.CutPrefix(sc.Text(), "data: ")
if !ok {
continue
}
if data == "[DONE]" {
break
}
now := time.Now()
if prev.IsZero() {
res.TTFT = now.Sub(t0)
} else {
res.ITLs = append(res.ITLs, now.Sub(prev))
}
prev = now
}
return res, sc.Err()
}
6. The metrics (requirement 5)#
REPORT:
TTFT: p50, p90, p95, p99, BUCKETED BY PROMPT LENGTH
ITL: p50, p90, p95, p99 (and ITL vs token position, if it varies)
E2E: bucketed by output length
Throughput: output tokens/sec, input tokens/sec, requests/sec
Goodput: requests/sec meeting BOTH TTFT and ITL SLOs
Cost: $/M input tokens, $/M output tokens
Efficiency: achieved / roofline; average running batch
ALSO REPORT:
the arrival rate or concurrency
the length distribution
the full server configuration
hardware and software versions
how long the measurement ran
the variance across runsGoodput is the metric that prevents gaming. A system can maximize throughput by making everyone slow; goodput cannot be improved that way.
Bucketing TTFT by prompt length is essential. An unbucketed p99 TTFT is dominated by your longest prompts and tells you nothing about system health.
7. Controlled environment (requirement 6)#
# Before benchmarking:
nvidia-smi -pm 1 # persistence mode
nvidia-smi -lgc 1410,1410 # lock SM clocks (find your value)
nvidia-smi -lmc <memclock> # lock memory clocks if supported
# ... run benchmark ...
nvidia-smi -rgc # reset
# Also:
# - no other processes on the GPU (nvidia-smi --query-compute-apps)
# - CPU governor set to performance
# - no other tenants on the node
# - document driver, CUDA, framework versions
Without locked clocks, run-to-run variance is 10-20% and you cannot detect a 10% improvement. This is the difference between a benchmark and an anecdote.
8. The complete harness#
// Structure of a correct benchmark.
func benchmark(ctx context.Context, url string, workload []Workload, arrivalRate float64, duration, warmup time.Duration) Report {
var (
mu sync.Mutex
results []Result
wg sync.WaitGroup
)
start := time.Now()
// OPEN LOOP: launch requests at rate λ regardless of completions.
// (A closed loop — N workers each waiting for their reply — slows down when the
// server does, and hides exactly the overload you are trying to measure.)
for i := 0; time.Since(start) < warmup+duration; i++ {
time.Sleep(time.Duration(rand.ExpFloat64() / arrivalRate * float64(time.Second)))
w, sent := workload[i%len(workload)], time.Since(start)
wg.Add(1)
go func() {
defer wg.Done()
r, err := measure(ctx, url, payloadFor(w))
if err != nil || sent < warmup { // discard warmup
return
}
r.PromptTokens = w.PromptTokens
mu.Lock()
results = append(results, r)
mu.Unlock()
}()
}
wg.Wait()
return analyze(results, duration)
}
func analyze(results []Result, duration time.Duration) Report {
good, outTokens := 0, 0
for _, r := range results {
outTokens += len(r.ITLs) + 1
if r.TTFT < ttftSLO && percentile(r.ITLs, 95) < itlSLO {
good++ // goodput counts only requests that met the SLO
}
}
return Report{
TTFTByPromptBucket: bucketPercentiles(results),
ThroughputTokS: float64(outTokens) / duration.Seconds(),
GoodputReqS: float64(good) / duration.Seconds(),
Requests: len(results),
}
}
Build this once. It is the most reusable artifact in performance work, and every comparison you make for years will run through it.
9. Reporting honestly#
GOOD REPORT:
"Llama-3-70B, vLLM 0.x.y, 8×H100 SXM (NVSwitch), TP=8, BF16,
max_num_seqs=256, max_num_batched_tokens=4096, chunked prefill on,
prefix caching on, CUDA graphs on.
Workload: 2,000 requests, prompt lengths lognormal(5.8, 1.1) clipped
[10, 32768] (p50=330, p95=2,100, p99=8,400), outputs correlated,
(p50=180, p95=890). Open-loop Poisson arrivals at 12 req/s.
Warmup 60 s, measurement 600 s. Clocks locked at 1410 MHz.
Results (3 runs, mean ± stdev):
TTFT p50: 210 ± 8 ms; p95: 780 ± 34 ms (prompts ≤ 2k tokens)
TTFT p95: 3,400 ± 210 ms (prompts > 8k tokens)
ITL p50: 22 ± 0.4 ms; p99: 61 ± 4 ms
Output throughput: 4,180 ± 90 tok/s
Goodput (TTFT<1s for ≤2k prompts, ITL p95<50ms): 10.8 req/s of 12
Cost: $0.53/M output tokens at $2.50/GPU-hour, 100% utilization
Average running batch: 71
DRAM_ACTIVE: 0.78"
BAD REPORT:
"We got 4,180 tokens/sec."The bad report is not wrong; it’s unfalsifiable and unusable. Nobody can reproduce it, compare against it, or decide anything from it.
10. Common mistakes#
Fixed-length prompts and outputs. The most common and most misleading.
Closed-loop only. Hides queueing behavior.
No warmup. Overstates by 20-50%.
Server-side measurement only. Misses ingress problems.
Unbucketed TTFT percentiles. Dominated by the length tail.
Unlocked clocks. 10-20% variance; can’t detect real changes.
One run. No error bars.
Reporting throughput without the latency at which it was achieved.
Comparing systems with different configurations. Tune both properly first.
Not reporting the workload distribution.
Benchmarking with ignore_eos and fixed output lengths, which is convenient and removes the
variance that makes scheduling hard.
11. Hands-on exercise#
A. Build the harness. Implement the benchmark from section 8. It must: support open and closed loop, sample from a configurable distribution, measure client-side, and produce the full report from section 9. This is Project 11’s foundation.
B. Show the closed/open difference. Benchmark the same system closed-loop at concurrency 64 and open-loop at the arrival rate that produces an average of 64 in flight. Compare the latency distributions. Explain the difference.
C. Show the warmup effect. Measure throughput in the first 10 seconds, 10-60 seconds, and after 5 minutes. Quantify the overstatement from a short benchmark.
D. Show the clock effect. Run the same benchmark 5 times with unlocked clocks and 5 times with locked clocks. Compare the variance.
E. Show the distribution effect. Benchmark with fixed 1,024-token prompts and with a realistic distribution having the same mean. Compare throughput and p99 latency. How different are they?
F. Reproduce a published benchmark. Take a published number for a system you can run. Can you reproduce it? What did you have to match? What was unspecified?
12. Interview questions#
- What makes an inference benchmark valid? Name five requirements.
- Closed loop vs open loop — when do you use each, and what does each hide?
- Why must TTFT percentiles be bucketed by prompt length?
- What is goodput and why is it harder to game than throughput?
- Why lock GPU clocks for benchmarking?
- Why does a fixed-length workload mislead?
- What would you require in a benchmark report before believing it?
13. Further reading#
- [REFERENCE] vLLM
benchmarks/benchmark_serving.py— a good starting point - [REFERENCE] MLPerf Inference rules — how the industry standardizes this
- [FUNDAMENTAL] Any queueing theory treatment of open vs closed systems
- [FUNDAMENTAL] Dean & Barroso, “The Tail at Scale”
- Next: 08 — GPU profiling with Nsight