Below the API

Benchmarking Inference Correctly

Advanced Advanced 1h 30m Difficulty 4/5

Prerequisites I.05, VIII.14

★ A wrong benchmark is worse than no benchmark, because it produces confident wrong decisions.


1. The requirements for a valid benchmark#

1. REALISTIC WORKLOAD SHAPE
   Real prompt and output length distributions, not fixed lengths.

2. REALISTIC ARRIVAL PATTERN
   Poisson or replayed traces, not "fire N requests at once."

3. STEADY STATE
   Warm up. Discard the first N seconds. Measure long enough.

4. CLIENT-SIDE MEASUREMENT
   Through the real ingress path.

5. THE RIGHT METRICS
   TTFT and ITL percentiles separately, bucketed by prompt length.
   Throughput AND cost. Goodput, not just throughput.

6. CONTROLLED ENVIRONMENT
   Locked clocks, persistence mode, no other tenants, documented config.

7. REPRODUCIBLE
   Fixed seeds where possible, recorded configuration, versioned.

8. HONEST REPORTING
   All the numbers, including the unfavorable ones. Error bars.

Most published inference benchmarks violate 1, 2, and 5. Most internal ones violate 3 and 6.


2. The workload shape (requirement 1)#

WRONG:
  All prompts 1,024 tokens, all outputs 128 tokens.
  → tells you about one point in a 2-D space, and not one your traffic occupies.

RIGHT:
  Sample from your measured distribution:
    prompt lengths:  lognormal or empirical, with a heavy tail
    output lengths:  lognormal, correlated with prompt length and task
    
  Or better: REPLAY REAL TRACES (anonymized lengths, not content).
// A reasonable synthetic distribution for chat
type Workload struct{ PromptTokens, OutputTokens int }

func sampleWorkload(rng *rand.Rand, n int) []Workload {
	clip := func(x, lo, hi float64) int { return int(math.Max(lo, math.Min(hi, x))) }
	out := make([]Workload, n)
	for i := range out {
		// prompt: heavy-tailed, mean ~600, p99 ~8000
		prompt := math.Exp(5.8 + 1.1*rng.NormFloat64())
		// output: correlated with prompt (longer questions → longer answers)
		output := math.Exp(4.6 + 0.15*math.Log(prompt/600) + 0.9*rng.NormFloat64())
		out[i] = Workload{clip(prompt, 10, 32768), clip(output, 1, 4096)}
	}
	return out
}

Record and publish the distribution you used. A benchmark without its length distribution is uninterpretable.


3. The arrival pattern (requirement 2)#

CLOSED LOOP (N concurrent clients, each sending the next request
             immediately after the previous completes)
  → measures maximum throughput at a fixed concurrency
  → does NOT measure latency under realistic load
  → the system never queues, because clients self-throttle
  → USE FOR: capacity ceiling

OPEN LOOP (requests arrive at rate λ, independent of completions)
  → measures what users experience
  → reveals queueing behavior and the latency knee
  → can overload the system (that's the point)
  → USE FOR: SLO validation, finding the operating limit

TRACE REPLAY (recorded arrival timestamps)
  → most realistic; captures burstiness and diurnal patterns
  → USE FOR: final validation

Report which you used. A “1,000 tok/s” number from a closed-loop test at concurrency 256 and one from an open-loop test at the SLO limit are different quantities.

The standard mistake: running closed-loop, reporting the throughput, and implying the latency is achievable at that throughput in production. It usually isn’t, because open-loop arrival creates queueing that closed-loop hides.


4. Steady state (requirement 3)#

Sources of non-steady-state behavior:
  - CUDA kernel JIT loading (first use of each kernel)
  - cuBLAS autotuning (first call per shape)
  - allocator pool growth
  - CUDA graph capture
  - torch.compile
  - prefix cache warming
  - GPU clock ramp-up (boost clocks decay under sustained load)
  - thermal equilibrium

WARMUP: at least 60 seconds of representative load, or 500 requests.
MEASURE: at least 5 minutes, or 2,000 requests, whichever is longer.

The clock ramp is real and often missed. A 30-second benchmark runs at boost clocks; a 2-hour production workload runs at sustained clocks, 10-20% lower. Short benchmarks systematically overstate performance.


5. Client-side measurement (requirement 4)#

SERVER-SIDE TTFT:  what the engine measured
CLIENT-SIDE TTFT:  what the user experiences

The difference includes: network RTT, TLS, proxy buffering, LB overhead,
                         client-side parsing.

A buffering proxy makes server-side TTFT look perfect while client-side
TTFT is 20x worse (Section II.08).

Always measure both and report both. If they differ by more than the network RTT, investigate.

// Correct client-side measurement: timestamps taken as each chunk ARRIVES.
type Result struct {
	TTFT         time.Duration
	ITLs         []time.Duration
	PromptTokens int
}

func measure(ctx context.Context, url string, payload []byte) (Result, error) {
	req, _ := http.NewRequestWithContext(ctx, "POST", url, bytes.NewReader(payload))
	req.Header.Set("Content-Type", "application/json")
	t0 := time.Now()
	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return Result{}, err
	}
	defer resp.Body.Close()

	var res Result
	var prev time.Time
	sc := bufio.NewScanner(resp.Body)
	for sc.Scan() {
		data, ok := strings.CutPrefix(sc.Text(), "data: ")
		if !ok {
			continue
		}
		if data == "[DONE]" {
			break
		}
		now := time.Now()
		if prev.IsZero() {
			res.TTFT = now.Sub(t0)
		} else {
			res.ITLs = append(res.ITLs, now.Sub(prev))
		}
		prev = now
	}
	return res, sc.Err()
}

6. The metrics (requirement 5)#

REPORT:
  TTFT:       p50, p90, p95, p99, BUCKETED BY PROMPT LENGTH
  ITL:        p50, p90, p95, p99  (and ITL vs token position, if it varies)
  E2E:        bucketed by output length
  Throughput: output tokens/sec, input tokens/sec, requests/sec
  Goodput:    requests/sec meeting BOTH TTFT and ITL SLOs
  Cost:       $/M input tokens, $/M output tokens
  Efficiency: achieved / roofline; average running batch

ALSO REPORT:
  the arrival rate or concurrency
  the length distribution
  the full server configuration
  hardware and software versions
  how long the measurement ran
  the variance across runs

Goodput is the metric that prevents gaming. A system can maximize throughput by making everyone slow; goodput cannot be improved that way.

Bucketing TTFT by prompt length is essential. An unbucketed p99 TTFT is dominated by your longest prompts and tells you nothing about system health.


7. Controlled environment (requirement 6)#

# Before benchmarking:
nvidia-smi -pm 1                       # persistence mode
nvidia-smi -lgc 1410,1410              # lock SM clocks (find your value)
nvidia-smi -lmc <memclock>             # lock memory clocks if supported
# ... run benchmark ...
nvidia-smi -rgc                        # reset

# Also:
#  - no other processes on the GPU (nvidia-smi --query-compute-apps)
#  - CPU governor set to performance
#  - no other tenants on the node
#  - document driver, CUDA, framework versions

Without locked clocks, run-to-run variance is 10-20% and you cannot detect a 10% improvement. This is the difference between a benchmark and an anecdote.


8. The complete harness#

// Structure of a correct benchmark.
func benchmark(ctx context.Context, url string, workload []Workload, arrivalRate float64, duration, warmup time.Duration) Report {
	var (
		mu      sync.Mutex
		results []Result
		wg      sync.WaitGroup
	)
	start := time.Now()

	// OPEN LOOP: launch requests at rate λ regardless of completions.
	// (A closed loop — N workers each waiting for their reply — slows down when the
	// server does, and hides exactly the overload you are trying to measure.)
	for i := 0; time.Since(start) < warmup+duration; i++ {
		time.Sleep(time.Duration(rand.ExpFloat64() / arrivalRate * float64(time.Second)))
		w, sent := workload[i%len(workload)], time.Since(start)
		wg.Add(1)
		go func() {
			defer wg.Done()
			r, err := measure(ctx, url, payloadFor(w))
			if err != nil || sent < warmup { // discard warmup
				return
			}
			r.PromptTokens = w.PromptTokens
			mu.Lock()
			results = append(results, r)
			mu.Unlock()
		}()
	}
	wg.Wait()
	return analyze(results, duration)
}

func analyze(results []Result, duration time.Duration) Report {
	good, outTokens := 0, 0
	for _, r := range results {
		outTokens += len(r.ITLs) + 1
		if r.TTFT < ttftSLO && percentile(r.ITLs, 95) < itlSLO {
			good++ // goodput counts only requests that met the SLO
		}
	}
	return Report{
		TTFTByPromptBucket: bucketPercentiles(results),
		ThroughputTokS:     float64(outTokens) / duration.Seconds(),
		GoodputReqS:        float64(good) / duration.Seconds(),
		Requests:           len(results),
	}
}

Build this once. It is the most reusable artifact in performance work, and every comparison you make for years will run through it.


9. Reporting honestly#

GOOD REPORT:
  "Llama-3-70B, vLLM 0.x.y, 8×H100 SXM (NVSwitch), TP=8, BF16,
   max_num_seqs=256, max_num_batched_tokens=4096, chunked prefill on,
   prefix caching on, CUDA graphs on.
   
   Workload: 2,000 requests, prompt lengths lognormal(5.8, 1.1) clipped
   [10, 32768] (p50=330, p95=2,100, p99=8,400), outputs correlated,
   (p50=180, p95=890). Open-loop Poisson arrivals at 12 req/s.
   Warmup 60 s, measurement 600 s. Clocks locked at 1410 MHz.
   
   Results (3 runs, mean ± stdev):
     TTFT p50: 210 ± 8 ms;  p95: 780 ± 34 ms  (prompts ≤ 2k tokens)
     TTFT p95: 3,400 ± 210 ms (prompts > 8k tokens)
     ITL p50: 22 ± 0.4 ms;  p99: 61 ± 4 ms
     Output throughput: 4,180 ± 90 tok/s
     Goodput (TTFT<1s for ≤2k prompts, ITL p95<50ms): 10.8 req/s of 12
     Cost: $0.53/M output tokens at $2.50/GPU-hour, 100% utilization
     Average running batch: 71
     DRAM_ACTIVE: 0.78"

BAD REPORT:
  "We got 4,180 tokens/sec."

The bad report is not wrong; it’s unfalsifiable and unusable. Nobody can reproduce it, compare against it, or decide anything from it.


10. Common mistakes#

Fixed-length prompts and outputs. The most common and most misleading.

Closed-loop only. Hides queueing behavior.

No warmup. Overstates by 20-50%.

Server-side measurement only. Misses ingress problems.

Unbucketed TTFT percentiles. Dominated by the length tail.

Unlocked clocks. 10-20% variance; can’t detect real changes.

One run. No error bars.

Reporting throughput without the latency at which it was achieved.

Comparing systems with different configurations. Tune both properly first.

Not reporting the workload distribution.

Benchmarking with ignore_eos and fixed output lengths, which is convenient and removes the variance that makes scheduling hard.


11. Hands-on exercise#

A. Build the harness. Implement the benchmark from section 8. It must: support open and closed loop, sample from a configurable distribution, measure client-side, and produce the full report from section 9. This is Project 11’s foundation.

B. Show the closed/open difference. Benchmark the same system closed-loop at concurrency 64 and open-loop at the arrival rate that produces an average of 64 in flight. Compare the latency distributions. Explain the difference.

C. Show the warmup effect. Measure throughput in the first 10 seconds, 10-60 seconds, and after 5 minutes. Quantify the overstatement from a short benchmark.

D. Show the clock effect. Run the same benchmark 5 times with unlocked clocks and 5 times with locked clocks. Compare the variance.

E. Show the distribution effect. Benchmark with fixed 1,024-token prompts and with a realistic distribution having the same mean. Compare throughput and p99 latency. How different are they?

F. Reproduce a published benchmark. Take a published number for a system you can run. Can you reproduce it? What did you have to match? What was unspecified?


12. Interview questions#

  1. What makes an inference benchmark valid? Name five requirements.
  2. Closed loop vs open loop — when do you use each, and what does each hide?
  3. Why must TTFT percentiles be bucketed by prompt length?
  4. What is goodput and why is it harder to game than throughput?
  5. Why lock GPU clocks for benchmarking?
  6. Why does a fixed-length workload mislead?
  7. What would you require in a benchmark report before believing it?

13. Further reading#

  • [REFERENCE] vLLM benchmarks/benchmark_serving.py — a good starting point
  • [REFERENCE] MLPerf Inference rules — how the industry standardizes this
  • [FUNDAMENTAL] Any queueing theory treatment of open vs closed systems
  • [FUNDAMENTAL] Dean & Barroso, “The Tail at Scale”
  • Next: 08 — GPU profiling with Nsight

↑↓ navigate ↵ open