Section V covered what continuous batching is. This file covers the server-side knobs: what they control, how to tune them, and what happens when you get them wrong.
1. The knobs#
Every continuous-batching engine exposes roughly these:
max_num_seqs ceiling on concurrent sequences
max_num_batched_tokens ceiling on tokens processed per iteration
max_model_len per-request context ceiling
gpu_memory_utilization fraction of GPU claimed for the KV pool
enable_chunked_prefill split long prefills across steps
scheduling_policy FCFS / priority
preemption_mode recompute / swapThey interact. Understanding the interactions is the difference between a tuned and an untuned deployment.
2. What each one actually does#
max_num_seqs#
The maximum number of sequences in the running set.
Too low: you leave throughput on the table; the GPU runs at a smaller batch
than memory allows.
Too high: KV memory runs out mid-flight → preemption thrashing; and ITL
rises past your SLO.
Set it to: min(memory-allowed concurrency at your p95 context,
latency-allowed concurrency at your ITL SLO)Compute the first from Section V.06. Measure the second: increase batch until ITL exceeds your SLO.
A common failure: setting it to the memory maximum. Then a few long-context requests arrive, KV runs out, and the engine preempts continuously. Leave 20-30% headroom.
max_num_batched_tokens#
The total tokens (prefill + decode) processed in one iteration. This is the main prefill/decode balance knob.
Small (512-1024):
✓ decode steps stay fast → good ITL
✗ prefill is slow (many small chunks) → worse TTFT
✗ prefill can't reach the compute-bound regime
Large (8192-32768):
✓ prefill throughput is high → good TTFT
✗ a step containing a big prefill chunk is slow → ITL spikes
Typical good values: 2048-8192, depending on your ITL SLO.Rule of thumb:
max_num_batched_tokens ≈ (ITL_SLO_ms / prefill_ms_per_1k_tokens) × 1000For a model doing 30 ms per 1,000 prefill tokens with a 60 ms ITL SLO:
(60/30) × 1000 = 2000. Start there, then measure.
gpu_memory_utilization#
The fraction of GPU memory vLLM (or equivalent) claims at startup for weights + KV pool + graphs.
0.85 conservative; leaves room for fragmentation and activation spikes
0.90 the usual default
0.95 more KV cache, but a long prefill's activation spike can OOMTest with your longest realistic prompt before raising it. The activation peak during a 128k-token prefill is much larger than during decode, and the memory profiler that sizes the KV pool may not have seen it.
enable_chunked_prefill#
Splits prefill into chunks that ride along with decode steps.
ON: ITL is smooth; TTFT for long prompts slightly worse; throughput slightly better
OFF: long prefills block all decodes; ITL spikesTurn it on if you have any long prompts and any ITL SLO. Default in recent vLLM versions.
3. The interactions#
max_num_seqs ↑ → more KV needed → may need gpu_memory_utilization ↑
→ or max_model_len ↓
max_num_batched_tokens ↑ → better prefill throughput
→ worse ITL for concurrent decodes
→ more activation memory (may need lower gpu_mem_util)
max_model_len ↑ → larger per-sequence block tables
→ fewer sequences fit
→ larger worst-case activation
chunked_prefill ON → allows a smaller max_num_batched_tokens without hurting
prefill throughput as muchThe dependency that surprises people: max_model_len affects capacity even for short
requests, because block tables are sized for the maximum and the memory profiler reserves for
the worst case. Setting it to the model’s 128k maximum when your traffic is 4k can halve your
usable concurrency.
4. Tuning procedure#
1. MEASURE your traffic: p50/p95/p99 prompt length, output length, arrival rate.
2. SET max_model_len to p99 prompt + p99 output, plus ~20%. Not the model max.
3. COMPUTE memory-allowed concurrency (Section V.06) at p95 context.
Set max_num_seqs to 70-80% of that.
4. SET max_num_batched_tokens from the ITL formula above. Enable chunked prefill.
5. LOAD TEST at your expected peak. Measure p95 TTFT, p95 ITL, throughput,
preemption rate, and running batch size.
6. ADJUST:
preemption rate > 1% → reduce max_num_seqs
running batch << max → you're traffic-limited, not config-limited
ITL p95 > SLO → reduce max_num_batched_tokens
TTFT p95 > SLO with low
queue_wait → increase max_num_batched_tokens
TTFT p95 > SLO with high
queue_wait → you need more capacity, not tuning
7. REPEAT at 1.5x expected peak to verify headroom.Step 6’s last two lines are the important distinction: tuning cannot fix insufficient capacity. Separate the two before spending days on config.
5. Worked example#
Workload: chat, p95 prompt 3,000 tokens, p95 output 600 tokens, peak 40 req/s
Model: Llama-3-8B, one H100
SLO: p95 TTFT < 800 ms, p95 ITL < 50 ms
1. max_model_len = 3000 + 600 + 20% = 4,400 → round to 4,608
2. KV per token = 128 KiB
Budget = 80 − 16.1 (weights) − 2 (graphs) − 3 (headroom) = 58.9 GB
Per sequence at 4,608 = 0.60 GB
Memory-allowed concurrency = 98
→ max_num_seqs = 72 (73% of the maximum)
3. Prefill rate ≈ 45,000 tok/s → 22 ms per 1,000 tokens
max_num_batched_tokens ≈ (50/22) × 1000 = 2,270 → use 2,048
enable_chunked_prefill = True
4. Check throughput:
decode step bytes = 16.1 + 72 × 4608 × 128 KiB = 16.1 + 42.5 = 58.6 GB
T_step = 58.6/3350 = 17.5 ms theoretical, ~25 ms real
→ ITL 25 ms ✓ (SLO 50)
→ throughput = 72/0.025 = 2,880 tok/s
→ at 600 output tokens per request: 4.8 req/s completing
✗ PEAK IS 40 req/s — we need ~8 replicas.
5. Revised: 8 replicas of this configuration, plus headroom → 10-11.Note that the tuning was fine and the capacity was the problem. That’s the common case, and the arithmetic finds it in ten minutes.
6-9. Under the hood, performance, production, mistakes#
Under the hood — what the engine logs at startup:
INFO: # GPU blocks: 30,131, # CPU blocks: 2,048
INFO: Maximum concurrency for 4608 tokens per request: 104.6x
INFO: Capturing CUDA graphs: 100%|...| 35/35 [00:24<00:00]
INFO: Graph capturing finished in 24 secs, took 1.42 GiBRead these lines every time. # GPU blocks × block_size ÷ max_model_len is your
memory-allowed concurrency, stated by the engine. If it’s much lower than you expected, check
max_model_len and gpu_memory_utilization.
Performance — the effect of each knob (8B model, H100, chat workload):
Configuration Throughput p95 TTFT p95 ITL
max_num_seqs=16 1,100 tok/s 340 ms 19 ms
max_num_seqs=72 2,880 610 ms 25 ms
max_num_seqs=98 (memory max) 2,910 1,200 ms 34 ms ← preemption
max_num_batched_tokens=512 2,650 1,400 ms 21 ms
max_num_batched_tokens=8192 3,050 480 ms 68 ms ← ITL SLO broken
chunked_prefill=off 2,900 590 ms 95 ms ← ITL spikesEvery row is a real tradeoff. The middle configuration is the one that meets both SLOs.
Production:
- Tune per workload. A RAG service and a chat service want different settings for the same model.
- Re-tune when traffic changes. Prompt lengths drift as products evolve.
- Monitor running batch size and preemption rate as your primary tuning signals.
- Document the configuration and the reasoning. Six months later nobody remembers why
max_num_batched_tokensis 2048. - Set
max_model_lenfrom traffic, not from the model card.
Mistakes:
max_num_seqsat the memory maximum. Preemption thrashing.max_model_lenat the model’s maximum. Halves usable concurrency.gpu_memory_utilizationat 0.97. OOM on a long prefill.- Chunked prefill off with long prompts and an ITL SLO.
- Tuning to fix a capacity problem.
- Copying someone else’s configuration for a different workload.
10. Hands-on exercise#
A. Run the tuning procedure. For a model and workload you have, execute all seven steps. Document each decision and its justification.
B. Sweep the knobs. Reproduce the table in section 7 on your hardware: sweep max_num_seqs
and max_num_batched_tokens independently, measuring throughput, p95 TTFT, and p95 ITL. Plot
the Pareto frontier.
C. Induce preemption. Set max_num_seqs to the memory maximum and send long-context
requests. Observe the preemption counter and the latency impact.
D. max_model_len cost. Start the same server with max_model_len = 4096 and 131072.
Compare the reported GPU block counts and the implied concurrency.
E. Chunked prefill. Measure p95 ITL with chunked prefill on and off, with a workload mixing short streaming requests and occasional 32k prompts.
11. Interview questions#
- What does
max_num_batched_tokenscontrol and how would you choose it? - Why shouldn’t
max_num_seqsbe set to the memory maximum? - How does
max_model_lenaffect capacity even for short requests? - How do you distinguish a tuning problem from a capacity problem?
- What does the engine’s startup “# GPU blocks” line tell you?
- When would you disable chunked prefill?
- Walk me through tuning a server for a chat workload with a 50 ms ITL SLO.
12. Further reading#
- [REFERENCE] vLLM engine arguments documentation
- [ESTABLISHED] Agrawal et al., “Sarathi-Serve” (2024) — the chunked prefill tradeoff, analyzed
- [REFERENCE] TensorRT-LLM
max_num_tokenstuning guidance - Next: 05 — Backpressure, timeouts, retries