Below the API

Batching in Servers

Intermediate Advanced 1h Difficulty 3/5

Prerequisites V.08, V.09, 03

Section V covered what continuous batching is. This file covers the server-side knobs: what they control, how to tune them, and what happens when you get them wrong.


1. The knobs#

Every continuous-batching engine exposes roughly these:

max_num_seqs               ceiling on concurrent sequences
max_num_batched_tokens     ceiling on tokens processed per iteration
max_model_len              per-request context ceiling
gpu_memory_utilization     fraction of GPU claimed for the KV pool
enable_chunked_prefill     split long prefills across steps
scheduling_policy          FCFS / priority
preemption_mode            recompute / swap

They interact. Understanding the interactions is the difference between a tuned and an untuned deployment.


2. What each one actually does#

max_num_seqs#

The maximum number of sequences in the running set.

Too low:   you leave throughput on the table; the GPU runs at a smaller batch
           than memory allows.
Too high:  KV memory runs out mid-flight → preemption thrashing; and ITL
           rises past your SLO.

Set it to: min(memory-allowed concurrency at your p95 context,
               latency-allowed concurrency at your ITL SLO)

Compute the first from Section V.06. Measure the second: increase batch until ITL exceeds your SLO.

A common failure: setting it to the memory maximum. Then a few long-context requests arrive, KV runs out, and the engine preempts continuously. Leave 20-30% headroom.

max_num_batched_tokens#

The total tokens (prefill + decode) processed in one iteration. This is the main prefill/decode balance knob.

Small (512-1024):
  ✓ decode steps stay fast → good ITL
  ✗ prefill is slow (many small chunks) → worse TTFT
  ✗ prefill can't reach the compute-bound regime

Large (8192-32768):
  ✓ prefill throughput is high → good TTFT
  ✗ a step containing a big prefill chunk is slow → ITL spikes

Typical good values: 2048-8192, depending on your ITL SLO.

Rule of thumb:

max_num_batched_tokens ≈ (ITL_SLO_ms / prefill_ms_per_1k_tokens) × 1000

For a model doing 30 ms per 1,000 prefill tokens with a 60 ms ITL SLO: (60/30) × 1000 = 2000. Start there, then measure.

gpu_memory_utilization#

The fraction of GPU memory vLLM (or equivalent) claims at startup for weights + KV pool + graphs.

0.85  conservative; leaves room for fragmentation and activation spikes
0.90  the usual default
0.95  more KV cache, but a long prefill's activation spike can OOM

Test with your longest realistic prompt before raising it. The activation peak during a 128k-token prefill is much larger than during decode, and the memory profiler that sizes the KV pool may not have seen it.

enable_chunked_prefill#

Splits prefill into chunks that ride along with decode steps.

ON:   ITL is smooth; TTFT for long prompts slightly worse; throughput slightly better
OFF:  long prefills block all decodes; ITL spikes

Turn it on if you have any long prompts and any ITL SLO. Default in recent vLLM versions.


3. The interactions#

max_num_seqs ↑     → more KV needed → may need gpu_memory_utilization ↑
                                     → or max_model_len ↓

max_num_batched_tokens ↑  → better prefill throughput
                          → worse ITL for concurrent decodes
                          → more activation memory (may need lower gpu_mem_util)

max_model_len ↑    → larger per-sequence block tables
                   → fewer sequences fit
                   → larger worst-case activation

chunked_prefill ON → allows a smaller max_num_batched_tokens without hurting
                     prefill throughput as much

The dependency that surprises people: max_model_len affects capacity even for short requests, because block tables are sized for the maximum and the memory profiler reserves for the worst case. Setting it to the model’s 128k maximum when your traffic is 4k can halve your usable concurrency.


4. Tuning procedure#

1. MEASURE your traffic: p50/p95/p99 prompt length, output length, arrival rate.

2. SET max_model_len to p99 prompt + p99 output, plus ~20%. Not the model max.

3. COMPUTE memory-allowed concurrency (Section V.06) at p95 context.
   Set max_num_seqs to 70-80% of that.

4. SET max_num_batched_tokens from the ITL formula above. Enable chunked prefill.

5. LOAD TEST at your expected peak. Measure p95 TTFT, p95 ITL, throughput,
   preemption rate, and running batch size.

6. ADJUST:
   preemption rate > 1%        → reduce max_num_seqs
   running batch << max        → you're traffic-limited, not config-limited
   ITL p95 > SLO               → reduce max_num_batched_tokens
   TTFT p95 > SLO with low
     queue_wait                → increase max_num_batched_tokens
   TTFT p95 > SLO with high
     queue_wait                → you need more capacity, not tuning

7. REPEAT at 1.5x expected peak to verify headroom.

Step 6’s last two lines are the important distinction: tuning cannot fix insufficient capacity. Separate the two before spending days on config.


5. Worked example#

Workload: chat, p95 prompt 3,000 tokens, p95 output 600 tokens, peak 40 req/s
Model: Llama-3-8B, one H100
SLO: p95 TTFT < 800 ms, p95 ITL < 50 ms

1. max_model_len = 3000 + 600 + 20% = 4,400 → round to 4,608

2. KV per token = 128 KiB
   Budget = 80 − 16.1 (weights) − 2 (graphs) − 3 (headroom) = 58.9 GB
   Per sequence at 4,608 = 0.60 GB
   Memory-allowed concurrency = 98
   → max_num_seqs = 72   (73% of the maximum)

3. Prefill rate ≈ 45,000 tok/s → 22 ms per 1,000 tokens
   max_num_batched_tokens ≈ (50/22) × 1000 = 2,270 → use 2,048
   enable_chunked_prefill = True

4. Check throughput:
   decode step bytes = 16.1 + 72 × 4608 × 128 KiB = 16.1 + 42.5 = 58.6 GB
   T_step = 58.6/3350 = 17.5 ms theoretical, ~25 ms real
   → ITL 25 ms ✓ (SLO 50)
   → throughput = 72/0.025 = 2,880 tok/s
   → at 600 output tokens per request: 4.8 req/s completing
   ✗ PEAK IS 40 req/s — we need ~8 replicas.

5. Revised: 8 replicas of this configuration, plus headroom → 10-11.

Note that the tuning was fine and the capacity was the problem. That’s the common case, and the arithmetic finds it in ten minutes.


6-9. Under the hood, performance, production, mistakes#

Under the hood — what the engine logs at startup:

INFO: # GPU blocks: 30,131, # CPU blocks: 2,048
INFO: Maximum concurrency for 4608 tokens per request: 104.6x
INFO: Capturing CUDA graphs: 100%|...| 35/35 [00:24<00:00]
INFO: Graph capturing finished in 24 secs, took 1.42 GiB

Read these lines every time. # GPU blocks × block_size ÷ max_model_len is your memory-allowed concurrency, stated by the engine. If it’s much lower than you expected, check max_model_len and gpu_memory_utilization.

Performance — the effect of each knob (8B model, H100, chat workload):

Configuration                              Throughput  p95 TTFT  p95 ITL
max_num_seqs=16                             1,100 tok/s   340 ms   19 ms
max_num_seqs=72                             2,880         610 ms   25 ms
max_num_seqs=98 (memory max)                2,910       1,200 ms   34 ms  ← preemption
max_num_batched_tokens=512                  2,650       1,400 ms   21 ms
max_num_batched_tokens=8192                 3,050         480 ms   68 ms  ← ITL SLO broken
chunked_prefill=off                         2,900         590 ms   95 ms  ← ITL spikes

Every row is a real tradeoff. The middle configuration is the one that meets both SLOs.

Production:

  • Tune per workload. A RAG service and a chat service want different settings for the same model.
  • Re-tune when traffic changes. Prompt lengths drift as products evolve.
  • Monitor running batch size and preemption rate as your primary tuning signals.
  • Document the configuration and the reasoning. Six months later nobody remembers why max_num_batched_tokens is 2048.
  • Set max_model_len from traffic, not from the model card.

Mistakes:

  • max_num_seqs at the memory maximum. Preemption thrashing.
  • max_model_len at the model’s maximum. Halves usable concurrency.
  • gpu_memory_utilization at 0.97. OOM on a long prefill.
  • Chunked prefill off with long prompts and an ITL SLO.
  • Tuning to fix a capacity problem.
  • Copying someone else’s configuration for a different workload.

10. Hands-on exercise#

A. Run the tuning procedure. For a model and workload you have, execute all seven steps. Document each decision and its justification.

B. Sweep the knobs. Reproduce the table in section 7 on your hardware: sweep max_num_seqs and max_num_batched_tokens independently, measuring throughput, p95 TTFT, and p95 ITL. Plot the Pareto frontier.

C. Induce preemption. Set max_num_seqs to the memory maximum and send long-context requests. Observe the preemption counter and the latency impact.

D. max_model_len cost. Start the same server with max_model_len = 4096 and 131072. Compare the reported GPU block counts and the implied concurrency.

E. Chunked prefill. Measure p95 ITL with chunked prefill on and off, with a workload mixing short streaming requests and occasional 32k prompts.


11. Interview questions#

  1. What does max_num_batched_tokens control and how would you choose it?
  2. Why shouldn’t max_num_seqs be set to the memory maximum?
  3. How does max_model_len affect capacity even for short requests?
  4. How do you distinguish a tuning problem from a capacity problem?
  5. What does the engine’s startup “# GPU blocks” line tell you?
  6. When would you disable chunked prefill?
  7. Walk me through tuning a server for a chat workload with a 50 ms ITL SLO.

12. Further reading#

  • [REFERENCE] vLLM engine arguments documentation
  • [ESTABLISHED] Agrawal et al., “Sarathi-Serve” (2024) — the chunked prefill tradeoff, analyzed
  • [REFERENCE] TensorRT-LLM max_num_tokens tuning guidance
  • Next: 05 — Backpressure, timeouts, retries

↑↓ navigate ↵ open