Below the API

vLLM Architecture

Intermediate Advanced 1h 30m Difficulty 4/5

Prerequisites V.09, V.10, 01, 03

★ Read this file with the source open. vLLM is the reference implementation of everything in Section V, and reading it is the fastest way to make those concepts concrete.


1. What is it?#

An open-source LLM inference engine that introduced PagedAttention and popularized continuous batching. It is the de facto default for self-hosted LLM serving.

Design goals, in priority order:
  1. Throughput (tokens/sec/GPU)
  2. Memory efficiency (concurrency per GB)
  3. Broad model support
  4. Ease of use (OpenAI-compatible out of the box)

Note what’s not first: single-request latency. vLLM optimizes for serving many users. For absolute lowest latency at batch 1, TensorRT-LLM often wins.

Diagram — vLLM at a glance#

flowchart TB
  subgraph FE["API server process"]
    HTTP["OpenAI-compatible HTTP"] --> PRE["Input processing<br/>tokenize, templates, multimodal"]
    OUTP["Output processing<br/>detokenize, stream"]
  end
  subgraph CORE["Engine core"]
    SCH["Scheduler<br/>waiting and running queues"] --> KVM["KV cache manager<br/>block pool, prefix hashes"]
    SCH --> EXE["Executor"]
  end
  subgraph WK["GPU workers"]
    MR["Model runner<br/>build batch, forward, sample"] --> ATT["Attention backend<br/>paged KV"]
  end
  PRE -->|"IPC"| SCH
  EXE --> MR
  MR -->|"sampled tokens"| SCH
  SCH -->|"IPC"| OUTP

  class HTTP,PRE,OUTP io
  class SCH,EXE queue
  class KVM memory
  class MR,ATT compute

2. The architecture#

┌──────────────────────────────────────────────────────────────┐
│ API SERVER PROCESS (Python, asyncio)                         │
│   OpenAI-compatible routes                                    │
│   tokenizer pool, detokenizer                                 │
│   request validation, chat templates                          │
└───────────────────────┬──────────────────────────────────────┘
                        │ ZeroMQ / msgpack IPC
┌───────────────────────▼──────────────────────────────────────┐
│ ENGINE CORE PROCESS                                           │
│  ┌────────────────────────────────────────────────────────┐  │
│  │ SCHEDULER                                               │  │
│  │   waiting deque / running list                          │  │
│  │   budget: max_num_seqs, max_num_batched_tokens          │  │
│  │   chunked prefill, preemption, priorities               │  │
│  └───────────────────┬────────────────────────────────────┘  │
│  ┌───────────────────▼────────────────────────────────────┐  │
│  │ KV CACHE MANAGER (BlockManager)                         │  │
│  │   block pool, free list, refcounts                      │  │
│  │   prefix cache hash table, LRU eviction                 │  │
│  │   copy-on-write for forked sequences                    │  │
│  └───────────────────┬────────────────────────────────────┘  │
│  ┌───────────────────▼────────────────────────────────────┐  │
│  │ MODEL EXECUTOR / WORKER                                 │  │
│  │   input preparation (flat tensors, block tables)        │  │
│  │   CUDA graph replay (decode) / eager (prefill)          │  │
│  │   attention backend (FlashAttention / FlashInfer)       │  │
│  └───────────────────┬────────────────────────────────────┘  │
│  ┌───────────────────▼────────────────────────────────────┐  │
│  │ SAMPLER (GPU)                                           │  │
│  └────────────────────────────────────────────────────────┘  │
└───────────────────────┬──────────────────────────────────────┘
                        │ NCCL (if TP > 1)
              ┌─────────┴─────────┬─────────┐
          worker rank 1      rank 2     rank 3

3. The files to read, in order#

1. vllm/engine/llm_engine.py            the main step() loop — START HERE
2. vllm/core/scheduler.py               scheduling policy, budgets, preemption
3. vllm/core/block_manager.py           (or block/ dir) paged allocation, prefix cache
4. vllm/worker/model_runner.py          input preparation, CUDA graph capture
5. vllm/attention/backends/             the attention kernel dispatch
6. csrc/attention/attention_kernels.cu  the paged attention CUDA kernel
7. csrc/layernorm_kernels.cu            a small, readable fused kernel
8. vllm/model_executor/layers/sampler.py  GPU sampling
9. vllm/entrypoints/openai/api_server.py  the API layer

File 1 is 500 lines and contains the whole design. Read it before anything else.

(Directory layout shifts between versions; the concepts don’t. Find the equivalents in your version.)


4. The key design decisions#

Decision: paged KV cache#

Covered in Section V.10. The consequences that show up throughout the codebase:

  • block_tables appear in every kernel signature.
  • Memory is preallocated at startup; the pool never grows.
  • Sequences can share blocks (refcounts), enabling prefix caching and parallel sampling.
  • The scheduler’s admission check is “does the block manager have free blocks?”

Decision: iteration-level scheduling#

The scheduler runs before every forward pass, not per request. Consequences:

  • The scheduler’s own CPU cost is on the critical path (~0.5-3 ms at high batch).
  • Requests can be admitted or preempted at any step.
  • Batch composition changes constantly, which is why CUDA graphs need a capture list.

Decision: separate API and engine processes#

Because of the GIL. Consequences:

  • IPC serialization on every request/response (optimized with msgpack and shared memory for large payloads).
  • The API process can be scaled independently.
  • A crash in one doesn’t take down the other.

Decision: chunked prefill (default on in recent versions)#

Prefill is split into chunks that share iterations with decode. Consequences:

  • max_num_batched_tokens becomes the main tuning knob.
  • ITL is smooth even with long prompts.
  • Prefill and decode are mixed in one forward pass, requiring varlen kernels.

Decision: preemption by recomputation (default)#

When KV runs out, discard a sequence’s blocks and re-prefill later, rather than swapping to CPU. Because prefill is fast and PCIe is slow (Section II.09).


5. The scheduler, in detail#

# Conceptually (the real code is more elaborate)
def schedule(self):
    budget = SchedulingBudget(
        token_budget=self.max_num_batched_tokens,
        max_num_seqs=self.max_num_seqs)

    # 1. Continue running sequences (decode), preempting if memory is short
    running, preempted = self._schedule_running(budget)

    # 2. Bring back swapped sequences if there's room
    swapped_in = self._schedule_swapped(budget)

    # 3. Admit new sequences from the waiting queue
    prefills = self._schedule_prefills(budget)

    return SchedulerOutputs(
        scheduled_seq_groups=prefills + swapped_in + running,
        num_prefill_groups=len(prefills),
        num_batched_tokens=budget.num_batched_tokens,
        blocks_to_swap_in=..., blocks_to_swap_out=..., blocks_to_copy=...)

Order matters: running sequences are scheduled first (protecting in-flight work), then swapped-in, then new prefills. This prevents new admissions from starving existing requests.

The budget object enforces both max_num_seqs and max_num_batched_tokens simultaneously — a step stops adding work when either is exhausted.


6. The block manager and prefix caching#

# Conceptually
class BlockManager:
    def allocate(self, seq_group):
        num_blocks = ceil(seq.num_tokens / block_size)
        if self.enable_prefix_caching:
            # walk the chained hashes; reuse blocks that match
            for i, block_hash in enumerate(compute_chained_hashes(seq.token_ids)):
                cached = self.cached_blocks.get(block_hash)
                if cached is None: break
                cached.ref_count += 1
                block_table.append(cached)
            # allocate fresh blocks for the remainder
        ...

    def free(self, seq):
        for block in seq.block_table:
            block.ref_count -= 1
            if block.ref_count == 0:
                if block.is_cached: self.lru.add(block)   # keep for prefix reuse
                else:               self.free_list.append(block)

Freed blocks with a content hash go to an LRU cache rather than straight to the free list. That’s the whole implementation of automatic prefix caching — the paging machinery was already there.


7. Performance characteristics#

STRENGTHS
  ✓ throughput at high batch — the design target
  ✓ memory efficiency (96%+ KV utilization)
  ✓ prefix caching (automatic, free)
  ✓ multi-LoRA
  ✓ broad and fast model support — usually first to support new architectures
  ✓ easy to deploy; OpenAI-compatible

WEAKNESSES
  ✗ Python scheduler cost at very high batch (being addressed in V1 architecture)
  ✗ batch-1 latency behind TensorRT-LLM
  ✗ prefill kernel selection less aggressively tuned than TRT
  ✗ per-step CPU overhead in input preparation

The vLLM V1 engine rearchitecture addresses several of these (a leaner scheduler, better async handling, reduced per-step CPU work). Check which version you’re running; the performance characteristics differ meaningfully.


8. Production configuration#

vllm serve meta-llama/Meta-Llama-3-70B-Instruct \
  --tensor-parallel-size 8 \
  --max-model-len 8192 \
  --max-num-seqs 256 \
  --max-num-batched-tokens 4096 \
  --gpu-memory-utilization 0.90 \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --kv-cache-dtype fp8 \
  --quantization fp8 \
  --disable-log-requests \
  --served-model-name llama-3-70b \
  --port 8000

Each flag maps to something in this curriculum:

  • --tensor-parallel-size → Section IX
  • --max-model-len, --max-num-seqs, --max-num-batched-tokens → Section VIII.04
  • --gpu-memory-utilization → Section V.10
  • --enable-prefix-caching → Section V.11
  • --enable-chunked-prefill → Section XIII.05
  • --kv-cache-dtype fp8 → Section VII.12
  • --quantization fp8 → Section VII.05

Read the startup log: block count, graph capture, and the resolved configuration.


9. Common mistakes#

Not enabling prefix caching. It’s often the single largest win and it’s a flag.

--max-model-len at the model’s maximum. Halves usable concurrency.

--enforce-eager left on. Loses 25%.

Ignoring the startup log. It tells you your capacity directly.

Treating it as a black box. When it’s slow, you need to know which component.

Not reading scheduler.py. It’s 700 lines and it’s the best available documentation of how LLM serving actually works.


10. Hands-on exercise#

A. Read the loop. Read llm_engine.py’s step() and trace one request through it. Draw the sequence diagram.

B. Read the scheduler. Read scheduler.py. Answer: in what order are sequence types scheduled? What triggers preemption? How is the token budget enforced? Where would you add a priority class?

C. Read the block manager. Find the prefix-cache hash computation. Verify it chains (includes the prefix). Find where freed blocks go to the LRU.

D. Instrument it. Add timing instrumentation to step() for each phase (schedule, prepare, execute, sample, update). Run at batch 8 and batch 256. Where does CPU time go?

E. Configure and tune. Deploy vLLM for a realistic workload, apply the tuning procedure from Section VIII.04, and document your final configuration with justifications.

F. Break it deliberately. Set --max-num-seqs too high and observe preemption. Set --max-model-len too high and observe the reduced block count. Understanding failure modes is how you recognize them in production.


11. Interview questions#

  1. Describe vLLM’s process architecture and why it’s split that way.
  2. What does vLLM’s scheduler do on each iteration, and in what order?
  3. How is prefix caching implemented on top of the block manager?
  4. Why does vLLM prefer recomputation over swapping for preemption?
  5. What does --max-num-batched-tokens control and how does it interact with chunked prefill?
  6. What are vLLM’s strengths and weaknesses relative to TensorRT-LLM?
  7. What does the startup log tell you about your capacity?

12. Further reading#

  • [ESTABLISHED] Kwon et al., “Efficient Memory Management for LLM Serving with PagedAttention” (SOSP 2023)
  • [REFERENCE] vLLM documentation and source
  • [REFERENCE] vLLM V1 engine design documents
  • Next: 12 — SGLang architecture

↑↓ navigate ↵ open