1. What is it?#
The components every LLM inference server has, and how they fit together.
┌──────────────────────────────────────┐
HTTP/gRPC ────► │ API LAYER │
│ parse, validate, auth, rate limit │
│ chat template, tokenize │
└──────────────┬───────────────────────┘
│ Request objects
┌──────────────▼───────────────────────┐
│ SCHEDULER │
│ waiting queue / running set │
│ admission, preemption, priorities │
│ prefix cache lookup │
└──────────────┬───────────────────────┘
│ batch metadata
┌──────────────▼───────────────────────┐
│ KV CACHE MANAGER │
│ block pool, block tables, refcounts │
│ prefix hash table, eviction │
└──────────────┬───────────────────────┘
│
┌──────────────▼───────────────────────┐
│ MODEL EXECUTOR │
│ forward pass, CUDA graphs │
│ distributed communication (TP/PP) │
└──────────────┬───────────────────────┘
│ logits
┌──────────────▼───────────────────────┐
│ SAMPLER │
│ penalties, temperature, top-k/p │
│ stop conditions, constrained decode │
└──────────────┬───────────────────────┘
│ token ids
┌──────────────▼───────────────────────┐
│ DETOKENIZER + STREAM │
│ incremental decode, SSE framing │
└──────────────────────────────────────┘Every production engine has exactly these six components. The differences are in the policies inside each.
Diagram — The request path through a serving system#
flowchart LR
C["Client"] --> GW["Gateway<br/>auth, limits, routing"]
GW --> API["API server<br/>validate, tokenize, chat template"]
API --> Q["Queue<br/>admission control"]
Q --> SCH["Scheduler<br/>continuous batching"]
SCH --> EX["Model executor<br/>forward pass on GPU"]
EX <--> KV[("KV cache manager")]
EX --> SMP["Sampler"]
SMP --> DET["Detokenize + stream"]
DET -->|"SSE"| C
SMP -.->|"next step"| SCH
class GW,API,DET io
class Q,SCH queue
class EX,SMP compute
class KV memory
class C neutral2. Why it’s structured this way#
Because of the constraints established in Section V:
CONSTRAINT → COMPONENT
Requests have unknown, variable cost → scheduler with per-step decisions
Memory is the binding capacity limit → KV cache manager owns admission
Decode is memory-bound → batching is the executor's job
Tokens must be streamed → detokenizer is incremental and stateful
CPU work competes with GPU driving → API layer separated from the engine loop
Requests are long-lived and stateful → per-request state lives with the schedulerThe architecture is a direct consequence of the physics. That’s why every engine converged on roughly the same shape.
3. Simple analogy#
A hospital emergency department.
- API layer — reception: check in, verify insurance, take vitals.
- Scheduler — triage: who goes next, who waits, who is turned away.
- KV cache manager — bed management: how many beds, who occupies which, when they free.
- Executor — the treatment rooms: where work actually happens.
- Sampler — the decision at each step: what to do next for this patient.
- Detokenizer/stream — communicating with the patient and family as things progress.
The bottleneck is beds, not doctors. Triage exists because beds are scarce. That’s exactly the KV-cache-bound situation of an LLM server.
4. Tiny example — the main loop#
Every engine’s heart:
type Engine struct {
model Model
kv *BlockManager
waiting, running []*Request
maxBatch int
maxBatchedTokens int
}
func (e *Engine) Add(r *Request) {
r.Arrival = time.Now()
e.waiting = append(e.waiting, r)
}
func (e *Engine) Step() (finished []*Request) {
// --- 1. SCHEDULE: decide what runs this iteration ---
prefill, decode := e.schedule()
if len(prefill)+len(decode) == 0 {
return nil
}
// --- 2. PREPARE: build the flat tensors and metadata ---
batch := e.prepareInputs(prefill, decode)
// --- 3. EXECUTE: one forward pass ---
logits := e.model.Forward(batch) // a CUDA graph replay if decode-only
// --- 4. SAMPLE ---
tokens := e.sample(logits, batch.Sampling)
// --- 5. UPDATE: append tokens, grow KV, detect completion ---
finished = e.update(append(prefill, decode...), tokens)
// --- 6. FREE ---
for _, r := range finished {
e.kv.Free(r.BlockTable)
e.running = slices.DeleteFunc(e.running, func(x *Request) bool { return x == r })
}
return finished
}
func (e *Engine) schedule() (prefill, decode []*Request) {
// admit new requests while memory and the token budget allow
decode = slices.Clone(e.running)
budget := e.maxBatchedTokens - len(decode) // decode uses 1 token each
for len(e.waiting) > 0 && len(e.running) < e.maxBatch {
r := e.waiting[0]
need := min(r.RemainingPromptTokens(), budget)
if need <= 0 || !e.kv.CanAllocate(e.kv.BlocksNeeded(need)) {
break
}
e.waiting = e.waiting[1:]
r.BlockTable = e.kv.AllocateWithPrefixCache(r)
prefill, e.running = append(prefill, r), append(e.running, r)
budget -= need
}
return prefill, decode
}
That’s a working continuous-batching engine. Real ones add: chunked prefill (splitting a
prefill across steps), preemption, priorities, multi-LoRA, speculative decoding, and careful
optimization of prepare_inputs — but the loop is this.
5. Technical explanation of each component#
API layer#
Responsibilities:
- Protocol handling (HTTP/1.1 + SSE, HTTP/2, gRPC)
- Auth, rate limiting, quota
- Request validation (max lengths, parameter ranges)
- Chat template application
- Tokenization (thread pool, releases the GIL)
- Response streaming and detokenization
- Metrics emission
Design notes:
- Runs async (asyncio/tokio) — thousands of long-lived connections
- SEPARATE PROCESS from the engine, usually. Why? The GIL: HTTP parsing
and tokenization must not block the engine's kernel-launch loop.
- Communicates with the engine over ZeroMQ, shared memory, or an async queueScheduler#
State:
waiting: deque of admitted-but-not-started requests
running: list of requests with allocated KV
swapped: preempted requests (if using swap-based preemption)
Per-step decisions:
1. Which waiting requests to admit
2. How many prefill tokens to process (chunked prefill budget)
3. Whether to preempt anyone
4. Batch composition
Policies (this is where engines differ):
FCFS, priority, fair-share, shortest-first
prefill-priority vs decode-priority
chunked vs whole prefillKV cache manager#
block pool: preallocated GPU memory, divided into fixed blocks
free list: available block ids
block tables: per-sequence logical→physical mapping
refcounts: for copy-on-write sharing
prefix hash table: content hash → block id, for prefix caching
eviction: LRU over refcount-0 cached blocksThis component owns capacity. Admission is really “does the KV manager have room?”
Model executor#
- Owns the weights and the KV cache tensors
- Builds input tensors from the scheduler's metadata
- Runs the forward pass (eager or CUDA graph replay)
- Handles distributed execution (TP ranks, PP stages)
- Returns logits
In a TP setup: one executor process per rank, all running in lockstep,
communicating via NCCL. The scheduler runs on rank 0 and broadcasts decisions.Sampler#
Fully on GPU (Section V.13). Handles per-request parameters via
vectorized per-row operations.
Also: stop-string detection, constrained decoding masks, logprob extraction.Detokenizer#
Stateful per request (Section V.02). Must handle multi-byte UTF-8
spanning tokens, and hold back partial stop sequences.
Usually runs in the API process, not the engine process.6. Under the hood — the process topology#
A typical vLLM deployment with TP=4:
Process: API server (Python, asyncio)
├─ uvicorn worker(s)
├─ tokenizer thread pool
└─ ZeroMQ ──┐
│
Process: Engine core (rank 0) ← scheduler lives here
├─ scheduler
├─ KV manager
├─ model executor (rank 0)
└─ NCCL ────┬──────┬──────┐
│ │ │
Process: rank 1 rank 2 rank 3 ← executors only, driven by rank 0Why separate processes rather than threads?
- The GIL. The engine loop must not be blocked by HTTP work.
- TP ranks need separate CUDA contexts and separate NCCL ranks.
- Fault isolation — an API-layer crash shouldn’t take down the model.
Cost: IPC serialization on every request and response. Engines optimize this heavily (shared memory for large payloads, msgpack rather than pickle).
7-9. Performance, production, mistakes#
Performance — where CPU time goes in the engine process, at batch 128:
prepare_inputs (building tensors, block tables) 30-45%
scheduler bookkeeping 15-25%
sampling metadata construction 10-15%
kernel launches (with CUDA graphs) 5-10%
output processing 10-15%Note that almost none of it is kernel launching once graphs are on. The bottleneck shifts to Python-side data structure manipulation, which is why engines have moved scheduler internals to C++/Rust or heavily optimized the Python (flat arrays, cached tensors, incremental updates).
Production:
- Monitor each component separately. Queue depth (scheduler), KV utilization (cache manager), step time (executor), CPU per process.
- The API process and engine process have different scaling characteristics. You may need several API workers per engine.
prepare_inputscost scales with batch size — at very large batch it can dominate. Profile it.- Fault isolation matters: an engine crash should be detected and the process restarted, with in-flight requests failed cleanly rather than hanging.
Mistakes:
- Running the API and engine in one Python process. GIL contention destroys ITL.
- Not bounding the waiting queue. Unbounded memory growth and requests that complete after the client left.
- Blocking the async event loop in the API layer.
- Treating the engine as a black box. When it’s slow, you need to know which component.
10. Hands-on exercise#
A. Build it. Implement the six components as separate classes, with the main loop from section 4. Use a small model. Support: streaming, per-request sampling params, and stop conditions. This is Project 06 and the foundation for 07-09.
B. Map a real engine. Read vLLM’s source and locate each of the six components. Draw the call graph for one request from HTTP arrival to first token.
C. Measure the components. Instrument a real engine to report per-component time per step. Where does CPU time actually go at batch 8 vs batch 256?
D. Process topology. For a running vLLM server with TP>1, list the processes and threads
(ps -eLf, py-spy dump on each). Draw the topology. Which process would you profile for an
ITL problem? For a TTFT problem?
11. Interview questions#
- Name the six components of an inference server and what each owns.
- Why are the API layer and the engine separate processes?
- Which component owns capacity, and why?
- Where does CPU time go in the engine process at high batch?
- What is
prepare_inputsand why does it matter? - In a TP=4 deployment, where does the scheduler run and how do the other ranks know what to do?
- What breaks if you run the API and engine in one Python process?
12. Further reading#
- [REFERENCE] vLLM source:
entrypoints/openai/api_server.py,engine/llm_engine.py,core/scheduler.py,worker/model_runner.py - [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — the architecture that started this
- [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023) §5 — vLLM’s system design
- Next: 02 — APIs