1. The idea#
Run prefill and decode on separate GPU pools, and transfer the KV cache between them.
COLOCATED (standard) DISAGGREGATED
one pool does both prefill pool → KV transfer → decode pool
prefill and decode compete each pool tuned for its own bottleneck
chunked prefill mediates ┌──────────────────┐
request ─────────▶│ PREFILL POOL │
│ compute-bound │
│ high TP, big │
│ batch, FLOPs- │
│ optimized HW │
└────────┬─────────┘
│ KV cache transfer
│ (GB per request)
┌────────▼─────────┐
│ DECODE POOL │──▶ streamed tokens
│ memory-bound │
│ low TP, huge │
│ batch, bandwidth│
│ -optimized HW │
└──────────────────┘Diagram — Prefill and decode on separate workers#
sequenceDiagram
participant C as Client
participant R as Router
participant P as Prefill worker
participant D as Decode worker
C->>R: request
R->>P: prompt
Note over P: compute-bound and bursty
P->>D: KV cache + first token over RDMA, NVLink or TCP
Note over D: memory-bound and steady
loop until EOS
D-->>C: token
end
Note over P,D: Pays off only if moving the KV is faster than recomputing the prefill2. Problem → Why → Optimization#
PROBLEM Prefill and decode have opposite characteristics (Section V.03)
but share hardware and configuration. Every tuning decision is
a compromise between them.
WHY They're different workloads:
prefill: compute-bound, benefits from FLOPs and high TP
decode: memory-bound, benefits from bandwidth and huge batch
A single configuration serves neither optimally.
OPTIMIZE Separate them. Tune each pool independently. Scale each
independently.
TRADE-OFFS
✓ each pool optimally configured (TP degree, batch size, hardware)
✓ independent scaling: prefill-heavy traffic scales the prefill pool
✓ no interference: long prefills cannot affect decode ITL at all
✓ better hardware matching (FLOPs-rich for prefill, bandwidth-rich
for decode)
✗ KV TRANSFER: GB per request over the interconnect
✗ substantially more complex: two pools, transfer machinery, a
scheduler that spans both
✗ TTFT includes the transfer time
✗ layout compatibility constraints between pools
WHEN TO USE Very large scale, extreme prefill/decode ratio imbalance,
and a fast interconnect. This is a "hundreds of GPUs"
technique.
WHEN NOT TO Almost everywhere else. Chunked prefill (file 05) captures
most of the benefit for a fraction of the complexity.Be clear about this: chunked prefill is the answer for most deployments. Disaggregation is what you do when chunked prefill’s compromise is no longer good enough — which is a large-scale problem.
3. The KV transfer — the crux#
After prefill, the decode pool needs the sequence's KV cache.
SIZE
Llama-3-70B, GQA-8, FP8 KV, 4,000-token prompt:
4000 × 160 KiB = 640 MB per request
TRANSFER TIME
NVLink (same node): 640 MB / 370 GB/s = 1.7 ms
InfiniBand NDR (400G): 640 MB / 50 GB/s = 12.8 ms
RoCE 100G: 640 MB / 12 GB/s = 53 ms
Ethernet 25G (no RDMA): 640 MB / 3 GB/s = 213 ms ← unusable
COMPARE TO RECOMPUTING IT
prefill 4,000 tokens at 28,000 tok/s = 143 ms
→ transfer is CHEAPER than recompute on NVLink and InfiniBand
→ this is the same arithmetic as Section IX.11, and it's what
makes disaggregation viable at allDisaggregation requires RDMA-class interconnect. On plain Ethernet the transfer costs more than the prefill it saves.
Layer-wise streaming#
The transfer can OVERLAP with prefill:
as each layer's KV is computed, send it to the decode pool
→ by the time prefill finishes, most of the KV has arrived
→ effective transfer time ≈ the last layer's transfer only
layer 0 KV computed → send ─┐
layer 1 KV computed → send ─┤ overlapped with
... ├ continuing prefill
layer 79 KV computed → send ┘
→ transfer latency largely hiddenThis is the key implementation technique and what makes disaggregation’s TTFT competitive. Without it, TTFT = prefill + full transfer.
4. Why it wins (when it does)#
INDEPENDENT CONFIGURATION
PREFILL POOL DECODE POOL
TP degree 8 (more FLOPs) 2-4 (less comm overhead)
batch size small (a few huge (hundreds — decode
long prompts fill needs batch for intensity)
the GPU)
max_num_seqs low very high
hardware FLOPs-optimized bandwidth-optimized
(H100) (H200)
KV cache size small (transient) large (holds all decodes)
INDEPENDENT SCALING
RAG traffic surge → scale the prefill pool only
long generations → scale the decode pool only
→ in a colocated deployment you'd scale everything
NO INTERFERENCE
a 128k prefill cannot affect decode ITL, because it's on
different hardware entirely
→ p99 ITL becomes very stableREPORTED GAINS (DistServe, Splitwise, and similar)
1.5-4x goodput at the same SLO, or
tighter SLOs at the same throughput
The gain is largest when:
- the prefill/decode ratio is extreme
- SLOs are tight on both TTFT and ITL
- you have enough scale that pool-level tuning matters5. What makes it hard#
1. LAYOUT COMPATIBILITY
KV computed at TP=8 has a different head split than KV consumed
at TP=4. Either match the TP degrees (losing the independent
configuration benefit) or reshape during transfer (costly).
→ most implementations require matching TP degrees, which
removes one of the main benefits.
2. SCHEDULING ACROSS POOLS
a request occupies a prefill slot, then a decode slot.
The scheduler must reason about both, and about the transfer.
Backpressure from the decode pool must reach the prefill pool.
3. THE TRANSFER MACHINERY
RDMA setup, buffer management, flow control, error handling.
This is real systems engineering.
4. FAILURE MODES
what if the decode pool has no capacity when prefill finishes?
→ hold the KV (occupying prefill memory), or discard and retry
what if the transfer fails mid-flight?
→ the request is in an inconsistent state
5. PREFIX CACHING ACROSS POOLS
the prefix cache lives in the prefill pool. Fine.
But the decode pool's KV for a completed request could serve
as a prefix cache for a follow-up turn... in the wrong pool.
→ multi-turn conversation handling is genuinely complicated
6. COST OF THE SPLIT
two pools means two sets of model weights resident.
→ for a 70B model at FP8: 71 GB × 2 pools = 142 GB just for weights
→ this is a real cost, and it's why small deployments shouldn't
do thisPoint 6 is the one that rules it out for small deployments. Duplicating the weights across pools is only affordable when each pool is many GPUs.
6. The systems#
DistServe (2024) the academic formulation; goodput-optimal
placement of prefill and decode
Splitwise (Microsoft) phase splitting with hardware heterogeneity —
different GPU types per pool
Mooncake (Moonshot) KV-cache-centric architecture with a
disaggregated KV store as a first-class component
DeepSeek's system (per their reports) large-scale prefill/decode
separation with expert parallelism
vLLM / SGLang experimental disaggregation support via
"KV connectors"Mooncake’s framing is worth noting: rather than “prefill pool sends KV to decode pool,” they treat the KV cache as a shared distributed store that both pools access. This unifies disaggregation with prefix caching and KV tiering (Section IX.11) into one abstraction.
7. Is it right for you?#
DECISION CHECKLIST
□ Do you have RDMA-class interconnect (NVLink or InfiniBand)?
NO → stop. The transfer is too expensive.
□ Is chunked prefill already enabled and tuned?
NO → do that first. It's 90% of the benefit for 5% of the effort.
□ Are you still seeing prefill/decode interference after chunked prefill?
Measure: p99 ITL vs p50 ITL. If p99 < 2× p50, you don't have
an interference problem.
□ Is your prefill/decode ratio extreme?
RAG (10:1 input:output) or long-generation (1:20) benefit most.
Balanced workloads benefit least.
□ Are you at a scale where two pools each have many GPUs?
Weight duplication must be a small fraction of your fleet.
□ Do you have the engineering capacity for the transfer machinery,
cross-pool scheduling, and the failure modes?
ALL YES → evaluate it.
ANY NO → chunked prefill is your answer.Most readers of this file should conclude “not yet.” That’s the correct conclusion, and knowing why is the point.
8. Production implications#
- Chunked prefill first. Always. It’s the 90% solution.
- Disaggregation requires RDMA. Verify GPUDirect RDMA (Section IX.08).
- Layer-wise KV streaming is essential to hide the transfer latency.
- Matching TP degrees between pools is usually required — factor that into the benefit calculation.
- Budget for weight duplication across pools.
- The cross-pool scheduler is the hard part, not the transfer.
- Watch the field. Support in mainstream engines is improving; the calculus may change.
9. Common mistakes#
Adopting it before chunked prefill. Far more complexity for marginal additional benefit.
Without RDMA. The transfer costs more than the prefill.
Without layer-wise streaming. TTFT includes the full transfer.
Assuming independent TP degrees work. Layout compatibility usually forces matching.
Underestimating the weight duplication cost.
Not handling the “decode pool full when prefill finishes” case.
Deploying it at a scale where the complexity dominates.
10. Hands-on exercise#
A. Compute the transfer economics. For your model and interconnect, compute: KV size per request at your p50 and p95 prompt lengths, transfer time, and prefill time. Is transfer cheaper than recompute?
B. Measure your interference. With chunked prefill enabled, measure p50 and p99 ITL. Is p99 > 2× p50? If not, you don’t have the problem disaggregation solves.
C. Simulate the benefit. Build a simulator with separate prefill and decode pools, each independently configured. Compare goodput against a colocated pool with chunked prefill, for several prefill/decode ratios. Where does disaggregation win?
D. Layer-wise streaming. Implement (or simulate) layer-wise KV transfer overlapped with prefill. Measure the effective transfer latency versus the naive end-of-prefill transfer.
E. The weight duplication cost. For a deployment you know, compute the GPU cost of duplicating weights across two pools. At what fleet size does it become a small fraction?
F. Read the papers. Read DistServe and Splitwise. What assumptions does each make about the interconnect and the workload? Do they hold for you?
11. Interview questions#
- What is disaggregated serving and what problem does it solve?
- How large is the KV transfer, and when is it cheaper than recomputing?
- What is layer-wise KV streaming and why is it necessary?
- Why do the two pools usually need matching TP degrees, and what does that cost?
- What must be true before you’d consider disaggregation?
- Why is chunked prefill usually the right answer instead?
- What’s the hardest engineering part of disaggregation?
12. Further reading#
- [EMERGING] Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving” (OSDI 2024)
- [EMERGING] Patel et al., “Splitwise: Efficient Generative LLM Inference Using Phase Splitting” (ISCA 2024)
- [EMERGING] Qin et al., “Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving” (2024)
- [REFERENCE] vLLM KV connector / disaggregated prefill documentation
- Next: 07 — KV offloading and transfer