PidokuInfra

Disaggregated Prefill/Decode

Expert Advanced 1h 30m Difficulty 5/5 Topic 06 of 12

Prerequisites V.03, 05, IX.11


1. The idea#

Run prefill and decode on separate GPU pools, and transfer the KV cache between them.

COLOCATED (standard)                 DISAGGREGATED
  one pool does both                   prefill pool → KV transfer → decode pool
  prefill and decode compete           each pool tuned for its own bottleneck
  chunked prefill mediates
                     ┌──────────────────┐
   request ─────────▶│  PREFILL POOL    │
                     │  compute-bound   │
                     │  high TP, big    │
                     │  batch, FLOPs-   │
                     │  optimized HW    │
                     └────────┬─────────┘
                              │ KV cache transfer
                              │ (GB per request)
                     ┌────────▼─────────┐
                     │  DECODE POOL     │──▶ streamed tokens
                     │  memory-bound    │
                     │  low TP, huge    │
                     │  batch, bandwidth│
                     │  -optimized HW   │
                     └──────────────────┘

Diagram — Prefill and decode on separate workers#

sequenceDiagram
    participant C as Client
    participant R as Router
    participant P as Prefill worker
    participant D as Decode worker
    C->>R: request
    R->>P: prompt
    Note over P: compute-bound and bursty
    P->>D: KV cache + first token over RDMA, NVLink or TCP
    Note over D: memory-bound and steady
    loop until EOS
        D-->>C: token
    end
    Note over P,D: Pays off only if moving the KV is faster than recomputing the prefill

2. Problem → Why → Optimization#

PROBLEM   Prefill and decode have opposite characteristics (Section V.03)
          but share hardware and configuration. Every tuning decision is
          a compromise between them.

WHY       They're different workloads:
            prefill: compute-bound, benefits from FLOPs and high TP
            decode:  memory-bound, benefits from bandwidth and huge batch
          A single configuration serves neither optimally.

OPTIMIZE  Separate them. Tune each pool independently. Scale each
          independently.

TRADE-OFFS
  ✓ each pool optimally configured (TP degree, batch size, hardware)
  ✓ independent scaling: prefill-heavy traffic scales the prefill pool
  ✓ no interference: long prefills cannot affect decode ITL at all
  ✓ better hardware matching (FLOPs-rich for prefill, bandwidth-rich
    for decode)
  ✗ KV TRANSFER: GB per request over the interconnect
  ✗ substantially more complex: two pools, transfer machinery, a
    scheduler that spans both
  ✗ TTFT includes the transfer time
  ✗ layout compatibility constraints between pools

WHEN TO USE   Very large scale, extreme prefill/decode ratio imbalance,
              and a fast interconnect. This is a "hundreds of GPUs"
              technique.
WHEN NOT TO   Almost everywhere else. Chunked prefill (file 05) captures
              most of the benefit for a fraction of the complexity.

Be clear about this: chunked prefill is the answer for most deployments. Disaggregation is what you do when chunked prefill’s compromise is no longer good enough — which is a large-scale problem.


3. The KV transfer — the crux#

After prefill, the decode pool needs the sequence's KV cache.

SIZE
  Llama-3-70B, GQA-8, FP8 KV, 4,000-token prompt:
    4000 × 160 KiB = 640 MB per request

TRANSFER TIME
  NVLink (same node):      640 MB / 370 GB/s = 1.7 ms
  InfiniBand NDR (400G):   640 MB / 50 GB/s  = 12.8 ms
  RoCE 100G:               640 MB / 12 GB/s  = 53 ms
  Ethernet 25G (no RDMA):  640 MB / 3 GB/s   = 213 ms   ← unusable

COMPARE TO RECOMPUTING IT
  prefill 4,000 tokens at 28,000 tok/s = 143 ms

  → transfer is CHEAPER than recompute on NVLink and InfiniBand
  → this is the same arithmetic as Section IX.11, and it's what
    makes disaggregation viable at all

Disaggregation requires RDMA-class interconnect. On plain Ethernet the transfer costs more than the prefill it saves.

Layer-wise streaming#

The transfer can OVERLAP with prefill:

  as each layer's KV is computed, send it to the decode pool
  → by the time prefill finishes, most of the KV has arrived
  → effective transfer time ≈ the last layer's transfer only

  layer 0 KV computed → send ─┐
  layer 1 KV computed → send ─┤ overlapped with
  ...                         ├ continuing prefill
  layer 79 KV computed → send ┘
  
  → transfer latency largely hidden

This is the key implementation technique and what makes disaggregation’s TTFT competitive. Without it, TTFT = prefill + full transfer.


4. Why it wins (when it does)#

INDEPENDENT CONFIGURATION

                        PREFILL POOL        DECODE POOL
  TP degree             8 (more FLOPs)      2-4 (less comm overhead)
  batch size            small (a few         huge (hundreds — decode
                        long prompts fill    needs batch for intensity)
                        the GPU)
  max_num_seqs          low                  very high
  hardware              FLOPs-optimized      bandwidth-optimized
                        (H100)               (H200)
  KV cache size         small (transient)    large (holds all decodes)
  
INDEPENDENT SCALING
  RAG traffic surge → scale the prefill pool only
  long generations → scale the decode pool only
  → in a colocated deployment you'd scale everything

NO INTERFERENCE
  a 128k prefill cannot affect decode ITL, because it's on
  different hardware entirely
  → p99 ITL becomes very stable
REPORTED GAINS (DistServe, Splitwise, and similar)
  1.5-4x goodput at the same SLO, or
  tighter SLOs at the same throughput

  The gain is largest when:
    - the prefill/decode ratio is extreme
    - SLOs are tight on both TTFT and ITL
    - you have enough scale that pool-level tuning matters

5. What makes it hard#

1. LAYOUT COMPATIBILITY
   KV computed at TP=8 has a different head split than KV consumed
   at TP=4. Either match the TP degrees (losing the independent
   configuration benefit) or reshape during transfer (costly).
   → most implementations require matching TP degrees, which
     removes one of the main benefits.

2. SCHEDULING ACROSS POOLS
   a request occupies a prefill slot, then a decode slot.
   The scheduler must reason about both, and about the transfer.
   Backpressure from the decode pool must reach the prefill pool.

3. THE TRANSFER MACHINERY
   RDMA setup, buffer management, flow control, error handling.
   This is real systems engineering.

4. FAILURE MODES
   what if the decode pool has no capacity when prefill finishes?
   → hold the KV (occupying prefill memory), or discard and retry
   what if the transfer fails mid-flight?
   → the request is in an inconsistent state

5. PREFIX CACHING ACROSS POOLS
   the prefix cache lives in the prefill pool. Fine.
   But the decode pool's KV for a completed request could serve
   as a prefix cache for a follow-up turn... in the wrong pool.
   → multi-turn conversation handling is genuinely complicated

6. COST OF THE SPLIT
   two pools means two sets of model weights resident.
   → for a 70B model at FP8: 71 GB × 2 pools = 142 GB just for weights
   → this is a real cost, and it's why small deployments shouldn't
     do this

Point 6 is the one that rules it out for small deployments. Duplicating the weights across pools is only affordable when each pool is many GPUs.


6. The systems#

DistServe (2024)        the academic formulation; goodput-optimal
                        placement of prefill and decode
Splitwise (Microsoft)   phase splitting with hardware heterogeneity —
                        different GPU types per pool
Mooncake (Moonshot)     KV-cache-centric architecture with a
                        disaggregated KV store as a first-class component
DeepSeek's system       (per their reports) large-scale prefill/decode
                        separation with expert parallelism
vLLM / SGLang           experimental disaggregation support via
                        "KV connectors"

Mooncake’s framing is worth noting: rather than “prefill pool sends KV to decode pool,” they treat the KV cache as a shared distributed store that both pools access. This unifies disaggregation with prefix caching and KV tiering (Section IX.11) into one abstraction.


7. Is it right for you?#

DECISION CHECKLIST

□ Do you have RDMA-class interconnect (NVLink or InfiniBand)?
    NO → stop. The transfer is too expensive.

□ Is chunked prefill already enabled and tuned?
    NO → do that first. It's 90% of the benefit for 5% of the effort.

□ Are you still seeing prefill/decode interference after chunked prefill?
    Measure: p99 ITL vs p50 ITL. If p99 < 2× p50, you don't have
    an interference problem.

□ Is your prefill/decode ratio extreme?
    RAG (10:1 input:output) or long-generation (1:20) benefit most.
    Balanced workloads benefit least.

□ Are you at a scale where two pools each have many GPUs?
    Weight duplication must be a small fraction of your fleet.

□ Do you have the engineering capacity for the transfer machinery,
  cross-pool scheduling, and the failure modes?

ALL YES → evaluate it.
ANY NO  → chunked prefill is your answer.

Most readers of this file should conclude “not yet.” That’s the correct conclusion, and knowing why is the point.


8. Production implications#

  • Chunked prefill first. Always. It’s the 90% solution.
  • Disaggregation requires RDMA. Verify GPUDirect RDMA (Section IX.08).
  • Layer-wise KV streaming is essential to hide the transfer latency.
  • Matching TP degrees between pools is usually required — factor that into the benefit calculation.
  • Budget for weight duplication across pools.
  • The cross-pool scheduler is the hard part, not the transfer.
  • Watch the field. Support in mainstream engines is improving; the calculus may change.

9. Common mistakes#

Adopting it before chunked prefill. Far more complexity for marginal additional benefit.

Without RDMA. The transfer costs more than the prefill.

Without layer-wise streaming. TTFT includes the full transfer.

Assuming independent TP degrees work. Layout compatibility usually forces matching.

Underestimating the weight duplication cost.

Not handling the “decode pool full when prefill finishes” case.

Deploying it at a scale where the complexity dominates.


10. Hands-on exercise#

A. Compute the transfer economics. For your model and interconnect, compute: KV size per request at your p50 and p95 prompt lengths, transfer time, and prefill time. Is transfer cheaper than recompute?

B. Measure your interference. With chunked prefill enabled, measure p50 and p99 ITL. Is p99 > 2× p50? If not, you don’t have the problem disaggregation solves.

C. Simulate the benefit. Build a simulator with separate prefill and decode pools, each independently configured. Compare goodput against a colocated pool with chunked prefill, for several prefill/decode ratios. Where does disaggregation win?

D. Layer-wise streaming. Implement (or simulate) layer-wise KV transfer overlapped with prefill. Measure the effective transfer latency versus the naive end-of-prefill transfer.

E. The weight duplication cost. For a deployment you know, compute the GPU cost of duplicating weights across two pools. At what fleet size does it become a small fraction?

F. Read the papers. Read DistServe and Splitwise. What assumptions does each make about the interconnect and the workload? Do they hold for you?


11. Interview questions#

  1. What is disaggregated serving and what problem does it solve?
  2. How large is the KV transfer, and when is it cheaper than recomputing?
  3. What is layer-wise KV streaming and why is it necessary?
  4. Why do the two pools usually need matching TP degrees, and what does that cost?
  5. What must be true before you’d consider disaggregation?
  6. Why is chunked prefill usually the right answer instead?
  7. What’s the hardest engineering part of disaggregation?

12. Further reading#

  • [EMERGING] Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving” (OSDI 2024)
  • [EMERGING] Patel et al., “Splitwise: Efficient Generative LLM Inference Using Phase Splitting” (ISCA 2024)
  • [EMERGING] Qin et al., “Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving” (2024)
  • [REFERENCE] vLLM KV connector / disaggregated prefill documentation
  • Next: 07 — KV offloading and transfer

↑↓ navigate↵ openesc close