1. What is it?#
Split the model by layers across GPUs. GPU 0 holds layers 0-19, GPU 1 holds 20-39, and so on. Activations flow from one stage to the next.
GPU 0 GPU 1 GPU 2 GPU 3
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ layers │───►│ layers │───►│ layers │───►│ layers │──► logits
│ 0-19 │ │ 20-39 │ │ 40-59 │ │ 60-79 │
└──────────┘ └──────────┘ └──────────┘ └──────────┘
↑
activations (batch × seq × d) sent point-to-point — SMALLDiagram — Layers split in sequence#
flowchart LR IN["Tokens"] --> S0["GPU 0<br/>first half of the layers"] S0 -->|"activations<br/>batch x d_model"| S1["GPU 1<br/>second half of the layers"] S1 --> OUT["Logits, next token"] OUT -.->|"every token crosses every stage, in order"| IN class S0,S1 compute class IN,OUT neutral
2. Problem → Why → Optimization#
PROBLEM TP doesn't scale past ~8 GPUs (file 03), and TP requires NVLink,
which usually doesn't extend across nodes.
WHY TP's AllReduce cost is fixed per layer and grows relatively as N grows.
OPTIMIZE Split by layers instead. Communication is a point-to-point send of
the activation tensor between adjacent stages — tiny compared to
an AllReduce, and it happens once per stage boundary rather than
twice per layer.
TRADE-OFFS
✓ tiny communication (activations only, ~1/160th of TP's traffic)
✓ works over slow interconnects — this is why it's the multi-node choice
✓ enables models far larger than one node
✗ PIPELINE BUBBLES: stages idle while waiting
✗ does NOT reduce per-token latency (the token still traverses all layers)
✗ load balancing across stages is fiddly
WHEN TO USE Multi-node, when the model is too large for one node's TP group.
WHEN NOT TO Single node (use TP), or when you need lower latency.3. Simple analogy#
An assembly line versus four workers each building a whole car.
Assembly line (PP): each worker does one stage. A car takes the same total time to build, but four cars are in progress at once — throughput is 4x if the line is full.
The bubble: when only one car is on the line, three workers are idle. The line is only efficient when it’s full — which means many concurrent items.
That’s the essential property: PP improves throughput, not latency, and only when there’s enough work to fill the pipeline.
4. The bubble, quantified#
Naive PP with 4 stages and 1 microbatch:
time →
GPU 0: [S0] idle idle idle
GPU 1: [S1] idle idle
GPU 2: [S2] idle
GPU 3: [S3]
Utilization: 4 stages × 1 unit of work / (4 GPUs × 4 time units) = 25%With microbatches, the pipeline fills:
4 stages, 8 microbatches:
GPU 0: [1][2][3][4][5][6][7][8]
GPU 1: [1][2][3][4][5][6][7][8]
GPU 2: [1][2][3][4][5][6][7][8]
GPU 3: [1][2][3][4][5][6][7][8]
└bubble┘ └bubble┘
Bubble fraction = (stages - 1) / (microbatches + stages - 1)
= 3 / (8 + 3) = 27%Bubble fraction for P stages and M microbatches: (P-1)/(M+P-1)
P=4, M=4 → 43%
P=4, M=8 → 27%
P=4, M=32 → 9%
P=4, M=128 → 2.3%
P=8, M=32 → 18%
P=8, M=128 → 5.2%The rule: you need M ≫ P. For PP=4 you want at least 32 microbatches — meaning at least 32 concurrent sequences. PP is only efficient at high concurrency, which makes it a throughput-serving technique, not a low-latency one.
In LLM decode, what is a microbatch?#
For decode, a natural microbatch is a group of sequences. With 128 concurrent sequences and PP=4, you can split into 8 microbatches of 16 — giving a 27% bubble — or 32 microbatches of 4, giving 9% but with more per-microbatch overhead.
This is a real tuning parameter and it’s why PP is more finicky than TP.
5. Technical explanation#
Communication cost — the reason PP works across nodes#
PP: send the activation tensor between adjacent stages
size = batch × seq × d × bytes
Llama-3-70B decode, batch 32, d=8192, BF16:
32 × 1 × 8192 × 2 = 512 KB
3 stage boundaries (PP=4) → 1.5 MB per token, point-to-point
TP: 160 AllReduces per token
at batch 32: 143 MB per token, collective
→ PP moves ~100x less data than TP.Over a 400 Gb/s InfiniBand link (50 GB/s): 1.5 MB / 50 GB/s = 30 µs. Negligible.
This is why the standard large-model layout is TP within a node and PP across nodes.
Load balancing across stages#
Naive: equal layers per stage. But stages aren't equal:
- stage 0 also has the embedding
- the last stage also has the final norm and the LM head (expensive! V×d)
- some architectures have uneven layers
Result: the last stage is often 20-40% slower, and the whole pipeline
runs at the slowest stage's speed.
Fix: assign FEWER transformer layers to the first and last stages.
e.g. PP=4 on 80 layers: [18, 21, 21, 20] instead of [20, 20, 20, 20]Engines expose this (--pipeline-parallel-size plus a partition specification, or automatic
balancing based on profiling). Check whether yours balances automatically — the default equal
split is often 15-25% off optimal.
Interaction with continuous batching#
This is where PP gets genuinely complicated:
With continuous batching, sequences join and leave every step.
With PP, several microbatches are in flight across stages simultaneously.
→ The scheduler must decide the batch for microbatch M+1 while
microbatch M is still traversing the pipeline.
→ Freeing KV for a finished sequence must wait until it exits the pipeline.
→ Preemption is harder: the sequence may be mid-pipeline.Engines handle this with a scheduling lookahead and careful bookkeeping. It works, but it’s a source of complexity and of subtle bugs. PP is meaningfully harder to get right than TP, and that’s part of why it’s less commonly used for inference than for training.
6. Under the hood — the standard multi-node layout#
2 nodes × 8 GPUs = 16 GPUs, model = 405B FP16 (810 GB)
TP=8 within each node (over NVLink), PP=2 across nodes (over InfiniBand):
Node 0: GPUs 0-7, layers 0-63, TP=8 → 810/2/8 = 50.6 GB per GPU
Node 1: GPUs 8-15, layers 64-125, TP=8
Communication:
within node: 2 AllReduces per layer over NVLink (fast)
across node: 1 activation send per token over InfiniBand (tiny)This is the canonical layout for models that exceed one node. TP where the interconnect is fast, PP where it isn’t.
7. Performance#
Llama-3-405B FP16, 2 nodes × 8 H100:
Configuration ITL Throughput (b=64) Notes
TP=16 (across nodes) — — AllReduce over IB: unusable
TP=8 × PP=2 32 ms 2,000 tok/s the right answer
PP=16 — — 16 stages: huge bubbles
At batch 8 (low concurrency):
TP=8 × PP=2 38 ms 210 tok/s bubble fraction ~30%
→ PP is inefficient at low concurrency8. Production implications#
- Use PP only across nodes, and only when TP within a node isn’t enough.
- Balance the stages. The last stage has the LM head; give it fewer layers.
- PP needs high concurrency. If your traffic is low, the bubbles dominate. Consider a smaller model or quantization instead.
- Try quantization first. FP8 halves the model and may eliminate the need for PP entirely — a large operational simplification (Section IX.01, example C).
- Monitor per-stage utilization. An imbalanced pipeline shows as one stage at 100% and the rest lower.
- Failure handling is harder: a node failure kills the pipeline, and restarting requires reloading both nodes.
9. Common mistakes#
Using PP within a node. TP is strictly better there (lower latency, no bubbles).
PP with low concurrency. Bubbles dominate.
Equal layer split. The last stage is slower; balance it.
Too many stages. Bubble fraction grows with P; keep PP ≤ 4 for inference.
Expecting PP to reduce latency. It doesn’t; the token still traverses every layer.
Not trying quantization first.
10. Hands-on exercise#
A. Compute the bubbles. For P ∈ {2,4,8} and M ∈ {4,8,32,128}, compute the bubble fraction. Plot it. What M do you need for < 10% bubbles at P=4?
B. Measure PP. If you have multi-node access, run a model with TP=8×PP=2 and measure ITL and throughput at batch 8, 32, 128. Compute the effective bubble fraction from the measurements.
C. Stage balancing. Profile per-stage time in a PP deployment. How imbalanced is the default split? Compute the optimal layer assignment and measure the improvement.
D. Communication comparison. Compute the per-token communication volume for TP=16 and for TP=8×PP=2 on a 405B model. Confirm the ~100x difference.
E. Quantization alternative. For a model that requires PP at FP16, check whether FP8 lets it fit in one node with TP=8. Compare total throughput and operational complexity.
11. Interview questions#
- What is pipeline parallelism and how does it differ from tensor parallelism?
- What is a pipeline bubble? Give the formula and compute it for PP=4, M=8.
- Why does PP work across nodes when TP doesn’t?
- Why is the last pipeline stage usually the bottleneck, and how do you fix it?
- Why does PP not improve per-token latency?
- How does PP complicate continuous batching?
- What’s the canonical layout for a model too large for one node, and why?
12. Further reading#
- [ESTABLISHED] Huang et al., “GPipe” (2018) — the bubble analysis
- [ESTABLISHED] Narayanan et al., “PipeDream” (2019) and Megatron-LM’s pipeline schedules
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022)
- [REFERENCE] vLLM and TensorRT-LLM pipeline parallelism documentation
- Next: 05 — Expert parallelism