The idea in one minute#
There are three basic ways to use several GPUs for one model. Replicate it: each GPU holds a full copy and serves different requests. Split each layer across GPUs (tensor parallelism): they compute pieces of every layer and combine results. Split the layers among GPUs (pipeline parallelism): data flows through them like stations on a line.
Replication scales throughput with no communication at all. The other two exist only for models that do not fit on one GPU, and both pay a communication tax.
A picture#
flowchart TB
subgraph R["Replicas: independent copies"]
direction LR
LB["Load balancer"] --> RA["GPU 0<br/>whole model"]
LB --> RB["GPU 1<br/>whole model"]
end
subgraph TP["Tensor parallel: every layer split"]
direction LR
TA["GPU 0<br/>half of each layer"] <-->|"exchange every layer"| TB["GPU 1<br/>half of each layer"]
end
subgraph PP["Pipeline parallel: layers split"]
direction LR
PA["GPU 0<br/>layers 1 to 40"] -->|"hand over once"| PB["GPU 1<br/>layers 41 to 80"]
end
class LB queue
class RA,RB,TA,TB,PA,PB computeHow it really works#
| Replicas | Tensor parallel | Pipeline parallel | |
|---|---|---|---|
| Each GPU holds | Whole model | A slice of every layer | A contiguous range of layers |
| Fits a model larger than one GPU | No | Yes | Yes |
| Communication | None | Twice per layer, between all GPUs | Once per stage boundary |
| Effect on latency | Unchanged | Slightly worse than one (imaginary) big GPU | Unchanged per token, but GPUs wait for each other |
| Effect on throughput | Scales linearly | Sub-linear: 2 GPUs ≈ 1.6–1.9x | Good only if stages stay busy |
| Needs fast interconnect | No | Yes — NVLink class | Modest |
| Failure of one GPU | Others keep serving | The whole group stops | The whole pipeline stops |
Replicas: always the first choice#
If the model fits on one GPU, give each GPU its own copy and balance requests across them. Throughput scales perfectly, a failure removes only that capacity, and no special wiring is needed. Only reach for the other schemes when the model (plus a useful amount of KV cache) does not fit.
Tensor parallelism#
Each weight matrix is cut into N slices, one per GPU. Every GPU computes its slice of the
layer; then the partial results must be combined, which requires an all-reduce: every GPU
ends up with the sum of everyone’s contribution. This happens about twice per layer.
Its efficiency is compute ÷ (compute + communication). With NVLink and large batches,
communication is small and efficiency is 85–95%. Without NVLink, or at very small batches where
the fixed latency dominates, it collapses (lesson 01’s program shows the arithmetic).
Pipeline parallelism#
GPU 0 runs the first layers and passes the activations to GPU 1. Communication is light, but for a single sequence only one GPU works at a time — the other waits. It pays off when many sequences are in flight so every stage stays busy, and it is the usual fallback when the interconnect is too slow for tensor parallelism.
Collectives and NCCL#
The standard group-communication patterns are called collectives:
| Collective | Result |
|---|---|
| Broadcast | One GPU’s data is copied to all |
| All-gather | Every GPU ends up with everyone’s pieces, concatenated |
| Reduce | One GPU gets the sum of all |
| All-reduce | Every GPU gets the sum of all |
NVIDIA’s NCCL library implements them, choosing rings or trees over the fastest links it finds (NVLink, then PCIe, then the network). Frameworks call NCCL; you rarely do so directly. But “NCCL error” in a log now tells you which layer of the system is unhappy: GPU-to-GPU communication.
When more GPUs make it slower#
Adding GPUs to a split model adds communication to every layer and shrinks each GPU’s share of compute. Past some point the tax outgrows the gain. If a model fits on four GPUs, running it tensor-parallel on eight often yields lower throughput than two independent four-GPU replicas.
Rule: split as little as needed to fit, then replicate.
Code#
A ring all-reduce among goroutine “GPUs”, and the efficiency formula.
// allreduce.go — the collective behind tensor parallelism, and what it costs.
package main
import (
"fmt"
"sync"
)
// allReduce: each rank starts with its own vector and ends with the element-wise sum of all.
// Values travel around a ring; each hop is one message over a link.
func allReduce(local [][]float64) (hops int) {
n := len(local)
sum := append([]float64{}, local[0]...)
for r := 1; r < n; r++ { // reduce phase: accumulate around the ring
for i := range sum {
sum[i] += local[r][i]
}
hops++
}
var wg sync.WaitGroup
for r := 0; r < n; r++ { // broadcast phase: everyone receives the total
wg.Add(1)
go func() { defer wg.Done(); copy(local[r], sum) }()
hops++
}
wg.Wait()
return hops - 1 // the last rank already holds the sum
}
func main() {
gpus := [][]float64{{1, 2}, {10, 20}, {100, 200}, {1000, 2000}}
hops := allReduce(gpus)
fmt.Println("every GPU now holds", gpus[0], "after", hops, "hops")
// Efficiency of tensor parallelism = compute / (compute + communication).
fmt.Println("\nGPUs compute/token comm/token speedup vs 1 GPU")
const computeMs = 40.0
for _, c := range []struct {
n int
commMs float64
}{{2, 1.5}, {4, 3}, {8, 6}, {8, 60}} {
perToken := computeMs/float64(c.n) + c.commMs
fmt.Printf("%4d %10.1f ms %7.1f ms %8.2fx\n", c.n, computeMs/float64(c.n), c.commMs, computeMs/perToken)
}
}The last row is eight GPUs on a slow interconnect: slower than one GPU.
Remember this#
- Fits on one GPU → replicate. Nothing else scales as well.
- Tensor parallelism: split every layer; needs NVLink-class links; all-reduce twice per layer.
- Pipeline parallelism: split by layers; light communication; GPUs idle unless many sequences are in flight.
- Split as little as needed to fit, then replicate.
Try it#
- Run
allreduce.go. Change it to a tree (pairs combine in parallel). How many rounds does it take for 8 GPUs versus the ring? - With
computeMs = 40, what communication time makes 4 GPUs no faster than 2? - You have 8 GPUs and a model that fits on 2. Compare one 8-way group against four 2-way replicas for throughput and for the impact of a single GPU failure.
Check yourself#
- Why are replicas preferred whenever the model fits?
- What is an all-reduce and why does tensor parallelism need it?
- Explain how adding GPUs can reduce throughput.