PidokuInfra

Splitting Work Across GPUs

Advanced 1h Difficulty 4/5 Topic 02 of 04

Prerequisites 01, V.03

The idea in one minute#

There are three basic ways to use several GPUs for one model. Replicate it: each GPU holds a full copy and serves different requests. Split each layer across GPUs (tensor parallelism): they compute pieces of every layer and combine results. Split the layers among GPUs (pipeline parallelism): data flows through them like stations on a line.

Replication scales throughput with no communication at all. The other two exist only for models that do not fit on one GPU, and both pay a communication tax.

A picture#

flowchart TB
  subgraph R["Replicas: independent copies"]
    direction LR
    LB["Load balancer"] --> RA["GPU 0<br/>whole model"]
    LB --> RB["GPU 1<br/>whole model"]
  end
  subgraph TP["Tensor parallel: every layer split"]
    direction LR
    TA["GPU 0<br/>half of each layer"] <-->|"exchange every layer"| TB["GPU 1<br/>half of each layer"]
  end
  subgraph PP["Pipeline parallel: layers split"]
    direction LR
    PA["GPU 0<br/>layers 1 to 40"] -->|"hand over once"| PB["GPU 1<br/>layers 41 to 80"]
  end
  class LB queue
  class RA,RB,TA,TB,PA,PB compute

How it really works#

ReplicasTensor parallelPipeline parallel
Each GPU holdsWhole modelA slice of every layerA contiguous range of layers
Fits a model larger than one GPUNoYesYes
CommunicationNoneTwice per layer, between all GPUsOnce per stage boundary
Effect on latencyUnchangedSlightly worse than one (imaginary) big GPUUnchanged per token, but GPUs wait for each other
Effect on throughputScales linearlySub-linear: 2 GPUs ≈ 1.6–1.9xGood only if stages stay busy
Needs fast interconnectNoYes — NVLink classModest
Failure of one GPUOthers keep servingThe whole group stopsThe whole pipeline stops

Replicas: always the first choice#

If the model fits on one GPU, give each GPU its own copy and balance requests across them. Throughput scales perfectly, a failure removes only that capacity, and no special wiring is needed. Only reach for the other schemes when the model (plus a useful amount of KV cache) does not fit.

Tensor parallelism#

Each weight matrix is cut into N slices, one per GPU. Every GPU computes its slice of the layer; then the partial results must be combined, which requires an all-reduce: every GPU ends up with the sum of everyone’s contribution. This happens about twice per layer.

Its efficiency is compute ÷ (compute + communication). With NVLink and large batches, communication is small and efficiency is 85–95%. Without NVLink, or at very small batches where the fixed latency dominates, it collapses (lesson 01’s program shows the arithmetic).

Pipeline parallelism#

GPU 0 runs the first layers and passes the activations to GPU 1. Communication is light, but for a single sequence only one GPU works at a time — the other waits. It pays off when many sequences are in flight so every stage stays busy, and it is the usual fallback when the interconnect is too slow for tensor parallelism.

Collectives and NCCL#

The standard group-communication patterns are called collectives:

CollectiveResult
BroadcastOne GPU’s data is copied to all
All-gatherEvery GPU ends up with everyone’s pieces, concatenated
ReduceOne GPU gets the sum of all
All-reduceEvery GPU gets the sum of all

NVIDIA’s NCCL library implements them, choosing rings or trees over the fastest links it finds (NVLink, then PCIe, then the network). Frameworks call NCCL; you rarely do so directly. But “NCCL error” in a log now tells you which layer of the system is unhappy: GPU-to-GPU communication.

When more GPUs make it slower#

Adding GPUs to a split model adds communication to every layer and shrinks each GPU’s share of compute. Past some point the tax outgrows the gain. If a model fits on four GPUs, running it tensor-parallel on eight often yields lower throughput than two independent four-GPU replicas.

Rule: split as little as needed to fit, then replicate.

Code#

A ring all-reduce among goroutine “GPUs”, and the efficiency formula.

Go
// allreduce.go — the collective behind tensor parallelism, and what it costs.
package main

import (
	"fmt"
	"sync"
)

// allReduce: each rank starts with its own vector and ends with the element-wise sum of all.
// Values travel around a ring; each hop is one message over a link.
func allReduce(local [][]float64) (hops int) {
	n := len(local)
	sum := append([]float64{}, local[0]...)
	for r := 1; r < n; r++ { // reduce phase: accumulate around the ring
		for i := range sum {
			sum[i] += local[r][i]
		}
		hops++
	}
	var wg sync.WaitGroup
	for r := 0; r < n; r++ { // broadcast phase: everyone receives the total
		wg.Add(1)
		go func() { defer wg.Done(); copy(local[r], sum) }()
		hops++
	}
	wg.Wait()
	return hops - 1 // the last rank already holds the sum
}

func main() {
	gpus := [][]float64{{1, 2}, {10, 20}, {100, 200}, {1000, 2000}}
	hops := allReduce(gpus)
	fmt.Println("every GPU now holds", gpus[0], "after", hops, "hops")

	// Efficiency of tensor parallelism = compute / (compute + communication).
	fmt.Println("\nGPUs  compute/token  comm/token  speedup vs 1 GPU")
	const computeMs = 40.0
	for _, c := range []struct {
		n      int
		commMs float64
	}{{2, 1.5}, {4, 3}, {8, 6}, {8, 60}} {
		perToken := computeMs/float64(c.n) + c.commMs
		fmt.Printf("%4d  %10.1f ms  %7.1f ms  %8.2fx\n", c.n, computeMs/float64(c.n), c.commMs, computeMs/perToken)
	}
}

The last row is eight GPUs on a slow interconnect: slower than one GPU.

Remember this#

  • Fits on one GPU → replicate. Nothing else scales as well.
  • Tensor parallelism: split every layer; needs NVLink-class links; all-reduce twice per layer.
  • Pipeline parallelism: split by layers; light communication; GPUs idle unless many sequences are in flight.
  • Split as little as needed to fit, then replicate.

Try it#

  1. Run allreduce.go. Change it to a tree (pairs combine in parallel). How many rounds does it take for 8 GPUs versus the ring?
  2. With computeMs = 40, what communication time makes 4 GPUs no faster than 2?
  3. You have 8 GPUs and a model that fits on 2. Compare one 8-way group against four 2-way replicas for throughput and for the impact of a single GPU failure.

Check yourself#

  1. Why are replicas preferred whenever the model fits?
  2. What is an all-reduce and why does tensor parallelism need it?
  3. Explain how adding GPUs can reduce throughput.

↑↓ navigate↵ openesc close