Below the API

Sharing One GPU

Advanced Advanced 50 min Difficulty 3/5

Prerequisites III.04, IV.05

The idea in one minute#

A flagship GPU is often far bigger than one small workload needs. There are three ways to let several processes share it: time-slicing (take turns), MPS (run together, no walls), and MIG (cut the GPU into separate smaller GPUs, with walls). They trade efficiency against isolation — how much one tenant can slow down, crash or spy on another.

An analogy#

Sharing a commercial kitchen:

  • Time-slicing: each restaurant gets the whole kitchen for an hour. Simple, but the kitchen idles while they swap, and only one cooks at a time.
  • MPS: everyone cooks at once in the same room. Efficient, but someone’s fire sets off everyone’s alarm.
  • MIG: the kitchen is rebuilt as several small locked kitchens. Nobody can interfere — but a small kitchen stays small even when the one next door is empty.

A picture#

flowchart TB
  subgraph TS["Time-slicing"]
    direction LR
    A1["Process A<br/>whole GPU"] -->|"context switch"| A2["Process B<br/>whole GPU"]
  end
  subgraph MPS["MPS"]
    direction LR
    B1["Process A"] --> SRV["MPS server<br/>one shared context"]
    B2["Process B"] --> SRV
    SRV --> GPU1["GPU: kernels overlap"]
  end
  subgraph MIG["MIG"]
    direction LR
    C1["Process A"] --> S1["Slice 1<br/>own SMs, own memory"]
    C2["Process B"] --> S2["Slice 2<br/>own SMs, own memory"]
  end
  class A1,A2,B1,B2,C1,C2 compute
  class SRV queue
  class GPU1,S1,S2 memory

How it really works#

Default: time-slicing#

Each process gets its own GPU context (its private state on the device). When several processes are active, the GPU runs one context at a time and switches between them every few milliseconds. Each context switch has a cost, and while process A runs, B’s kernels wait — even if A is only using 5% of the SMs.

Memory is not time-sliced: every process’s allocations coexist, and there is no limit per process. One greedy process can make the others fail with out-of-memory.

MPS: Multi-Process Service#

MPS funnels several processes through one shared context, so their kernels can run on the GPU at the same time, filling SMs that one process would leave idle. For many small workloads this raises total throughput substantially.

The cost is isolation. Limits on each client’s memory and compute share are available but are soft: there is no hardware wall. Historically a fatal error in one client could take down the others. MPS is for processes that trust each other — typically replicas of your own service.

MIG: Multi-Instance GPU#

On supporting data-center GPUs (A100, H100 and later), MIG partitions the hardware into up to seven instances, each with dedicated SMs, dedicated memory, and its own slice of the memory system. To software each instance looks like a separate, smaller GPU.

Isolation is real: one instance cannot use another’s memory, slow it down, or crash it. The cost is rigidity: sizes come from a fixed menu, changing the layout requires the GPU to be idle, and an instance cannot borrow idle capacity from its neighbour.

Choosing#

Time-slicingMPSMIG
Kernels run concurrentlyNoYesYes
Memory isolationNoneSoft limitsHard
Performance isolationNoneWeakStrong
Fault isolationYes (separate contexts)WeakStrong
FlexibilityHighHighLow (fixed slice sizes)
Best forDev machines, bursty notebooksYour own small services packed togetherMultiple tenants, guaranteed capacity

Also possible, and often best: share inside one process. A single inference server that batches requests from many users (as LLM servers do) shares the GPU more efficiently than any of the three, because it combines work into the same kernels. The mechanisms above are for when workloads cannot live in one process.

Code#

A small simulator: two bursty workloads on one GPU under time-slicing versus MPS-style concurrent execution.

// sharing.go — why running together beats taking turns for small workloads.
package main

import "fmt"

type Job struct {
	Name     string
	Kernels  int     // kernels to run
	KernelMs float64 // duration of each
	SMShare  float64 // fraction of the GPU's SMs each kernel can use
}

func timeSliced(jobs []Job, sliceMs, switchMs float64) float64 {
	remaining := make([]float64, len(jobs))
	for i, j := range jobs {
		remaining[i] = float64(j.Kernels) * j.KernelMs
	}
	var clock float64
	for active := len(jobs); active > 0; {
		for i := range jobs {
			if remaining[i] <= 0 {
				continue
			}
			run := min(sliceMs, remaining[i])
			clock += run + switchMs // only this job runs; then pay for the switch
			if remaining[i] -= run; remaining[i] <= 0 {
				active--
			}
		}
	}
	return clock
}

func concurrent(jobs []Job) float64 {
	// Kernels overlap. If their combined SM demand fits, nobody waits.
	var demand, longest float64
	for _, j := range jobs {
		demand += j.SMShare
		longest = max(longest, float64(j.Kernels)*j.KernelMs)
	}
	return longest * max(1, demand) // oversubscribed SMs stretch everything
}

func main() {
	jobs := []Job{{"embedder", 1000, 2, 0.25}, {"reranker", 1000, 2, 0.30}, {"small LLM", 1000, 2, 0.35}}
	fmt.Printf("time-sliced: %6.0f ms\n", timeSliced(jobs, 2, 0.3))
	fmt.Printf("concurrent:  %6.0f ms\n", concurrent(jobs))
}

Remember this#

  • Time-slicing: one process at a time, no memory limits. Fine for development.
  • MPS: concurrent kernels, soft limits, weak isolation. For trusted co-tenants.
  • MIG: hardware partitions, strong isolation, fixed sizes.
  • Batching inside one server is the most efficient sharing of all.

Try it#

  1. Run sharing.go. Raise each job’s SMShare to 0.6. What happens to the concurrent time, and why?
  2. Add a memory field to Job and a per-slice memory limit. Model a MIG layout of three slices and show a job that fits under MPS but not under MIG.
  3. You operate a platform for external customers. Which mechanism, and why?

Check yourself#

  1. Why does time-slicing waste capacity for small workloads?
  2. What does MPS give up to gain efficiency?
  3. What makes MIG isolation “hard” rather than “soft”?

↑↓ navigate ↵ open