1. What is it?#
A stream is an ordered queue of GPU operations. Operations in the same stream execute in order; operations in different streams may overlap.
Stream 0: [kernel A]──[kernel B]──[kernel C] in order
Stream 1: [memcpy H2D]──[kernel D] in order, may overlap with stream 0
Stream 2: [memcpy D2H] may overlap with bothStreams are how you get concurrency between independent work: overlapping data transfer with compute, running a small kernel alongside a large one, or serving two models on one GPU.
2. Why does it exist?#
Because a single stream serializes everything, and the GPU has independent hardware units that could be working simultaneously:
Hardware units on an H100:
- SMs (compute)
- Copy engines (H2D and D2H, separate)
- NVLink/network engines
With one stream: copy, then compute, then copy. Serial.
With three streams: copy(i+1) while compute(i) while copy-back(i-1). Pipelined.3. Simple analogy#
A kitchen with one cook versus a kitchen with lanes.
One stream: the cook does everything in strict order — fetch ingredients, chop, cook, plate, serve, then fetch the next order’s ingredients. The oven sits idle while fetching.
Multiple streams: while dish A is in the oven, fetch ingredients for dish B and plate dish C. The equipment is used concurrently.
The synchronization primitives are the rules: “don’t plate dish A until it’s out of the oven.”
4. Tiny example#
Overlapping transfer and compute:
import torch, time
N = 100_000_000
h = torch.empty(N, dtype=torch.float32, pin_memory=True) # pinned!
d = torch.empty(N, dtype=torch.float32, device='cuda')
W = torch.randn(4096, 4096, device='cuda', dtype=torch.float16)
x = torch.randn(4096, 4096, device='cuda', dtype=torch.float16)
# --- SERIAL: same stream ---
torch.cuda.synchronize(); t0 = time.perf_counter()
d.copy_(h)
y = x @ W
torch.cuda.synchronize()
t_serial = time.perf_counter()-t0
# --- OVERLAPPED: separate streams ---
s_copy = torch.cuda.Stream()
torch.cuda.synchronize(); t0 = time.perf_counter()
with torch.cuda.stream(s_copy):
d.copy_(h, non_blocking=True)
y = x @ W # runs on the default stream, concurrently
torch.cuda.synchronize()
t_overlap = time.perf_counter()-t0
print(f"serial {t_serial*1e3:.1f} ms, overlapped {t_overlap*1e3:.1f} ms, "
f"saved {(1-t_overlap/t_serial)*100:.0f}%")
Verify the overlap actually happened in Nsight Systems — the copy and the kernel should appear on different rows, temporally overlapping.
5. Technical explanation#
The default stream, and why it’s tricky#
Legacy default stream (stream 0):
Implicitly synchronizes with all other blocking streams.
Any operation on the default stream waits for all other streams,
and other streams wait for it.
→ accidentally serializes everything
Per-thread default stream (compile with --default-stream per-thread):
Each host thread gets its own default stream. No implicit sync.
Non-blocking streams (cudaStreamNonBlocking):
Do not synchronize with the legacy default stream.PyTorch uses a per-device default stream and creates non-blocking streams via
torch.cuda.Stream(), so the classic footgun is mostly avoided — but if you mix raw CUDA code
in, be careful.
Events for cross-stream dependencies#
s1, s2 = torch.cuda.Stream(), torch.cuda.Stream()
ev = torch.cuda.Event()
with torch.cuda.stream(s1):
a = produce()
ev.record(s1) # mark: "a is ready"
with torch.cuda.stream(s2):
ev.wait(s2) # s2 waits for the event, but the CPU does NOT block
b = consume(a)
ev.wait(stream) inserts a GPU-side dependency. The CPU keeps running. This is how you build
pipelines without CPU involvement.
Contrast with ev.synchronize(), which blocks the CPU until the event fires. Use that only
when you actually need the CPU to wait.
Where streams help in inference#
1. WEIGHT LOADING at startup
Overlap disk read, H2D copy, and layout conversion across streams.
Cuts cold start meaningfully (Section II.07).
2. KV CACHE OFFLOAD / PREFETCH
Copy KV blocks between GPU and CPU on a side stream while compute proceeds.
Essential for CPU-offload designs (Section XIII.07).
3. MULTI-MODEL SERVING on one GPU
Each model on its own stream. Works, but they compete for SMs;
MPS or MIG gives better isolation.
4. SPECULATIVE DECODING
Draft model on one stream, target on another — limited benefit since
they're dependent, but the draft's small kernels can fill gaps.
5. DISAGGREGATED KV TRANSFER
Sending KV to a decode worker while continuing prefill (Section XIII.06).
6. PREFILL/DECODE OVERLAP
Some engines run prefill and decode on separate streams. Contended,
but can improve utilization.Priorities#
high = torch.cuda.Stream(priority=-1) # lower number = higher priority
low = torch.cuda.Stream(priority=0)
Priority affects which stream’s blocks get scheduled first when SM slots free up. It does not preempt running blocks. So a long-running kernel on the low-priority stream still delays the high-priority one — priorities help at block granularity, not instruction granularity.
Practical use: put latency-critical decode on a high-priority stream and background work (prefix cache warming, metrics computation) on a low-priority one.
MPS and time-slicing#
Default: multiple processes on one GPU are TIME-SLICED by the driver.
Only one process's kernels run at a time. Context switches cost ~10-100 µs.
MPS (Multi-Process Service):
A daemon merges multiple processes' work into one context, so their
kernels can run CONCURRENTLY on different SMs.
nvidia-cuda-mps-control -d
MIG: Hardware partitioning. Strongest isolation, fixed sizes.For serving several small models on one GPU, MPS is often the right answer: better utilization than time-slicing, more flexible than MIG. Downside: a fault in one process can affect the MPS server.
6. Under the hood#
Whether two kernels actually overlap depends on resources:
Kernel A uses 100% of SMs → kernel B waits regardless of streams
Kernel A uses 20% of SMs → kernel B can fill the remaining 80%So streams enable overlap; they don’t guarantee it. Small kernels overlap well; large ones don’t. A decode step’s small kernels can overlap with a prefill’s large ones only to the extent the prefill leaves SMs free — which it usually doesn’t.
Verify overlap in Nsight Systems: kernels on different rows should visibly overlap in time. If they’re serialized despite being on different streams, look for an implicit synchronization or resource contention.
7. Performance implications#
Use case Typical gain from streams
Weight loading overlap 20-40% faster cold start
KV offload prefetch makes offload viable at all
Multi-model on one GPU (MPS) 1.5-3x utilization vs time-slicing
Prefill/decode overlap 5-15% (contended)
H2D/compute overlap in a pipeline 30-50% for transfer-heavy workloads8. Production implications#
- Use pinned memory pools for anything you stream.
- Verify overlap in a profile. Assuming it happened is a common error.
- Consider MPS for multi-model, small-model serving. Measure it against time-slicing.
- Beware over-synchronizing. A single
torch.cuda.synchronize()in a hot path serializes everything you carefully pipelined. - Stream priorities are a weak tool. Don’t rely on them for latency isolation; use separate GPUs or MIG for hard guarantees.
9. Common mistakes#
Assuming different streams guarantee overlap. Resources must be available.
Using the legacy default stream in code that mixes with other streams.
Forgetting pinned memory. No async transfer without it.
Over-synchronizing. torch.cuda.synchronize() in a loop.
Expecting stream priority to preempt. It doesn’t.
Not verifying with a profiler.
10. Hands-on exercise#
A. Prove overlap. Run the example in section 4. Verify in Nsight Systems that the copy and
the kernel overlap. Then remove pin_memory=True and observe that they don’t.
B. Pipeline. Build a 3-stage pipeline (H2D → compute → D2H) with 3 streams and double buffering, processing 20 chunks. Compare total time to the serial version. Compute the theoretical speedup and compare.
C. Overlap limits. Launch a kernel that uses all SMs and a small kernel on another stream. Measure whether they overlap. Then reduce the first kernel’s grid size until overlap occurs.
D. MPS. If you have a spare GPU, run two model server processes with and without MPS. Compare aggregate throughput.
E. Find the sync. Take a piece of ML code, profile it, and find every synchronization point. Remove the unnecessary ones and measure.
11. Interview questions#
- What is a CUDA stream and what ordering does it guarantee?
- Why don’t kernels on different streams always overlap?
- What is an event and how does it differ from
synchronize()? - When is pinned memory required?
- What is MPS and when would you use it for inference?
- Give three places in an inference stack where streams matter.
- What does stream priority do and what doesn’t it do?
12. Further reading#
- [REFERENCE] CUDA C++ Programming Guide, “Streams and Events”
- [REFERENCE] NVIDIA MPS documentation
- [REFERENCE] Mark Harris, “How to Overlap Data Transfers in CUDA C/C++”
- Next: 07 — CUDA graphs