Below the API

Asynchronous Execution and Streams

Intermediate Intermediate 50 min Difficulty 3/5

Prerequisites 02

The idea in one minute#

Launching a kernel does not run it. It puts a command in a queue and returns immediately; the GPU works through the queue on its own time. The host only waits when it synchronizes — explicitly, or implicitly by asking for a result.

A queue of commands is called a stream. Commands in one stream run in order. Commands in different streams may overlap. Understanding this is the difference between correct and nonsense timings, and between a busy GPU and an idle one.

An analogy#

Ordering at a counter. You place the order and get a receipt at once — that is the launch returning. The kitchen cooks in the order tickets arrived. You only wait when you stand at the counter until your food appears — that is synchronizing.

If you time “how long until I got my receipt”, you have measured the cashier, not the kitchen.

A picture#

sequenceDiagram
  participant H as Host (Go)
  participant Q as Stream (queue)
  participant G as GPU
  H->>Q: launch kernel A
  Note right of H: returns in microseconds
  H->>Q: launch kernel B
  H->>Q: launch kernel C
  Q->>G: run A
  Note over H: host is free to do other work
  Q->>G: run B
  Q->>G: run C
  H->>Q: synchronize
  Note over H: host blocks here
  G-->>H: all done

How it really works#

Asynchronous by default#

In CUDA, kernel launches are asynchronous. Some calls block until earlier work is done:

  • an explicit synchronize (cudaDeviceSynchronize, cudaStreamSynchronize),
  • a normal (synchronous) copy from device to host — you cannot copy a result that does not exist yet,
  • anything that reads a GPU value on the host (printing a tensor element, converting to a Go number).

The timing trap#

t0 := now()
launch(kernel)        // returns immediately
elapsed := now() - t0 // you measured the launch (microseconds), not the kernel

Correct timing synchronizes before starting the clock and before stopping it. This is the single most common mistake in GPU benchmarks; when someone reports an impossibly fast result, check for it first.

Hidden synchronization kills throughput#

The opposite problem: code that synchronizes when it did not mean to. Reading one number from the GPU inside a loop forces the host to wait for the GPU every iteration, and the GPU to sit idle while the host does its part. A pipeline that could overlap host and device work becomes strictly alternating.

Rule: keep data on the device and do not look at it until you must.

Streams#

A stream is an ordered queue. Within a stream, operations run one after another. Between streams there is no ordering, so the GPU can overlap them — for example copying the next batch in while the current batch computes.

Events are markers you place in a stream to wait for a specific point or to measure the time between two points on the GPU’s own clock.

Streams allow overlap but do not create capacity. If one kernel already uses every SM, a second stream’s kernel waits its turn. Streams help when work is small enough to share the GPU, and for overlapping copies with compute.

Go already has this model#

A stream behaves like a goroutine fed by a buffered channel: sends return immediately, work is processed in order, and you wait only when you ask for a reply.

Code#

// stream.go — a stream is an ordered queue; launching is not running.
package main

import (
	"fmt"
	"time"
)

type Stream struct{ q chan func() }

func NewStream() *Stream {
	s := &Stream{q: make(chan func(), 1024)}
	go func() { // the "GPU" draining this stream in order
		for op := range s.q {
			op()
		}
	}()
	return s
}

// Launch enqueues work and returns immediately.
func (s *Stream) Launch(d time.Duration) { s.q <- func() { time.Sleep(d) } }

// Synchronize blocks until everything queued so far has finished.
func (s *Stream) Synchronize() {
	done := make(chan struct{})
	s.q <- func() { close(done) }
	<-done
}

func main() {
	s := NewStream()

	// WRONG: times the launch
	t0 := time.Now()
	s.Launch(50 * time.Millisecond)
	fmt.Println("without sync:", time.Since(t0).Round(time.Microsecond))

	// RIGHT: sync, start clock, launch, sync, stop clock
	s.Synchronize()
	t0 = time.Now()
	s.Launch(50 * time.Millisecond)
	s.Synchronize()
	fmt.Println("with sync:   ", time.Since(t0).Round(time.Millisecond))

	// Two streams overlap; one stream is strictly sequential.
	a, b := NewStream(), NewStream()
	t0 = time.Now()
	a.Launch(50 * time.Millisecond)
	b.Launch(50 * time.Millisecond)
	a.Synchronize()
	b.Synchronize()
	fmt.Println("two streams: ", time.Since(t0).Round(time.Millisecond))
}

Remember this#

  • Launching a kernel enqueues it. The host continues immediately.
  • To time GPU work, synchronize before starting and before stopping the clock.
  • Reading GPU data on the host is an implicit synchronization. Avoid it in hot loops.
  • A stream is an ordered queue; different streams may overlap.

Try it#

  1. In stream.go, launch ten 10 ms kernels on one stream, then on ten streams. Compare.
  2. Add an Event type with Record() and Elapsed(other) so you can time the gap between two points inside a stream.
  3. A decode loop reads the chosen token back to the host each step to check for end-of-text. Explain what that costs and suggest an alternative.

Check yourself#

  1. Why does an unsynchronized timing give a wrong (too small) result?
  2. Name two operations that synchronize implicitly.
  3. When do multiple streams not improve throughput?

↑↓ navigate ↵ open