The idea in one minute#
Launching a kernel does not run it. It puts a command in a queue and returns immediately; the GPU works through the queue on its own time. The host only waits when it synchronizes — explicitly, or implicitly by asking for a result.
A queue of commands is called a stream. Commands in one stream run in order. Commands in different streams may overlap. Understanding this is the difference between correct and nonsense timings, and between a busy GPU and an idle one.
An analogy#
Ordering at a counter. You place the order and get a receipt at once — that is the launch returning. The kitchen cooks in the order tickets arrived. You only wait when you stand at the counter until your food appears — that is synchronizing.
If you time “how long until I got my receipt”, you have measured the cashier, not the kitchen.
A picture#
sequenceDiagram participant H as Host (Go) participant Q as Stream (queue) participant G as GPU H->>Q: launch kernel A Note right of H: returns in microseconds H->>Q: launch kernel B H->>Q: launch kernel C Q->>G: run A Note over H: host is free to do other work Q->>G: run B Q->>G: run C H->>Q: synchronize Note over H: host blocks here G-->>H: all done
How it really works#
Asynchronous by default#
In CUDA, kernel launches are asynchronous. Some calls block until earlier work is done:
- an explicit synchronize (
cudaDeviceSynchronize,cudaStreamSynchronize), - a normal (synchronous) copy from device to host — you cannot copy a result that does not exist yet,
- anything that reads a GPU value on the host (printing a tensor element, converting to a Go number).
The timing trap#
t0 := now()
launch(kernel) // returns immediately
elapsed := now() - t0 // you measured the launch (microseconds), not the kernelCorrect timing synchronizes before starting the clock and before stopping it. This is the single most common mistake in GPU benchmarks; when someone reports an impossibly fast result, check for it first.
Hidden synchronization kills throughput#
The opposite problem: code that synchronizes when it did not mean to. Reading one number from the GPU inside a loop forces the host to wait for the GPU every iteration, and the GPU to sit idle while the host does its part. A pipeline that could overlap host and device work becomes strictly alternating.
Rule: keep data on the device and do not look at it until you must.
Streams#
A stream is an ordered queue. Within a stream, operations run one after another. Between streams there is no ordering, so the GPU can overlap them — for example copying the next batch in while the current batch computes.
Events are markers you place in a stream to wait for a specific point or to measure the time between two points on the GPU’s own clock.
Streams allow overlap but do not create capacity. If one kernel already uses every SM, a second stream’s kernel waits its turn. Streams help when work is small enough to share the GPU, and for overlapping copies with compute.
Go already has this model#
A stream behaves like a goroutine fed by a buffered channel: sends return immediately, work is processed in order, and you wait only when you ask for a reply.
Code#
// stream.go — a stream is an ordered queue; launching is not running.
package main
import (
"fmt"
"time"
)
type Stream struct{ q chan func() }
func NewStream() *Stream {
s := &Stream{q: make(chan func(), 1024)}
go func() { // the "GPU" draining this stream in order
for op := range s.q {
op()
}
}()
return s
}
// Launch enqueues work and returns immediately.
func (s *Stream) Launch(d time.Duration) { s.q <- func() { time.Sleep(d) } }
// Synchronize blocks until everything queued so far has finished.
func (s *Stream) Synchronize() {
done := make(chan struct{})
s.q <- func() { close(done) }
<-done
}
func main() {
s := NewStream()
// WRONG: times the launch
t0 := time.Now()
s.Launch(50 * time.Millisecond)
fmt.Println("without sync:", time.Since(t0).Round(time.Microsecond))
// RIGHT: sync, start clock, launch, sync, stop clock
s.Synchronize()
t0 = time.Now()
s.Launch(50 * time.Millisecond)
s.Synchronize()
fmt.Println("with sync: ", time.Since(t0).Round(time.Millisecond))
// Two streams overlap; one stream is strictly sequential.
a, b := NewStream(), NewStream()
t0 = time.Now()
a.Launch(50 * time.Millisecond)
b.Launch(50 * time.Millisecond)
a.Synchronize()
b.Synchronize()
fmt.Println("two streams: ", time.Since(t0).Round(time.Millisecond))
}
Remember this#
- Launching a kernel enqueues it. The host continues immediately.
- To time GPU work, synchronize before starting and before stopping the clock.
- Reading GPU data on the host is an implicit synchronization. Avoid it in hot loops.
- A stream is an ordered queue; different streams may overlap.
Try it#
- In
stream.go, launch ten 10 ms kernels on one stream, then on ten streams. Compare. - Add an
Eventtype withRecord()andElapsed(other)so you can time the gap between two points inside a stream. - A decode loop reads the chosen token back to the host each step to check for end-of-text. Explain what that costs and suggest an alternative.
Check yourself#
- Why does an unsynchronized timing give a wrong (too small) result?
- Name two operations that synchronize implicitly.
- When do multiple streams not improve throughput?