1. What is it?#
Measuring what the GPU actually did, at two levels:
Nsight SYSTEMS the timeline. What ran when, on which stream, with what gaps.
→ "where is the time going across the whole application?"
Nsight COMPUTE one kernel, in depth. Occupancy, memory throughput, stalls, roofline.
→ "why is THIS kernel slow?"Use Systems first to find the problem, Compute second to understand it.
2. Why does it exist?#
Because GPU performance is opaque. The kernel took 4 ms — was that good? Was it memory-bound? Was the GPU even busy? Was the CPU the bottleneck? Without a profiler you’re guessing, and intuition about GPU performance is reliably wrong.
3. Simple analogy#
Nsight Systems is the flight recorder; Nsight Compute is the engine teardown.
The recorder tells you the plane was slow between waypoints 4 and 5. The teardown tells you a turbine blade was fouled. You need the recorder first — tearing down the wrong engine wastes days.
4. Nsight Systems — the timeline#
nsys profile -t cuda,nvtx,osrt,cudnn,cublas \
--cuda-memory-usage=true \
-o report python serve.py
nsys stats report.nsys-rep # text summary
nsys-ui report.nsys-rep # GUI timeline
Key views and what they tell you:
CUDA API row cudaLaunchKernel, cudaMemcpy, cudaStreamSynchronize
→ long cudaStreamSynchronize = the CPU is waiting for the GPU
→ many closely-packed cudaLaunchKernel = you may be launch-bound
Kernels row every kernel, its duration, its stream
→ GAPS between kernels = the GPU is idle. Why?
Memory row H2D and D2H transfers
→ transfers on the critical path = design problem
Threads rows CPU threads and their state
→ a Python thread at 100% while kernels have gaps = launch-bound
NVTX row your own annotations ← add these!Annotate your code. Without NVTX ranges the timeline is an undifferentiated wall of kernels:
import torch.cuda.nvtx as nvtx
with nvtx.range("prefill"):
out = model(prompt_ids)
with nvtx.range("decode_loop"):
for i in range(n):
with nvtx.range(f"step"):
with nvtx.range("forward"): out = model(tok, cache)
with nvtx.range("sample"): tok = sample(out.logits)
Now the timeline reads as your program rather than as an opaque kernel sequence. This takes ten minutes and pays for itself immediately.
What to look for, in order#
1. Gaps between kernels → launch-bound, or CPU work on the critical path
2. Long cudaStreamSynchronize → implicit sync (an .item() somewhere)
3. Memory copies in the loop → should be zero in steady-state decode
4. Serialized streams → missing overlap
5. One kernel dominating → go to Nsight Compute for that kernel
6. Many tiny kernels → fusion or CUDA graph opportunity5. Nsight Compute — the kernel#
# Profile a specific kernel, full detail
ncu --set full -k regex:"flash.*" -c 5 -o kernel_report python bench.py
# Just the roofline
ncu --set roofline -k my_kernel ./app
# Specific metrics (fast, scriptable)
ncu --metrics \
sm__throughput.avg.pct_of_peak_sustained_elapsed,\
dram__throughput.avg.pct_of_peak_sustained_elapsed,\
l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\
l1tex__t_requests_pipe_lsu_mem_global_op_ld.sum,\
sm__warps_active.avg.pct_of_peak_sustained_active \
./app
The five numbers that diagnose any kernel#
1. sm__throughput.pct_of_peak compute utilization
2. dram__throughput.pct_of_peak memory utilization
3. sm__warps_active.pct_of_peak achieved occupancy
4. sectors / requests coalescing quality (4 = ideal)
5. dominant warp stall reason what it's waiting onThe diagnostic table:
sm% dram% diagnosis action
────────────────────────────────────────────────────────────────────────
high low compute bound lower precision, fewer FLOPs
low high memory bound, at the roof fewer bytes: quantize, fuse
low low LATENCY bound more occupancy, more ILP,
better access patterns
high high well balanced you're doneThe “low/low” case is the interesting one and the most common for badly-written kernels: the GPU is neither computing nor moving data at capacity — it’s stalled. Look at the stall reasons.
The stall reasons, again#
Long Scoreboard waiting on global memory → occupancy, coalescing, prefetch
Short Scoreboard waiting on shared memory → bank conflicts
Barrier __syncthreads → fewer barriers, more work between
MIO Throttle memory instruction queue → fewer, wider memory ops
Math Throttle ALU saturated → compute bound (this is good)
Not Selected ready, but not chosen → healthy; you have parallelism6. PyTorch profiler — the accessible option#
from torch.profiler import profile, record_function, ProfilerActivity, schedule
with profile(
activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
schedule=schedule(wait=1, warmup=1, active=3),
on_trace_ready=torch.profiler.tensorboard_trace_handler('./log'),
record_shapes=True, profile_memory=True, with_stack=True,
) as prof:
for _ in range(5):
with record_function("inference"):
model(x)
prof.step()
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=20))
prof.export_chrome_trace("trace.json") # view in chrome://tracing or perfetto.dev
Advantages: no external tools, understands PyTorch operators, shows shapes and memory. Good enough for 80% of questions.
The key diagnostic from its table:
Self CPU time total: 850 ms
Self CUDA time total: 240 ms
→ CPU-bound. The GPU was idle 72% of the time.7. A worked diagnosis#
SYMPTOM: decode ITL is 18 ms; the roofline predicts 6 ms.
STEP 1 — Nsight Systems
Observation: kernels occupy 7 ms of each 18 ms step; 11 ms of gaps.
Conclusion: not a kernel problem. Something else is eating 11 ms.
STEP 2 — look at the CPU rows
Observation: a Python thread is at 100% throughout the gaps.
350 cudaLaunchKernel calls per step.
Conclusion: launch/dispatch bound.
STEP 3 — check for implicit sync
Observation: one cudaStreamSynchronize per step, 4 ms.
NVTX shows it inside "sample".
Root cause: the sampler calls .item() to check the stop condition.
STEP 4 — fixes
a) move the stop check to the GPU (or check every 8 tokens) → -4 ms
b) enable CUDA graphs → -6 ms
New ITL: ~8 ms. Remaining gap vs the 6 ms roofline is kernel efficiency.
STEP 5 — Nsight Compute on the top kernel
Observation: dram__throughput 71% of peak; sectors/request 4.0.
Conclusion: well-coalesced, near the memory roof. 71% is respectable.
Remaining action: reduce bytes (quantize), not tune the kernel.That sequence — timeline first, CPU rows second, sync points third, kernel detail last — is the method. Follow it and you’ll rarely waste effort.
8. Production implications#
- Profile with realistic workloads. A synthetic fixed-length benchmark hides the problems your production traffic has.
- Add NVTX ranges permanently — they cost nothing when not profiling.
- Keep a reference profile per release. Comparing two profiles finds regressions instantly; reading one profile is much harder.
- Profiling overhead:
nsys5-15%,ncucan be 10-100x for a full metric set (it replays kernels). Usencuon isolated benchmarks,nsyson realistic runs. - Profile in the container you deploy, with the same driver and library versions.
9. Common mistakes#
Going straight to Nsight Compute. You’ll optimize a kernel that isn’t the problem.
Profiling without warmup. The first iterations include compilation, allocation, and autotuning.
Not annotating with NVTX. The timeline is unreadable without it.
Profiling a toy workload. Different shapes, different kernels, different conclusions.
Ignoring the CPU rows. Half of inference problems are CPU-side.
Comparing profiles from different driver/library versions and concluding your code regressed.
10. Hands-on exercise#
A. Annotate and profile. Add NVTX ranges to a generation loop (prefill, each decode step,
forward, sample, detokenize). Profile with nsys and produce a timeline. What fraction of wall
time is the GPU busy?
B. Find the gaps. Measure sum(kernel_durations) / wall_clock for your decode loop at batch
1 and batch 64. Explain the difference.
C. Diagnose a kernel. Pick the top kernel by time. Get the five numbers from section 5. Classify it using the diagnostic table. What would you do next?
D. Find an implicit sync. Deliberately add a .item() to a decode loop. Find it in the
Nsight Systems timeline. Measure its cost. Remove it.
E. Build the roofline. Use ncu --set roofline on three kernels of different character.
Place them on a roofline plot. Which has the most headroom?
F. Regression detection. Profile the same workload twice with a deliberate change (e.g. disabling CUDA graphs). Diff the summaries. How quickly can you identify what changed?
11. Interview questions#
- When do you use Nsight Systems vs Nsight Compute?
- What are the five metrics you’d check first for a slow kernel?
sm__throughputanddram__throughputare both 20%. What does that mean and what do you do?- How do you tell whether a workload is launch-bound?
- What is NVTX and why does it matter?
- Walk me through diagnosing an ITL that’s 3x the roofline prediction.
- Why is
ncuunsuitable for profiling a production server?
12. Further reading#
- [REFERENCE] Nsight Systems and Nsight Compute user guides
- [REFERENCE] Nsight Compute metrics reference (learn the naming scheme)
- [REFERENCE] PyTorch Profiler documentation and the HTA (Holistic Trace Analysis) tool
- [REFERENCE] Perfetto UI (ui.perfetto.dev) for viewing Chrome traces
- Next: Section VII — Inference Optimization