Section VI.12 introduced the tools. This file is a working recipe book for inference-specific profiling.
1. The two tools, and when to use each#
NSIGHT SYSTEMS (nsys) the timeline. "Where is the time going?"
overhead 5-15%. Use on realistic workloads.
→ step 3 of the methodology
NSIGHT COMPUTE (ncu) one kernel, deeply. "Why is THIS kernel slow?"
overhead 10-100x (it replays kernels).
Use on isolated microbenchmarks.
→ step 5 of the methodologyNever run ncu against a production server. It serializes and replays kernels; the server
will appear to hang.
2. Recipe 1 — profiling a decode loop with Nsight Systems#
nsys profile \
--trace=cuda,nvtx,osrt,cublas,cudnn \
--cuda-memory-usage=true \
--capture-range=cudaProfilerApi \
--capture-range-end=stop \
-o decode_profile \
python bench_decode.pyWith NVTX annotations in the code:
import torch.cuda.nvtx as nvtx
import torch.cuda.profiler as profiler
# warm up first
for _ in range(20): engine.step()
profiler.start() # ← capture starts here
for i in range(10):
with nvtx.range(f"step_{i}"):
with nvtx.range("schedule"): out = engine.schedule()
with nvtx.range("prepare"): batch = engine.prepare(out)
with nvtx.range("forward"): logits = engine.forward(batch)
with nvtx.range("sample"): toks = engine.sample(logits)
with nvtx.range("update"): engine.update(toks)
profiler.stop()The --capture-range=cudaProfilerApi flag is important: it profiles only the annotated
region, keeping the trace small and excluding warmup.
Reading the result#
nsys stats decode_profile.nsys-rep** CUDA GPU Kernel Summary:
Time(%) Total Time Instances Avg (ns) Name
------- ------------ --------- ---------- -----------------------------------
31.2 142,331,4 320 444,786 void cutlass::Kernel<...gemm...>
18.7 85,220,1 320 266,313 paged_attention_v2_kernel
12.4 56,510,8 640 88,298 void vllm::fused_add_rms_norm_kernel
9.1 41,432,0 320 129,475 void vllm::silu_and_mul_kernel
...
** CUDA API Summary:
Time(%) Total Time Num Calls Avg (ns) Name
------- ------------ --------- --------- ---------------------------
62.1 412,220,1 10 41,222,0 cudaGraphLaunch
8.2 54,110,2 3200 16,909 cudaLaunchKernel
...What to look for, in order:
1. sum(kernel time) vs the NVTX "step" range duration
→ the gap is CPU time or waiting
2. cudaStreamSynchronize / cudaMemcpy in the API summary
→ implicit syncs; each one is a stall
3. The top 5 kernels by total time
→ your optimization targets
4. cudaLaunchKernel count
→ if high and graphs are supposedly on, capture failed
5. Gaps in the GPU timeline (visual, in the GUI)
→ launch-bound3. Recipe 2 — the “where are the gaps” analysis#
# From the nsys SQLite export
import sqlite3
con = sqlite3.connect("decode_profile.sqlite")
# total kernel time
kt = con.execute("""
SELECT SUM(end - start) FROM CUPTI_ACTIVITY_KIND_KERNEL
""").fetchone()[0]
# wall time of the profiled region
wall = con.execute("""
SELECT MAX(end) - MIN(start) FROM CUPTI_ACTIVITY_KIND_KERNEL
""").fetchone()[0]
print(f"GPU busy: {kt/wall*100:.1f}% gaps: {(wall-kt)/1e6:.1f} ms")GPU busy: 94% → GPU-bound. Go to Nsight Compute.
GPU busy: 41% → 59% of the time the GPU is IDLE. Find out why:
- CPU too slow (check the thread rows)
- implicit sync (check the API row)
- waiting on a collective (check NCCL rows)
- waiting on the network (check the OS runtime rows)This single number — GPU busy percentage — resolves the “is it the GPU?” question definitively, and it’s the number to compute first from any timeline.
4. Recipe 3 — deep kernel analysis with Nsight Compute#
Isolate the kernel in a microbenchmark first, then:
ncu --set full \
--kernel-name regex:"paged_attention" \
--launch-count 3 \
--launch-skip 10 \
-o attn_report \
python microbench_attention.py# Or, for scriptable output:
ncu --metrics \
sm__throughput.avg.pct_of_peak_sustained_elapsed,\
dram__throughput.avg.pct_of_peak_sustained_elapsed,\
sm__warps_active.avg.pct_of_peak_sustained_active,\
l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\
l1tex__t_requests_pipe_lsu_mem_global_op_ld.sum,\
l1tex__data_bank_conflicts_pipe_lsu_mem_shared_op_ld.sum,\
launch__registers_per_thread,\
smsp__warp_issue_stalled_long_scoreboard_per_warp_active.pct \
--csv -k regex:"paged_attention" python microbench.pyThe interpretation checklist#
sm__throughput% > 70 → compute bound
dram__throughput% > 70 → memory bound
both < 40 → latency bound
sectors/requests = 4 → perfectly coalesced (32-bit loads)
= 32 → fully uncoalesced
warps_active% < 25 → low occupancy; check registers/shared memory
registers_per_thread > 128 → likely limiting occupancy
bank_conflicts > 0 → shared memory access pattern problem
long_scoreboard% > 40 → stalled on global memory5. Recipe 4 — profiling a multi-GPU (TP) run#
# Profile all ranks; nsys handles MPI/multi-process
nsys profile --trace=cuda,nvtx,nccl \
--output=tp_profile_%p \
torchrun --nproc_per_node=4 serve.pyWhat to look for:
1. NCCL kernel time per rank
→ should be similar across ranks; large variance means imbalance
2. Gaps before NCCL kernels
→ a rank arrived late; find out why (it did more work, or was descheduled)
3. NCCL kernel duration vs the compute between collectives
→ your TP overhead, measured
4. Are the ranks in lockstep?
→ in the multi-row timeline, the collectives should alignA rank consistently arriving late at collectives is the signature of load imbalance or CPU contention on that rank — a common problem when the scheduler runs on rank 0.
6. Recipe 5 — the PyTorch profiler (when Nsight is too heavy)#
from torch.profiler import profile, ProfilerActivity, schedule, tensorboard_trace_handler
with profile(
activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
schedule=schedule(skip_first=10, wait=1, warmup=2, active=5, repeat=1),
on_trace_ready=tensorboard_trace_handler('./tb_logs'),
record_shapes=True,
profile_memory=True,
with_stack=True,
) as prof:
for _ in range(20):
engine.step()
prof.step()
print(prof.key_averages(group_by_input_shape=True).table(
sort_by="self_cuda_time_total", row_limit=25))
prof.export_chrome_trace("trace.json") # view at ui.perfetto.devAdvantages: understands PyTorch operators, shows shapes, tracks memory, no external tool. Sufficient for 80% of investigations. Use Nsight when you need kernel-level detail or multi-process timelines.
The group_by_input_shape=True option is particularly useful for inference — it shows you
whether one shape is disproportionately expensive.
7. What a healthy inference profile looks like#
DECODE STEP, batch 64, well-tuned:
GPU busy: 92-96%
Top kernel (GEMM): 30-35% of GPU time
paged_attention: 15-25%
fused norms: 8-12%
activation: 5-8%
sampling: 2-4%
NCCL (if TP): 8-15%
cudaLaunchKernel count: ~1 (graph replay)
cudaStreamSynchronize: 1 per step (at the end)
Memory copies: ~0 in steady state
RED FLAGS:
✗ GPU busy < 70% → CPU or waiting
✗ many cudaLaunchKernel → graphs not active
✗ cudaMemcpy in the loop → data movement per step
✗ multiple cudaStreamSync → implicit syncs
✗ a kernel you don't recognize taking 20%+
✗ NCCL > 30% of time → interconnect problem or TP too highPrint this list and compare against it. Deviations are your investigation list.
8. Production implications#
- Add NVTX annotations permanently. They cost nothing when not profiling and make every future investigation faster.
- Profile in the deployment container, with the same driver and library versions.
- Keep a reference profile per release. Diffing two profiles finds regressions in minutes.
- Use
nsyson realistic workloads,ncuon microbenchmarks. - Automate the “GPU busy %” calculation and track it as a metric.
- Never run
ncuagainst production.
9. Common mistakes#
Profiling without warmup. Captures JIT, autotuning, and allocator growth.
Profiling a synthetic workload. Different shapes → different kernels → different conclusions.
Running ncu on a server. Appears to hang.
Not annotating with NVTX. The timeline is unreadable.
Reading only the kernel summary and ignoring the gaps.
Comparing profiles across different library versions and concluding your code regressed.
Profiling one rank of a multi-GPU job and missing the imbalance.
10. Hands-on exercise#
A. Annotate and profile. Add NVTX ranges to a real decode loop. Profile with nsys. Compute
the GPU busy percentage. Produce the kernel time breakdown.
B. Find a gap. Deliberately introduce a .item() in the decode loop. Find it in the
timeline. Measure its cost. Remove it and confirm.
C. Deep-dive a kernel. Take the top kernel by time. Isolate it in a microbenchmark. Run ncu --set full. Fill in the interpretation checklist from section 4. What’s the diagnosis?
D. The healthy-profile comparison. Profile your system and compare against section 7’s reference. List every deviation. Investigate the largest.
E. Multi-GPU. If you have TP hardware, profile all ranks. Are they in lockstep? What fraction of time is NCCL? Does it match your Section IX.09 prediction?
F. Regression detection. Profile before and after a configuration change. Diff the kernel summaries. How quickly can you identify what changed?
11. Interview questions#
- When do you use Nsight Systems vs Nsight Compute?
- How do you compute “GPU busy percentage” and what does it tell you?
- What does a healthy decode profile look like? Name four red flags.
- Why can’t you run
ncuagainst a production server? - In a multi-GPU profile, what does a rank arriving late at collectives indicate?
- What NVTX annotations would you add to an inference engine?
- Walk through diagnosing a decode step that’s 3x its roofline prediction.
12. Further reading#
- [REFERENCE] Nsight Systems and Nsight Compute user guides
- [REFERENCE] Nsight Compute metrics reference and the “Kernel Profiling Guide”
- [REFERENCE] PyTorch Profiler and Holistic Trace Analysis (HTA)
- [REFERENCE] Perfetto UI for viewing Chrome traces
- Next: 09 — CPU and Python profiling