Observability & Evaluation · Software component
GPU System Profiler
Software componentObservability & EvaluationObservability & Evaluationarc:SystemTimelineProfiler
A system-wide timeline profiler that captures GPU kernel execution, SM utilisation, memory bandwidth and allocations, CPU activity, OS runtime waits and annotated code ranges to localise inference bottlenecks.
Responsibility. Reveals which resource constrains inference performance.
Also known as: Timeline profiler, Bottleneck profiler, System-wide profiler, CPU-GPU timeline profiler, System-wide performance profiler, GPU System Profiler
Variant of Performance Profiler abstract
When to choose. Choose for routine and production diagnosis of where time is spent across CPU-GPU boundaries, since event-driven tracing with CPU sampling keeps overhead at about 1-3%.
Relationships
is configured by structural
is invoked by dependency
is triggered by dynamic
receives telemetry from dynamic
monitors assurance
alternative to variability
Design guidance
- SHOULD first measure SM utilisation: above ~80% indicates compute-bound (then check memory bandwidth); below ~80% requires locating idle gaps.
- SHOULD attribute idle gaps before kernel launches to CPU preprocessing, gaps inside annotated tool ranges to tool latency, and gaps at memory allocation to KV-cache pressure.
- SHOULD bracket tool calls and agent steps with annotation ranges so idle GPU intervals can be tied to external invocations.
- SHOULD trace accelerator API, application annotations and OS runtime while disabling high-frequency context-switch tracing to limit overhead.
- MAY run continuously in production given 1-3% overhead, escalating to detailed captures only when a specific issue needs diagnosis.
- SHOULD use the same profiler and workflow across development, production and edge platforms so diagnosis carries over.
- SHOULD run in containers with GPU access for reproducible profiling environments across development, CI/CD and production.
- SHOULD establish a baseline, change one variable at a time, ignore warm-up transients, and use production-like workloads.
- SHOULD profile a baseline to classify the bottleneck (compute, memory, I/O or synchronization) before applying a targeted optimization.
Quantitative guidance
As stated by the sources; verify before use.
- Symptoms: compute-bound SM > 80%; memory-bandwidth-limited DRAM utilisation < 70% with slow inference; KV-cache pressure > 90% GPU memory with batch sizes 1-4 (Ch4.2).
- ReAct step profile: 640 ms = thought generation 320 ms (85% GPU) + synchronous tool call 280 ms (0% GPU) + observation 40 ms (82% GPU) (Ch4.2 case study).
- Overhead typically under 3% (1-3% with default CUDA, NVTX and GPU-metric tracing) (Ch4.4).
- CPU sampling every 1-10 ms builds statistical profiles without continuous-monitoring overhead (Ch4.4).
- Example: 60% average GPU utilization was actually 95% for 6 s then 0% for 4 s, pointing to batching logic, not GPU capacity (Ch4.4).
- ReAct timeline: GPU at 85-87% during LLM thought generation, 0% during ~500 ms tool calls (~50% idle); overlapping tool calls with next-step prompt processing expected ~2x throughput (Ch4.4).
- Optimization targets: >90% GPU utilization on target kernels, >80% of theoretical memory bandwidth (Ref4.04).
Classification
- Patterns
- Profile-before-optimiseBottleneck diagnostic decision treeNVTX range annotationEvent-driven tracingPeriodic CPU call-stack samplingCanary profilingTriggered detailed profilingUnified timeline profilingMulti-node profilingSingle-variable analysis with warm-up and repeated runs
- Technologies
- NVIDIA Nsight SystemsNVTXCUDA memory trackingNVIDIA Nsight Systems 2025.5.1+CUDA Toolkit 12.6NVIDIA Container ToolkitDocker
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)Flexibility (ISO/IEC 25010)
- Risks mitigated
- Optimising the wrong resource (e.g., larger batches when tool latency is the bottleneck)Masking symptoms with configuration changesMisattributing tool/API latency to GPU capacityHidden CPU-GPU synchronization overheadUnnoticed GPU idle timeHost-device synchronization overheadMemory-bandwidth-bound kernelsPCIe bottlenecks
Sources
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ref4.04: NVIDIA, "Nsight Systems," NVIDIA Developer. Accessed: Sep. 27, 2026. [Online]. Available: https://developer.nvidia.com/nsight-systems
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
- Ref7.06: NVIDIA, "Performance Tuning Guide," Megatron Bridge Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/megatron-bridge/latest/performance-guide.html
- Ref7.14: "NVIDIA Agentic AI Platform Ecosystem Integration," unpublished reference note (14-NVIDIA-Ecosystem-Integration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note