Observability & Evaluation · Software component
Bottleneck Analyzer
Software componentObservability & EvaluationObservability & Evaluationarc:BottleneckAnalyzer
An analysis component that classifies a workload's dominant bottleneck (inference-, memory-, synchronization- or preprocessing-bound) from timeline signatures and maps it to optimization strategies.
Responsibility. Identifies the dominant performance bottleneck and its remedy.
Also known as: Timeline pattern analysis, Bottleneck diagnosis, GPU-level latency diagnosis, Coordination vs computational bottleneck analysis
Relationships
is configured by structural
is invoked by dependency
reads dependency
Design guidance
- SHOULD target only the current dominant bottleneck per iteration and re-measure, since fixing one bottleneck often exposes the next.
- SHOULD check batch sizes and dynamic batching first when GPU utilization is low.
- SHOULD address host-device copy spikes with pinned memory, on-device residency and prefetching.
- SHOULD address kernel-launch gaps with asynchronous streams, operation fusion or compiled operators.
- SHOULD address periodic stalls from KV-cache reallocation with paged or pre-allocated KV-cache memory.
- SHOULD classify inference slowdowns from correlated GPU utilization, GPU memory and queue-time metrics before changing batch sizes or model configurations.
- SHOULD record dependency wait time separately from execution time so coordination bottlenecks (wait >> execution) are distinguishable from computational ones.
Quantitative guidance
As stated by the sources; verify before use.
- Kernel launch latency typically 10-50 microseconds per launch (Ch4.4).
- Pinned (page-locked) host memory doubles transfer bandwidth versus pageable memory (Ch4.4).
- A single request may use only 10-20% of GPU cores (Ch4.4).
- Root cause of a memory-bound incident identified in ~5 minutes with GPU metrics vs an estimated 2+ hours of trial-and-error without them (Ch8.1).
- Dependency waits over a 10 s threshold are flagged as slow; state reads over 100 ms indicate lock contention (Ch8.2A).
Classification
- Patterns
- Iterative bottleneck eliminationTimeline signature matching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Cost efficiency
- Risks mitigated
- Adding GPUs when configuration is the constraintScattered micro-optimizations
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ch8.2A: T. Nguyen, "Error Rates and Reliability," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2A. ISBN: 9798244538229.