Observability & Evaluation · Software component
GPU Telemetry Exporter
Software componentObservability & EvaluationObservability & Evaluationarc:GPUTelemetryExporter
A node-level telemetry component that samples accelerator streaming-multiprocessor utilization, memory usage and bandwidth, and power, and exports them to metrics and profiling back ends.
Responsibility. Exports GPU utilization, memory and power metrics.
Also known as: GPU metrics sampler, GPU utilization metric source, DCGM Exporter, GPU metrics exporter, GPU telemetry agent, DCGM agent
Relationships
deployed on structural
- Edge GPU Device Ch4.6
- GPU Node abstract Ch4.4 Ch4.5 +3
exposes structural
- Metrics Endpoint abstract Ch8.1
emits telemetry to dynamic
is orchestrated by control
monitors assurance
- GPU Node abstract Ch4.4 Ch4.5 +6
- GPU Partition Ch7.6
Design guidance
- SHOULD be read alongside timeline data because average utilization hides intermittent idle periods during tool execution.
- SHOULD be deployed on every GPU node and discovered automatically by the metrics collector.
- SHOULD report GPU utilization, memory usage and fragmentation, temperature/thermal throttling, power and clock speeds.
- SHOULD monitor power throttling and temperature, keeping GPUs below ~85 C to avoid thermal throttling (Ref7.01).
- SHOULD report utilisation and memory per partition instance so per-profile anomalies are detectable (Ch7.6).
- SHOULD poll GPU compute utilization, memory utilization, power, temperature and PCIe throughput into a driver-level cache so multiple consumers do not multiply GPU query overhead.
- SHOULD label every GPU metric with a GPU identifier to expose per-GPU load imbalance on multi-GPU servers.
Quantitative guidance
As stated by the sources; verify before use.
- Default metrics include DCGM_FI_DEV_GPU_UTIL, FB_USED/FB_FREE, GPU_TEMP, POWER_USAGE, PCIE_RX/TX_THROUGHPUT; collection interval configurable (e.g., 1000 ms) (Ref4.05).
- Healthy inference workloads show 70-95% GPU utilization; high variance indicates inefficient batching (Ref4.05).
- Polling interval 1-10 s; production typically 10 s (Ch8.1).
- An 8-GPU server averaging 70% utilization with GPU 0 at 98% and GPU 7 at 15% indicates load imbalance (Ch8.1).
- Target 75-85% GPU utilization: < 50% signals right-sizing opportunity, 85-95% running hot, > 95% risk of throttling/failures; thermal target 70-80 C, throttling likely above 85 C (Ref8.05).
Classification
- Patterns
- Pull-based metrics expositionConfigurable metric sets
- Technologies
- nvidia-smiNVIDIA Nsight Systems GPU metricsNVIDIA DCGMNVIDIA DCGM ExporterHelmnvidia-smi dmonDCGM Python bindingsPrometheus client gauges
- Quality attributes
- Maintainability (ISO/IEC 25010)Cost efficiencyPerformance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Average utilization masking intermittent GPU idlenessUndetected GPU underutilisation or saturationUndetected thermal throttlingUndetected GPU memory leaksUndetected hardware errors
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ref4.05: NVIDIA, "Setting up Prometheus," NVIDIA GPU Telemetry Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-telemetry/latest/kube-prometheus.html
- Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
- Ref7.16: "Production Monitoring and Operations for Agentic AI," unpublished reference note (16-Production-Monitoring-Operations.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note