Observability & Evaluation · Software component

GPU Telemetry Exporter

Software componentObservability & EvaluationObservability & Evaluationarc:GPUTelemetryExporter

A node-level telemetry component that samples accelerator streaming-multiprocessor utilization, memory usage and bandwidth, and power, and exports them to metrics and profiling back ends.

Responsibility. Exports GPU utilization, memory and power metrics.

Also known as: GPU metrics sampler, GPU utilization metric source, DCGM Exporter, GPU metrics exporter, GPU telemetry agent, DCGM agent

deployed on; monitorsemits telemetry toemits telemetry todeployed onmonitorsis orchestrated byexposesGPU Node: deployed on; monitorsGPU NodeMetrics Collector: emits telemetry toMetrics CollectorTime-Series Metrics Store: emits telemetry toTime-Series Metrics StoreEdge GPU Device: deployed onEdge GPU DeviceGPU Partition: monitorsGPU PartitionAccelerator Operator: is orchestrated byAccelerator OperatorMetrics Endpoint: exposesMetrics Endpoint
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

exposes structural

emits telemetry to dynamic

is orchestrated by control

monitors assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Pull-based metrics expositionConfigurable metric sets
Technologies
nvidia-smiNVIDIA Nsight Systems GPU metricsNVIDIA DCGMNVIDIA DCGM ExporterHelmnvidia-smi dmonDCGM Python bindingsPrometheus client gauges
Quality attributes
Maintainability (ISO/IEC 25010)Cost efficiencyPerformance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Average utilization masking intermittent GPU idlenessUndetected GPU underutilisation or saturationUndetected thermal throttlingUndetected GPU memory leaksUndetected hardware errors

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  2. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  3. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.
  4. Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
  5. Ref4.05: NVIDIA, "Setting up Prometheus," NVIDIA GPU Telemetry Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-telemetry/latest/kube-prometheus.html
  6. Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html
  7. Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
  8. Ref7.16: "Production Monitoring and Operations for Agentic AI," unpublished reference note (16-Production-Monitoring-Operations.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  9. Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note