Infrastructure · Software component
Custom Metrics Adapter
Software componentInfrastructureInfrastructurearc:CustomMetricsAdapter
A bridge that exposes collected application metrics through the container orchestrator's metrics API so autoscalers can scale on inference-specific signals.
Responsibility. Makes monitoring-system metrics consumable by the orchestrator's autoscaler.
Also known as: Custom metrics API, Custom metrics exporter
Relationships
receives data from dynamic
sends data to dynamic
Design guidance
- SHOULD expose queue depth, inference latency and business metrics as scaling signals beyond built-in CPU/memory metrics.
Classification
- Technologies
- Kubernetes custom metrics APIPrometheus
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note