Infrastructure · Software component

Metric-Driven Autoscaler

Software componentInfrastructureInfrastructurearc:MetricDrivenAutoscaler

A reactive autoscaler that computes desired replicas from several live metrics (queue-to-compute ratio, GPU utilization, request rate) and applies asymmetric stabilization windows.

Responsibility. Computes desired replica counts from live demand metrics.

Also known as: Queue-based autoscaler, Multi-metric HPA, Economic-aware autoscaler, Horizontal pod autoscaler, CPU-utilisation autoscaler, HorizontalPodAutoscaler, Horizontal Pod Autoscaler on inference metrics, Queue-to-compute ratio autoscaler, Horizontal Pod Autoscaler (multi-metric)

Variant of Autoscaler abstract

When to choose. Choose for stateless agents with unpredictable, high-variance traffic (e.g., multi-tenant SaaS) that tolerate 60-90 s scale-up lag.

scalesscalesscalesscalesinvokesscalesreadsdeployed oninvokesmonitorsspecializesscalesalternative tois configured byscalesreceives data fromAgent Controller: scalesAgent ControllerLLM Inference Service: scalesLLM Inference ServiceInference Server: scalesInference ServerWorker Agent: scalesWorker AgentMetrics Collector: invokesMetrics CollectorAnswer Synthesizer: scalesAnswer SynthesizerTime-Series Metrics Store: readsTime-Series Metrics StoreContainer Orchestrator: deployed onContainer OrchestratorToken Cost Meter: invokesToken Cost MeterEvent Broker: monitorsEvent BrokerAutoscaler: specializesAutoscalerSpot GPU Node: scalesSpot GPU NodeScheduled Scaler: alternative toScheduled ScalerAutoscaling Policy: is configured byAutoscaling PolicyStateless Workload Controller: scalesStateless Workload Contr…Custom Metrics Adapter: receives data fromCustom Metrics Adapter
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is configured by structural

invokes dependency

reads dependency

receives data from dynamic

scales control

monitors assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Queue-based autoscalingMulti-metric max-replica selectionAsymmetric scale-up/scale-downStabilization windowsCost-per-request-aware scaling
Technologies
Kubernetes Horizontal Pod AutoscalerPrometheusKubernetes HPA
Quality attributes
Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Cost efficiency
Risks mitigated
Scaling oscillation (thrashing)Misleading CPU-utilization signalsRunaway scalingOverprovisioning for peak load (high idle cost)

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  3. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  4. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  5. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  6. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  7. Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
  8. Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
  9. Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
  10. Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  11. Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note