Infrastructure · Software component
Metric-Driven Autoscaler
Software componentInfrastructureInfrastructurearc:MetricDrivenAutoscaler
A reactive autoscaler that computes desired replicas from several live metrics (queue-to-compute ratio, GPU utilization, request rate) and applies asymmetric stabilization windows.
Responsibility. Computes desired replica counts from live demand metrics.
Also known as: Queue-based autoscaler, Multi-metric HPA, Economic-aware autoscaler, Horizontal pod autoscaler, CPU-utilisation autoscaler, HorizontalPodAutoscaler, Horizontal Pod Autoscaler on inference metrics, Queue-to-compute ratio autoscaler, Horizontal Pod Autoscaler (multi-metric)
Variant of Autoscaler abstract
When to choose. Choose for stateless agents with unpredictable, high-variance traffic (e.g., multi-tenant SaaS) that tolerate 60-90 s scale-up lag.
Relationships
deployed on structural
is configured by structural
invokes dependency
reads dependency
receives data from dynamic
scales control
monitors assurance
- Event Broker abstract Ch4.3
alternative to variability
Design guidance
- SHOULD use queue-to-compute ratio as the primary metric, GPU utilization as secondary and request rate as tertiary, taking the highest replica recommendation.
- SHOULD scale up quickly and scale down cautiously (e.g., 60 s vs. 300 s stabilization).
- SHOULD NOT scale on CPU utilization alone for LLM workloads.
- MAY scale up only when projected post-scaling cost per request is lower than current cost per request.
- SHOULD scale inference replicas on inference-aware signals (queue depth, P95 latency, GPU utilisation) rather than CPU or memory alone.
- SHOULD scale up quickly but scale down gradually with stabilisation windows to avoid thrashing.
- SHOULD scale up immediately and scale down gradually to prevent request queueing during spikes and thrashing during dips.
- SHOULD autoscale on combined resource and custom metrics (e.g., queue depth) rather than provisioning permanently for peak load.
Quantitative guidance
As stated by the sources; verify before use.
- Scale up at queue-to-compute 800-1000 milliunits, down at 400-500; GPU up at 75-80%, down at <=50%; request rate up at 50-75% of per-pod capacity (Ch1.8).
- desiredReplicas = ceil(currentReplicas x current/target), evaluated every 15 s (Ch1.8).
- Scale-up adds max(50% of replicas, 2 pods); scale-down removes 10% per minute after 5-minute stabilization (Ch1.8).
- Control loop evaluates metrics every 15 s by default (Ch4.3).
- Example targets: 70% CPU, 80% memory, 50 queued requests per pod; 3-20 replicas (Ch4.3).
- Worked example autoscaling target 40-80% CPU utilisation (Ch4.2).
- CPU-based HPAs typically scale above 70-80% average CPU (Ch4.4).
- Example triggers: queue depth >100 requests for 30s or P95 latency >500ms (Ch4.5).
- NIM HPA: 4-12 replicas at 70% target GPU utilisation; scale-up +50% per 60s (60s window), scale-down 1 pod per 120s (300s window) (Ch4.5).
- Queue-to-compute ratio (queue_time/compute_time): scale up above 1,000 milliunits, scale down below 500 (Ref4.03).
- E-commerce chat: 10 GPUs at 10,000 concurrent conversations vs 1 GPU at 100 (1000 requests/GPU), 90% off-peak cost reduction (Ref4.03).
- Autoscaling 1-8 GPUs: average $3,000/month vs peak $8,000, utilisation 80-90%, ~40% annual saving (Ref4.03).
- Spawns new inference instances when request latency exceeds thresholds or GPU utilization reaches 85% (Ch7.1B).
- Scale-up on CPU >70% or >100 req/s per pod, adding up to 100% more replicas every 30s (Ch7.2).
- Scale-down after 5-minute stabilization, removing up to 50% of replicas per minute (Ch7.2).
- 5 replicas at $2.50/h cost $300/day continuously vs $200/day autoscaled (33% savings); scale-to-zero for dev environments achieves 60-80% savings (Ch7.2).
- Example: 3-20 replicas targeting 70% CPU, 80% memory and an average queue depth of 30 per pod (Ref7.17).
- Scaling from 4 to 12 A100 GPUs at peak ($6-8 to $18-24/hour) averaged ~$12/hour with autoscaling (Ref7.17).
Classification
- Patterns
- Queue-based autoscalingMulti-metric max-replica selectionAsymmetric scale-up/scale-downStabilization windowsCost-per-request-aware scaling
- Technologies
- Kubernetes Horizontal Pod AutoscalerPrometheusKubernetes HPA
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Cost efficiency
- Risks mitigated
- Scaling oscillation (thrashing)Misleading CPU-utilization signalsRunaway scalingOverprovisioning for peak load (high idle cost)
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note