Infrastructure · Data artifact
Autoscaling Policy
Data artifactInfrastructureInfrastructurearc:AutoscalingPolicy
A declarative specification of replica bounds, scaling metrics and targets, and scale-up/scale-down behaviour policies for a workload.
Responsibility. Declares when, how far and how fast a workload may scale.
Also known as: HPA manifest, Scaling policy, Cost guardrail, HorizontalPodAutoscaler spec, HPA specification
Relationships
configures structural
Design guidance
- MUST set maxReplicas from budget constraints; unbounded ceilings risk runaway cost under DoS or bugs.
- SHOULD scale up fast and scale down slowly with separate policies and a 5+ minute scale-down stabilization window.
- SHOULD use workload metrics such as queue depth and request latency in addition to CPU and memory.
- SHOULD cap maxReplicas for cost control so that overload degrades latency rather than letting cost spiral.
- SHOULD scale up by the more aggressive of percent/pod policies and scale down by the more conservative.
Quantitative guidance
As stated by the sources; verify before use.
- Worked example: min 3, max 20 replicas; 20 x L4 at $0.60/h caps cost at $12/h (Ch1.8).
- maxReplicas 100 at $2/h could cost $4,800/day under runaway scaling (Ch1.8).
- Production-tested targets 60-75% CPU; example scale-up 50% per 60 s, scale-down 1 pod per 120 s after 300 s stabilization (Ch4.3).
- Scale out when 1,000+ tasks are pending or median latency exceeds 5 s (Ch4.3).
- Worked example: maintain 40-80% CPU; minimum 2 instances off-peak for availability, 3 at peak (Ch4.2).
- Worked HPA: minReplicas 2, maxReplicas 20, CPU target 70%, memory target 75%; scaleUp stabilization 60 s, max(50%/min, 4 pods/min); scaleDown stabilization 300 s, min(10%/5 min, 1 pod/5 min) (Ch4.7).
- Scale-up behaviour grows 3 replicas to 4-5 within a minute and to 7-8 the next minute under sustained load (Ch4.7).
- stabilizationWindowSeconds 0 for scale-up and 300 for scale-down (Ch7.2).
- minReplicas 1, maxReplicas 40, memory utilization target 75%, scale-down stabilization window 300 s, scale down 50% at a time (Ref8.05).
Classification
- Patterns
- Asymmetric behavior policyBudget-derived replica ceilingAsymmetric scale-up/scale-downStabilization windows
- Technologies
- Kubernetes HPA manifest
- Quality attributes
- Cost efficiencyReliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Runaway scaling from bugs or attacksScaling oscillationScaling thrash (flapping)Slow response to demand spikes
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note