Infrastructure · Software component
Instance Group Autoscaler
Software componentInfrastructureInfrastructurearc:InstanceGroupAutoscaler
An infrastructure control component that adds or removes compute instances in a cloud instance group when monitored metrics cross thresholds, so paid capacity tracks demand.
Responsibility. Adjusts the number of compute instances provisioned to match demand.
Also known as: Cloud auto-scaling group, Elastic instance scaling
Relationships
is configured by structural
invokes dependency
scales control
- CPU Compute Node Ch4.7
- GPU Node abstract Ch4.7
Design guidance
- SHOULD monitor CPU, request queue depth, or agent-specific metrics (inference latency, success rate) to add and remove instances.
Quantitative guidance
As stated by the sources; verify before use.
- Static provisioning for 1,000 req/s peak vs 100 req/s overnight wastes 90% of overnight resources (Ch4.7).
Classification
- Patterns
- Elastic scalingHysteresis (fast scale-up, delayed scale-down)
- Quality attributes
- Cost efficiencyPerformance efficiency (ISO/IEC 25010)
- Risks mitigated
- Paying for idle peak capacityInstance churn from aggressive scale-down
Sources
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.