Infrastructure · Software component
Autoscaler
Software componentInfrastructureInfrastructureVariation point (abstract)arc:Autoscaler
A control component that adjusts the number of replicas of a workload to match demand, turning fixed infrastructure cost into variable cost.
Responsibility. Adjusts replica counts of workloads to match demand.
Also known as: Pod autoscaler, Horizontal Pod Autoscaler (HPA), Elastic scaling controller, Autoscaling group
Variants
| Variant | When to choose |
|---|---|
| Latency-Target Autoscaler | Choose for latency-sensitive deployments with explicit response-time SLOs. |
| Metric-Driven Autoscaler | Choose for stateless agents with unpredictable, high-variance traffic (e.g., multi-tenant SaaS) that tolerate 60-90 s scale-up lag. |
| Queue-Depth Autoscaler | Choose for agents that primarily orchestrate external API or tool calls, where CPU does not reflect capacity. |
| Resource-Utilization Autoscaler | Choose CPU for compute-bound agents dominated by inference or reasoning loops, and memory for agents holding large context windows or document embeddings. |
| Scheduled Scaler | Choose when traffic patterns are extremely predictable, making schedule-based capacity simpler and equally effective. |
Relationships
deployed on structural
is configured by structural
emits telemetry to dynamic
is triggered by dynamic
scales control
is monitored by assurance
Design guidance
- MUST cap maximum replicas from budget (affordable hourly cost / instance hourly rate), not only technical limits.
- SHOULD keep a minimum replica count that provides baseline capacity and fault tolerance.
- MUST only horizontally scale stateless agent replicas; state kept locally forces session affinity that defeats scaling.
- SHOULD scale up quickly and scale down conservatively, using stabilization windows to prevent oscillation.
- SHOULD keep a minimum replica count (>=2) so warm, redundant capacity always exists during low traffic.
- SHOULD choose scaling metrics that reflect the agent's actual capacity constraint (CPU, memory, queue depth, P95 latency, or cost per query).
- SHOULD pre-warm new replicas (load weights, connections, caches) before they enter the load-balancing pool.
- SHOULD plan the transition from vertical to horizontal scaling before hitting single-server limits, since vertical scaling creates a single point of failure.
- SHOULD match scaling strategy to each layer's workload characteristics.
Quantitative guidance
As stated by the sources; verify before use.
- Autoscaling cited as delivering 40-90% cost reduction; worked example saved $70,956/year (67%) vs. static peak provisioning (Ch1.8).
- Black Friday example scaled 2 to 30 GPUs in ~90 s, saving $24,000/year vs. static peak provisioning (Ch1.8).
- Example targets: keep CPU below 70% (scale up above 70%, scale down below 50%); scale-up can launch replicas within 30 s (Ch4.7).
- Cold start (weights to GPU, DB connections, cache warm-up, health checks) takes 30 s to several minutes (Ch4.7).
- Diurnal example: 2 replicas overnight scaling to 10 at mid-morning peak (Ch4.7).
- Hysteresis guidance: scale up within 1-2 min of demand increase; wait 10-15 min after demand decrease before scaling down (Ch4.7).
- Priority-based resource allocation with load balancing reported 30% cost reduction and 25% performance improvement; AI-based workload balancing up to 40% infrastructure cost reduction vs static provisioning (Ch4.7, unnamed research).
- Naive scale-out can double infrastructure cost while improving throughput only 30% (Ch4.7).
- Vertical limits ~256 cores and 12 TB RAM per server with non-linear cost growth (Ch6.5).
Classification
- Patterns
- Horizontal autoscalingBudget-bounded max replicasAsymmetric scaling (fast scale-up, slow scale-down / hysteresis)Stabilization windowsMinimum warm replica floorPre-warming replicas before load-balancer admissionPriority-based resource allocationHorizontal scalingHybrid per-layer scaling (vertical GPU embedding, horizontal CPU retrieval, horizontal GPU generation)
- Technologies
- Kubernetes Horizontal Pod AutoscalerKubernetes Horizontal Pod Autoscaler (autoscaling/v2)
- Quality attributes
- Cost efficiencyPerformance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Over-provisioning costLatency degradation under spikesScaling oscillation / thrashingCold-start capacity gapStatic over-provisioning wasteRunaway scaling cost
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.