Infrastructure · Software component

Autoscaler

Software componentInfrastructureInfrastructureVariation point (abstract)arc:Autoscaler

A control component that adjusts the number of replicas of a workload to match demand, turning fixed infrastructure cost into variable cost.

Responsibility. Adjusts replica counts of workloads to match demand.

Also known as: Pod autoscaler, Horizontal Pod Autoscaler (HPA), Elastic scaling controller, Autoscaling group

scalesscalesscalesis triggered byemits telemetry todeployed onscalesis monitored byscalesis specialized byis specialized byis specialized byis specialized byis configured byis specialized byAgent Controller: scalesAgent ControllerLLM Inference Service: scalesLLM Inference ServiceWorker Agent: scalesWorker AgentAlert Manager: is triggered byAlert ManagerMetrics Collector: emits telemetry toMetrics CollectorContainer Orchestrator: deployed onContainer OrchestratorRAG Query Orchestrator: scalesRAG Query OrchestratorPlatform Operator: is monitored byPlatform OperatorDialogue Flow Manager: scalesDialogue Flow ManagerMetric-Driven Autoscaler: is specialized byMetric-Driven AutoscalerQueue-Depth Autoscaler: is specialized byQueue-Depth AutoscalerScheduled Scaler: is specialized byScheduled ScalerResource-Utilization Autoscaler: is specialized byResource-Utilization Aut…Autoscaling Policy: is configured byAutoscaling PolicyLatency-Target Autoscaler: is specialized byLatency-Target Autoscaler
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Latency-Target AutoscalerChoose for latency-sensitive deployments with explicit response-time SLOs.
Metric-Driven AutoscalerChoose for stateless agents with unpredictable, high-variance traffic (e.g., multi-tenant SaaS) that tolerate 60-90 s scale-up lag.
Queue-Depth AutoscalerChoose for agents that primarily orchestrate external API or tool calls, where CPU does not reflect capacity.
Resource-Utilization AutoscalerChoose CPU for compute-bound agents dominated by inference or reasoning loops, and memory for agents holding large context windows or document embeddings.
Scheduled ScalerChoose when traffic patterns are extremely predictable, making schedule-based capacity simpler and equally effective.

Relationships

deployed on structural

is configured by structural

emits telemetry to dynamic

is triggered by dynamic

scales control

is monitored by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Horizontal autoscalingBudget-bounded max replicasAsymmetric scaling (fast scale-up, slow scale-down / hysteresis)Stabilization windowsMinimum warm replica floorPre-warming replicas before load-balancer admissionPriority-based resource allocationHorizontal scalingHybrid per-layer scaling (vertical GPU embedding, horizontal CPU retrieval, horizontal GPU generation)
Technologies
Kubernetes Horizontal Pod AutoscalerKubernetes Horizontal Pod Autoscaler (autoscaling/v2)
Quality attributes
Cost efficiencyPerformance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Over-provisioning costLatency degradation under spikesScaling oscillation / thrashingCold-start capacity gapStatic over-provisioning wasteRunaway scaling cost

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  3. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
  4. Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.