Infrastructure · Software component
Queue-Depth Autoscaler
Software componentInfrastructureInfrastructurearc:QueueDepthAutoscaler
An autoscaler that sizes a replica fleet from the number of requests waiting for processing rather than from host resource utilization.
Responsibility. Scales replicas according to pending-request queue depth.
Also known as: Backlog-based autoscaling, Aggressive metric-based autoscaling
Variant of Autoscaler abstract
When to choose. Choose for agents that primarily orchestrate external API or tool calls, where CPU does not reflect capacity.
Relationships
is configured by structural
scales control
- Automatic Speech Recognizer abstract Ch7.5
monitors assurance
alternative to variability
Design guidance
- SHOULD be preferred over CPU for API-orchestration agents that spend most time idle awaiting external services.
- SHOULD scale up aggressively and scale down conservatively to prevent thrashing, accepting short-term over-provisioning.
Quantitative guidance
As stated by the sources; verify before use.
- Example: target queue depth 5 (vs. 10 conservative), 5-100 replicas, doubling every 15s with 10s stabilization; scale-down 10%/60s after a 300s window (Ch7.5).
Classification
- Patterns
- Queue-depth scaling
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Under-provisioning of I/O-bound agents that show low CPU
Sources
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note