Infrastructure · Software component
Latency-Target Autoscaler
Software componentInfrastructureInfrastructurearc:LatencyTargetAutoscaler
An autoscaler driven by custom service metrics, typically P95 latency approaching the SLO threshold, adding replicas before response-time objectives are breached.
Responsibility. Scales replicas to keep tail latency within the service-level objective.
Also known as: Custom-metric autoscaling, SLO-driven autoscaling
Variant of Autoscaler abstract
When to choose. Choose for latency-sensitive deployments with explicit response-time SLOs.
Relationships
invokes dependency
alternative to variability
Design guidance
- MAY scale on cost per query instead of latency when budget constraints dominate performance goals.
- Custom metrics require extra instrumentation; SHOULD be adopted when default resource metrics do not track the binding constraint.
Classification
- Patterns
- Custom-metric scalingCost-per-query scaling
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Latency SLO violations
Sources
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.