Infrastructure · Software component

Latency-Target Autoscaler

Software componentInfrastructureInfrastructurearc:LatencyTargetAutoscaler

An autoscaler driven by custom service metrics, typically P95 latency approaching the SLO threshold, adding replicas before response-time objectives are breached.

Responsibility. Scales replicas to keep tail latency within the service-level objective.

Also known as: Custom-metric autoscaling, SLO-driven autoscaling

Variant of Autoscaler abstract

When to choose. Choose for latency-sensitive deployments with explicit response-time SLOs.

invokesspecializesalternative toalternative toMetrics Collector: invokesMetrics CollectorAutoscaler: specializesAutoscalerQueue-Depth Autoscaler: alternative toQueue-Depth AutoscalerResource-Utilization Autoscaler: alternative toResource-Utilization Aut…
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

alternative to variability

Design guidance

Classification

Patterns
Custom-metric scalingCost-per-query scaling
Quality attributes
Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Latency SLO violations

Sources

  1. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.