Infrastructure · Software component
Metric-Aware Load Balancer
Software componentInfrastructureInfrastructurearc:MetricAwareLoadBalancer
A load balancer that routes each request to the replica with the best real-time health indicators (CPU, memory pressure, queue depth, recent latency) polled from replica metrics endpoints.
Responsibility. Routes requests by near-real-time replica load and latency metrics.
Also known as: Dynamic load balancing, Least-response-time routing
Variant of Load Balancer abstract
When to choose. Choose only when simpler heuristics (least connections, weighting) fail to maintain SLOs and observability infrastructure can supply fresh metrics.
Relationships
caches dependency
invokes dependency
alternative to variability
Design guidance
- SHOULD poll replica metrics every few seconds with caching rather than per-second collection from every replica.
- MUST define fallback behaviour and metric TTLs for missing or stale metrics to avoid starving healthy replicas or overloading unreachable ones.
Quantitative guidance
As stated by the sources; verify before use.
- Example: prefers a replica at 45% CPU / 500 ms P95 over one at 85% CPU / 2 s P95; metrics 30 s old are of limited value (Ch4.7).
Classification
- Patterns
- Metric-driven routingHierarchical metric polling with cachingMetric TTL and fallback behaviour
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Traffic to transiently degraded replicasCascading failure death spirals
Sources
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.