Infrastructure · Software component

Metric-Aware Load Balancer

Software componentInfrastructureInfrastructurearc:MetricAwareLoadBalancer

A load balancer that routes each request to the replica with the best real-time health indicators (CPU, memory pressure, queue depth, recent latency) polled from replica metrics endpoints.

Responsibility. Routes requests by near-real-time replica load and latency metrics.

Also known as: Dynamic load balancing, Least-response-time routing

Variant of Load Balancer abstract

When to choose. Choose only when simpler heuristics (least connections, weighting) fail to maintain SLOs and observability infrastructure can supply fresh metrics.

invokes; cachesspecializesalternative tois target of alternativeToalternative toalternative toReplica Metrics Endpoint: invokes; cachesReplica Metrics EndpointLoad Balancer: specializesLoad BalancerSession Affinity Load Balancer: alternative toSession Affinity Load Ba…Least-Connections Load Balancer: is target of alternativeToLeast-Connections Load B…Weighted Round-Robin Load Balancer: alternative toWeighted Round-Robin Loa…Round-Robin Load Balancer: alternative toRound-Robin Load Balancer
Direct neighbourhood (hover for relationship types)

Relationships

caches dependency

invokes dependency

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Metric-driven routingHierarchical metric polling with cachingMetric TTL and fallback behaviour
Quality attributes
Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Traffic to transiently degraded replicasCascading failure death spirals

Sources

  1. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.