Infrastructure · Software component
Load Balancer
Software componentInfrastructureInfrastructureVariation point (abstract)arc:LoadBalancer
A distribution component that spreads requests across multiple instances of an agent or service using health checks.
Responsibility. Distributes requests across healthy replicas.
Also known as: Service load balancer, Ingress load balancer, Traffic distribution layer
Variants
| Variant | When to choose |
|---|---|
| Cache-Aware Inference Router | Choose when inference replicas hold non-equivalent state (KV caches, loaded adapters) and multi-agent routing must track which pods cache which contexts. |
| DNS Load Balancer | — |
| Layer-4 Load Balancer | Choose when all instances are identical and all requests equivalent, so raw throughput and simplicity outweigh request-aware routing. |
| Layer-7 Load Balancer | Choose when request-based routing, application-aware health checking, TLS termination, or cookie-based session affinity is required. |
| Least-Connections Load Balancer | Choose as the default when request durations vary widely (e.g., 90% under 2 s, 10% over 10 s). |
| Metric-Aware Load Balancer | Choose only when simpler heuristics (least connections, weighting) fail to maintain SLOs and observability infrastructure can supply fresh metrics. |
| Round-Robin Load Balancer | Choose only when all requests have similar processing cost and all replicas identical capacity. |
| Session Affinity Load Balancer | Choose as the simplest cache-locality option for conversational agents whose follow-up messages come from the same client. |
| Weighted Round-Robin Load Balancer | Choose when replicas run on heterogeneous infrastructure with measurable capacity differences, or for weighted gradual rollouts. |
Relationships
deployed on structural
is configured by structural
invokes dependency
is routed to by dynamic
receives data from dynamic
routes to dynamic
is overridden by control
monitors assurance
Design guidance
- SHOULD use health checks to detect failed instances and fail over automatically, across availability zones.
- MUST remove instances that fail health checks from rotation so traffic flows only to instances able to serve.
- SHOULD NOT rely on plain round-robin when request complexity varies widely; prefer load-aware algorithms.
- SHOULD dampen least-response-time routing (windowed averages, limited shift rates) to avoid positive-feedback oscillation.
- MAY use consistent hashing on user/session/query keys when replicas keep local caches, accepting reduced load balance.
- MAY use weighted round-robin for heterogeneous GPU fleets or canary rollouts (e.g., 90:10 shifting to 0:100).
- MUST itself be redundant (active-passive, active-active, or DNS-based distribution) so it is not a single point of failure.
- SHOULD start with least-connections routing for variable-duration agent workloads, add weighting for heterogeneous hardware, and adopt dynamic metrics only when simple heuristics fail SLOs.
- SHOULD NOT use IP-hash/session affinity unless affinity yields concrete caching or connection-reuse benefits.
- SHOULD front clustered vector store nodes, health-check each node's readiness endpoint and remove failed nodes from rotation.
- SHOULD use structured health status (healthy / degraded) to reduce traffic gradually instead of removing replicas abruptly.
- SHOULD distribute requests across identical inference or agent replicas on multiple nodes for scaling and redundancy.
Quantitative guidance
As stated by the sources; verify before use.
- Aggressive health checking (5 s interval, 2 failures) detects failure in ~10 s; conservative (30 s, 3 failures) takes ~90 s (Ch1.8).
- Multi-region failover cited as enabling 99.99% uptime (Ch1.8).
- Horizontal scaling across identical nodes is cited as near-linear (3x throughput with 3 nodes) (Ref7.17).
Classification
- Patterns
- Homogeneous agent poolHealth-check failoverRound-robinLeast connectionsLeast response timeWeighted round-robinConsistent hashingCanary traffic shiftingMulti-region failoverHealth-check-based rotation removalActive-passive HA pairActive-active HA pairHealth-check-coupled routing
- Technologies
- Kubernetes ServiceKubernetes IngressIstioCloud provider load balancerAWS Elastic Load Balancer / Application Load BalancerGoogle Cloud Load BalancerAzure Load Balancernginx
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Single point of failureUneven load distributionRouting traffic to failed instancesUneven load concentration on one replicaLoad balancer as single point of failureCascading overload and retry death spirals
Sources
- Ch1.3: T. Nguyen, "Multi-Agent Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.3. ISBN: 9798244538229.
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ref4.08: NVIDIA, "Agentic AI in the Factory," NVIDIA Enterprise AI Factory Design Guide White Paper. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/ai-enterprise/planning-resource/ai-factory-white-paper/latest/agentic-ai-in-the-factory.html
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note