Infrastructure · Software component

Load Balancer

Software componentInfrastructureInfrastructureVariation point (abstract)arc:LoadBalancer

A distribution component that spreads requests across multiple instances of an agent or service using health checks.

Responsibility. Distributes requests across healthy replicas.

Also known as: Service load balancer, Ingress load balancer, Traffic distribution layer

monitors; invokesis routed to by; is specialized byroutes toroutes toroutes toroutes todeployed onroutes tois overridden byis specialized byreceives data fromis specialized byroutes tois specialized byis specialized byis specialized byis specialized byis specialized byReadiness Endpoint: monitors; invokesReadiness EndpointDNS Load Balancer: is routed to by; is specialized byDNS Load BalancerAgent Controller: routes toAgent ControllerLLM Inference Service: routes toLLM Inference ServiceInference Server: routes toInference ServerWorker Agent: routes toWorker AgentContainer Orchestrator: deployed onContainer OrchestratorRAG Query Orchestrator: routes toRAG Query OrchestratorPlatform Operator: is overridden byPlatform OperatorLayer-7 Load Balancer: is specialized byLayer-7 Load BalancerContainer Health Prober: receives data fromContainer Health ProberLayer-4 Load Balancer: is specialized byLayer-4 Load BalancerVector Store Query API: routes toVector Store Query APISession Affinity Load Balancer: is specialized bySession Affinity Load Ba…Metric-Aware Load Balancer: is specialized byMetric-Aware Load BalancerCache-Aware Inference Router: is specialized byCache-Aware Inference Ro…Least-Connections Load Balancer: is specialized byLeast-Connections Load B…Weighted Round-Robin Load Balancer: is specialized byWeighted Round-Robin Loa…+3 more (see relationships)
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Cache-Aware Inference RouterChoose when inference replicas hold non-equivalent state (KV caches, loaded adapters) and multi-agent routing must track which pods cache which contexts.
DNS Load Balancer—
Layer-4 Load BalancerChoose when all instances are identical and all requests equivalent, so raw throughput and simplicity outweigh request-aware routing.
Layer-7 Load BalancerChoose when request-based routing, application-aware health checking, TLS termination, or cookie-based session affinity is required.
Least-Connections Load BalancerChoose as the default when request durations vary widely (e.g., 90% under 2 s, 10% over 10 s).
Metric-Aware Load BalancerChoose only when simpler heuristics (least connections, weighting) fail to maintain SLOs and observability infrastructure can supply fresh metrics.
Round-Robin Load BalancerChoose only when all requests have similar processing cost and all replicas identical capacity.
Session Affinity Load BalancerChoose as the simplest cache-locality option for conversational agents whose follow-up messages come from the same client.
Weighted Round-Robin Load BalancerChoose when replicas run on heterogeneous infrastructure with measurable capacity differences, or for weighted gradual rollouts.

Relationships

deployed on structural

is configured by structural

invokes dependency

is routed to by dynamic

receives data from dynamic

routes to dynamic

is overridden by control

monitors assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Homogeneous agent poolHealth-check failoverRound-robinLeast connectionsLeast response timeWeighted round-robinConsistent hashingCanary traffic shiftingMulti-region failoverHealth-check-based rotation removalActive-passive HA pairActive-active HA pairHealth-check-coupled routing
Technologies
Kubernetes ServiceKubernetes IngressIstioCloud provider load balancerAWS Elastic Load Balancer / Application Load BalancerGoogle Cloud Load BalancerAzure Load Balancernginx
Quality attributes
Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Single point of failureUneven load distributionRouting traffic to failed instancesUneven load concentration on one replicaLoad balancer as single point of failureCascading overload and retry death spirals

Sources

  1. Ch1.3: T. Nguyen, "Multi-Agent Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.3. ISBN: 9798244538229.
  2. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  3. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
  4. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  5. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  6. Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
  7. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
  8. Ref4.08: NVIDIA, "Agentic AI in the Factory," NVIDIA Enterprise AI Factory Design Guide White Paper. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/ai-enterprise/planning-resource/ai-factory-white-paper/latest/agentic-ai-in-the-factory.html
  9. Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
  10. Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note