Infrastructure · Software component
Layer-7 Load Balancer
Software componentInfrastructureInfrastructurearc:L7LoadBalancer
An application-layer load balancer that parses HTTP requests and routes them by host, path, headers, cookies or body content, typically terminating TLS.
Responsibility. Routes HTTP requests to replica pools using application-level routing rules.
Also known as: Ingress controller, Application load balancer, API gateway reverse proxy, Service mesh traffic splitter, Canary traffic router, Ingress, Kubernetes Ingress
Variant of Load Balancer abstract
When to choose. Choose when request-based routing, application-aware health checking, TLS termination, or cookie-based session affinity is required.
Relationships
deployed on structural
exposes structural
is configured by structural
invokes dependency
is invoked by dependency
routes to dynamic
sends data to dynamic
alternative to variability
Design guidance
- SHOULD provide external access with path-based routing, authentication and rate limiting through a single endpoint.
- SHOULD perform application-aware health checks beyond TCP connectivity.
- MUST keep gateway logic simple (authenticate, rate limit, route, minimal format transformation, aggregate); business rules and state belong in backend services ('smart pipes' anti-pattern).
- SHOULD avoid making the gateway a single point of failure that knows too much about many services.
- SHOULD provide aggregation endpoints returning complete datasets for a use case in one request instead of chatty fine-grained APIs.
- MAY route premium-tier requests to a dedicated reserved-capacity pool and free-tier traffic to an autoscaling pool that scales down when idle.
- SHOULD split canary traffic by weighted random sampling of requests between stable and canary backends.
- SHOULD terminate TLS, apply rate limiting and set request timeouts matched to agent processing latency.
Quantitative guidance
As stated by the sources; verify before use.
- Adds several milliseconds of latency; throughput limited to tens or hundreds of thousands of requests/s on equivalent hardware (Ch1.8).
- A dashboard needing 23 separate API calls pays 20-100 ms per round-trip under real network conditions (Ch4.2).
- Example: /api/v1/chat routed to the conversational agent and /api/v1/analyze to the analytical agent through one entry point (Ch4.7).
Classification
- Patterns
- Path-based routingTLS terminationCookie-based session affinityLeast response timeRequest aggregation (API composition)Tiered routing to dedicated capacity poolsWeighted traffic splittingAsynchronous request acknowledgementHostname-based routing
- Technologies
- Kubernetes Ingressnginx-ingressTraefikHAProxyIstioLinkerdAmazon API GatewayCloud-native load balancersNGINX
- Quality attributes
- Maintainability (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Routing requests to unsuitable instancesChatty APIs requiring many client round-tripsClients coupled to backend topology
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.