Model Serving · Software component

Cache-Aware Inference Router

Software componentModel ServingModelsarc:CacheAwareInferenceRouter

A load balancer for inference replicas that selects the target instance using model-serving state such as KV-cache contents, per-instance queue length, accelerator load and loaded LoRA adapters, plus request priority and cost.

Responsibility. Routes each inference request to the instance best able to serve it quickly.

Also known as: Model-aware routing, Inference gateway

Variant of Load Balancer abstract

When to choose. Choose when inference replicas hold non-equivalent state (KV caches, loaded adapters) and multi-agent routing must track which pods cache which contexts.

routes to; monitorsspecializesmonitorsalternative toalternative toInference Server: routes to; monitorsInference ServerLoad Balancer: specializesLoad BalancerKV Cache Manager: monitorsKV Cache ManagerLayer-4 Load Balancer: alternative toLayer-4 Load BalancerSession Affinity Load Balancer: alternative toSession Affinity Load Ba…
Direct neighbourhood (hover for relationship types)

Relationships

routes to dynamic

monitors assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
KV-cache-aware routingQueue-length balancingLoRA-adapter affinityPriority and cost-aware balancing
Technologies
GKE Inference Gateway
Quality attributes
Performance efficiency (ISO/IEC 25010)
Risks mitigated
Cold-cache latency penaltyExpensive adapter swapping

Sources

  1. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.