Model Serving · Software component
Cache-Aware Inference Router
Software componentModel ServingModelsarc:CacheAwareInferenceRouter
A load balancer for inference replicas that selects the target instance using model-serving state such as KV-cache contents, per-instance queue length, accelerator load and loaded LoRA adapters, plus request priority and cost.
Responsibility. Routes each inference request to the instance best able to serve it quickly.
Also known as: Model-aware routing, Inference gateway
Variant of Load Balancer abstract
When to choose. Choose when inference replicas hold non-equivalent state (KV caches, loaded adapters) and multi-agent routing must track which pods cache which contexts.
Relationships
routes to dynamic
monitors assurance
alternative to variability
Design guidance
- SHOULD NOT treat stateful inference replicas as equivalent; route to exploit cache and adapter locality.
Quantitative guidance
As stated by the sources; verify before use.
- Warm-cache pod ~100 ms vs cold pod 2+ s (~20x) when context must be rebuilt (Ch4.3).
Classification
- Patterns
- KV-cache-aware routingQueue-length balancingLoRA-adapter affinityPriority and cost-aware balancing
- Technologies
- GKE Inference Gateway
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Cold-cache latency penaltyExpensive adapter swapping
Sources
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.