Model Serving · Data artifact
KV Cache Eviction Policy
Data artifactModel ServingModelsVariation point (abstract)arc:KVCacheEvictionPolicy
An abstract policy deciding which requests' KV caches to evict or defer when cache demand exceeds available memory.
Responsibility. Selects KV-cache victims under memory pressure.
Variants
| Variant | When to choose |
|---|---|
| Cost-Aware Hybrid KV Cache Eviction Policy | Choose in production when recency, priority and recomputation cost must be balanced together. |
| FIFO KV Cache Eviction Policy | Choose when all requests have similar priority. |
| LRU KV Cache Eviction Policy | Choose for interactive sessions where recent activity signals continued engagement and idle sessions can be suspended. |
| Priority KV Cache Eviction Policy | Choose when requests carry differentiated SLAs (e.g., premium versus free tier). |
Relationships
configures structural
Design guidance
- SHOULD adopt sophisticated eviction only when request importance varies widely or memory is tight enough to force frequent eviction.
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.