Model Serving · Data artifact

KV Cache Eviction Policy

Data artifactModel ServingModelsVariation point (abstract)arc:KVCacheEvictionPolicy

An abstract policy deciding which requests' KV caches to evict or defer when cache demand exceeds available memory.

Responsibility. Selects KV-cache victims under memory pressure.

configuresis specialized byis specialized byis specialized byis specialized byKV Cache Manager: configuresKV Cache ManagerCost-Aware Hybrid KV Cache Eviction Policy: is specialized byCost-Aware Hybrid KV Cac…FIFO KV Cache Eviction Policy: is specialized byFIFO KV Cache Eviction P…LRU KV Cache Eviction Policy: is specialized byLRU KV Cache Eviction Po…Priority KV Cache Eviction Policy: is specialized byPriority KV Cache Evicti…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Cost-Aware Hybrid KV Cache Eviction PolicyChoose in production when recency, priority and recomputation cost must be balanced together.
FIFO KV Cache Eviction PolicyChoose when all requests have similar priority.
LRU KV Cache Eviction PolicyChoose for interactive sessions where recent activity signals continued engagement and idle sessions can be suspended.
Priority KV Cache Eviction PolicyChoose when requests carry differentiated SLAs (e.g., premium versus free tier).

Relationships

configures structural

Design guidance

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.