Model Serving · Software component

KV Cache Manager

Software componentModel ServingModelsarc:KVCacheManager

An inference-engine component that retains attention key-value states, including for shared prompt prefixes, so later requests skip recomputing them.

Responsibility. Retains and reuses attention key-value states across requests.

Also known as: Prefix cache, Automatic prefix caching, Layer 3 cache, Prompt Cache, Prefix Cache, Prompt cache, Prompt caching, KV cache optimization, Paged KV cache, Prefix caching, Paged attention, KV cache quantization, Provider automatic prompt caching, Server-side cached context

deployed on; cachesdeployed onis configured bydeployed onis invoked byis monitored bywritesis monitored byinvokesis configured byLLM Inference Service: deployed on; cachesLLM Inference ServiceInference Server: deployed onInference ServerSystem Prompt Template: is configured bySystem Prompt TemplateLLM Generation Backend: deployed onLLM Generation BackendIn-Flight Batch Scheduler: is invoked byIn-Flight Batch SchedulerInference Engine Profiler: is monitored byInference Engine ProfilerKV Cache Store: writesKV Cache StoreCache-Aware Inference Router: is monitored byCache-Aware Inference Ro…KV Cache Allocator: invokesKV Cache AllocatorKV Cache Eviction Policy: is configured byKV Cache Eviction Policy
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is configured by structural

caches dependency

invokes dependency

is invoked by dependency

writes dependency

is monitored by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Automatic prefix cachingPrompt (prefix) cachingPrompt prefix cachingPaged attention (block-based non-contiguous allocation)Low-precision KV storageCross-request prefix sharingPaged attentionINT8 KV cache quantizationBeam-search cache sharingReuse of precomputed keys and values for repeated queries over the same contextPrompt cachingGrouped-query attention sharing key-value projections (Ref5.03)
Technologies
NVIDIA NIMTensorRT-LLMOpenAI automatic prompt caching
Quality attributes
Performance efficiency (ISO/IEC 25010)Cost efficiency
Risks mitigated
GPU memory exhaustion from per-request KV caches in multi-user deploymentsKV cache memory fragmentationRedundant retransmission and re-encoding of static context

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
  3. Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
  4. Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
  5. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  6. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  7. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  8. Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
  9. Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.
  10. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  11. Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
  12. Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
  13. Ref5.03: Jamba Team et al., "Jamba-1.5: Hybrid Transformer-Mamba Models at Scale," arXiv:2408.12570, 2024.