Model Serving · Data store

KV Cache Store

Data storeModel ServingModelsarc:KVCacheStore

An accelerator-memory store of per-sequence attention key and value projections reused during autoregressive generation instead of recomputing earlier tokens.

Responsibility. Holds attention key-value projections for in-flight sequences.

Also known as: KV cache

is written bydeployed onis monitored byis written byis read byis configured byis read byis written byInference Server: is written byInference ServerGPU Node: deployed onGPU NodeGPU System Profiler: is monitored byGPU System ProfilerKV Cache Manager: is written byKV Cache ManagerAttention Kernel: is read byAttention KernelEngine Build Configuration: is configured byEngine Build ConfigurationFused Block-wise Attention Kernel: is read byFused Block-wise Attenti…KV Cache Allocator: is written byKV Cache Allocator
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is configured by structural

is read by dependency

is written by dependency

is monitored by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Multi-query attention (MQA)Grouped-query attention (GQA)
Quality attributes
Performance efficiency (ISO/IEC 25010)

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  2. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  3. Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
  4. Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
  5. Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note