Model Serving · Data store
KV Cache Store
Data storeModel ServingModelsarc:KVCacheStore
An accelerator-memory store of per-sequence attention key and value projections reused during autoregressive generation instead of recomputing earlier tokens.
Responsibility. Holds attention key-value projections for in-flight sequences.
Also known as: KV cache
Relationships
deployed on structural
is configured by structural
is read by dependency
is written by dependency
is monitored by assurance
Design guidance
- MAY be quantized to FP8 or INT8 to halve per-token footprint and double concurrent capacity.
Quantitative guidance
As stated by the sources; verify before use.
- Often consumes 40-60% of GPU memory and bounds maximum batch size (Ch4.4).
- Llama-3.1-70B, 8,192 tokens, FP16: ~3.2 GB per sequence; batch 8 needs 25.6 GB (Ch4.4).
- Reduces per-token generation cost from O(n) to O(1) in sequence length (Ch4.4).
- Scales linearly with sequence length and batch size; ~2 GB per 2048-token request for Llama 2 7B (Ch7.4, Ref7.05).
- MQA reduces KV memory to 1/H of multi-head attention (Ref7.05).
Classification
- Patterns
- Multi-query attention (MQA)Grouped-query attention (GQA)
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note