Model Serving · Software component
KV Cache Manager
Software componentModel ServingModelsarc:KVCacheManager
An inference-engine component that retains attention key-value states, including for shared prompt prefixes, so later requests skip recomputing them.
Responsibility. Retains and reuses attention key-value states across requests.
Also known as: Prefix cache, Automatic prefix caching, Layer 3 cache, Prompt Cache, Prefix Cache, Prompt cache, Prompt caching, KV cache optimization, Paged KV cache, Prefix caching, Paged attention, KV cache quantization, Provider automatic prompt caching, Server-side cached context
Relationships
deployed on structural
is configured by structural
caches dependency
invokes dependency
- KV Cache Allocator abstract Ch4.4
is invoked by dependency
writes dependency
is monitored by assurance
Design guidance
- SHOULD be combined with response-level caches as the computation-reduction layer of a multi-layer cache.
- MUST place stable prompt components (system prompt, tool schemas) before variable components (user query, retrieved context) to maximise prefix reuse.
- SHOULD cache stable prompt components (system prompts, tool descriptions, common reference data) and separate them from per-interaction dynamic content.
- SHOULD allocate KV cache in non-contiguous blocks to reduce fragmentation and raise feasible batch sizes when long contexts cause memory pressure.
- SHOULD enable prefix caching when requests share system prompts or few-shot demonstrations; it adds matching overhead without benefit for unique prompts.
- SHOULD allocate KV cache in pages rather than contiguous arrays and quantize cached values to raise concurrent request capacity.
- MUST NOT treat prompt caching as reducing token budget or cognitive load; it saves computation and latency only.
- SHOULD be applied to static context repeated across many requests (system prompts, tool definitions, stable profiles); savings scale with conversation depth.
- SHOULD be combined with request-level tracking of cached tokens so the realized hit rate and savings are verified.
Quantitative guidance
As stated by the sources; verify before use.
- Prefix caching reduces token processing by 30-50% for workloads with substantial prefix overlap (Ch1.8).
- Prompt caching can reduce cost of cached content by up to 90%; 80% cacheable of 10K input cuts per-query cost $0.36->$0.09 (75%) (Ch3.4).
- Caching 8K tokens of system prompts/tool descriptions reduces base costs 60-70% (Ch3.4).
- Prompt cache hits reduce API costs by 15-30% while improving response speed (Ch3.10).
- KV cache grows linearly with sequence length; 70B model, batch 2, 2,048-token context: 14.2 GB (36% of 40 GB) (Ch4.2 case study).
- After paged KV cache + FP8 weights: batch 8 KV cache 7.8 GB (20%) (Ch4.2).
- Shared 2,000-token system prompt across 100 sessions: 250 GB without prefix caching vs ~2.5 GB plus per-session unique tokens with it (Ch4.4).
- FP16 KV cache consumes 2-4GB per request for a 7B model at 2048 tokens (Ch4.6).
- INT8 KV cache halves memory; paging improves memory utilisation ~30%; combined 4-6x more concurrent requests per GPU (Ch4.6).
- Cached input tokens billed at ~50% of standard price ($1.25 vs $2.50 per 1M); the chapter also cites up to 90% discounts on cached reads (Ch8.3).
- Deployment reached an 82% cache hit rate (an 18-turn conversation yields 17 hits, 94% within conversations) versus a 40% estimate; monthly cost fell $7,500 -> $5,325 (29%, $2,175) (Ch8.3).
- TikTok testing-agent case: 12,000 static tokens per request (3,500 instructions, 4,200 for 47 tool definitions, 2,800 few-shot, 1,500 retrieved docs) vs 180-token queries and 420-token responses; input consumed 78% of cost despite 4x output pricing (87.9% at GPT-4o rates) (Ch8.3).
Classification
- Patterns
- Automatic prefix cachingPrompt (prefix) cachingPrompt prefix cachingPaged attention (block-based non-contiguous allocation)Low-precision KV storageCross-request prefix sharingPaged attentionINT8 KV cache quantizationBeam-search cache sharingReuse of precomputed keys and values for repeated queries over the same contextPrompt cachingGrouped-query attention sharing key-value projections (Ref5.03)
- Technologies
- NVIDIA NIMTensorRT-LLMOpenAI automatic prompt caching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Cost efficiency
- Risks mitigated
- GPU memory exhaustion from per-request KV caches in multi-user deploymentsKV cache memory fragmentationRedundant retransmission and re-encoding of static context
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch5.9: T. Nguyen, "Working Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.9. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref5.03: Jamba Team et al., "Jamba-1.5: Hybrid Transformer-Mamba Models at Scale," arXiv:2408.12570, 2024.