Model Serving · Software component
Paged KV Cache Allocator
Software componentModel ServingModelsarc:PagedKVCacheAllocator
A KV-cache allocator that assigns fixed-size, possibly non-contiguous pages on demand from a free pool and maps them through a page table.
Responsibility. Allocates KV-cache memory in on-demand fixed-size pages.
Also known as: Paged Attention, Paged KV cache, PagedAttention
Variant of KV Cache Allocator abstract
When to choose. Choose by default for production serving, especially workloads with high variance in generation length; benefit is minimal for fixed-length outputs.
Relationships
deployed on structural
is invoked by dependency
alternative to variability
Design guidance
- SHOULD be used for batch sizes above ~8 and for memory-constrained or variable-length workloads.
Quantitative guidance
As stated by the sources; verify before use.
- Pages typically 16-32 tokens; page-table indirection costs under 1% compute for 2-3x memory efficiency (Ch4.4).
- Enables 2-3x higher batch sizes (batch 4 -> 12-16 typical) (Ch4.4).
- Legal-contract example: batch 6 -> 12, 5.2 -> 6.4 req/s (combined 2.67x), 79.2 GB; eliminated 45% KV waste (Ch4.4).
- Reduces KV-cache waste from ~95% to ~5% and improves throughput 2-4x on dynamic workloads with bit-identical accuracy (Ch7.4, Ref7.05).
- A100 40GB rises from 12 to 32-48 concurrent requests (Ch7.4).
Classification
- Patterns
- Virtual-memory-style pagingCross-request prefix page sharingVirtual-memory paging of KV cacheBlock indirection table
- Technologies
- vLLMTensorRT-LLMNVIDIA NIM
- Risks mitigated
- KV-cache fragmentationPeriodic stalls from cache reallocation
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note