Model Serving · Software component
Contiguous KV Cache Allocator
Software componentModel ServingModelsarc:ContiguousKVCacheAllocator
A KV-cache allocator that reserves one contiguous buffer per request sized for the maximum sequence length at arrival.
Responsibility. Reserves maximum-length contiguous KV-cache buffers per request.
Also known as: Pre-allocated KV buffers, Standard contiguous KV cache, Continuous KV cache, Static KV-cache allocation
Variant of KV Cache Allocator abstract
When to choose. Choose (with buffers pre-allocated for maximum length) only when sequence lengths are bounded and predictable, accepting higher initial memory use.
Relationships
alternative to variability
Design guidance
- MAY be used for latency-critical single-user serving with batch 1-4 and predictable sequence lengths, or when GPU memory is abundant.
- SHOULD NOT be used with 16+ concurrent variable-length sequences, where fragmentation returns to 80-90%.
Quantitative guidance
As stated by the sources; verify before use.
- Fragmentation wastes 40-60% of KV-cache memory; a 400-token generation in a 2,048-token slot wastes 80% (Ch4.4).
- Saves 2-5% per-token latency by avoiding indirection; crossover vs paging around batch size 8 (Ch7.4).
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.