Model Serving · Software component
KV Cache Allocator
Software componentModel ServingModelsVariation point (abstract)arc:KVCacheAllocator
An abstract inference-engine component that allocates and releases accelerator memory for each request's key-value cache.
Responsibility. Allocates KV-cache memory to in-flight requests.
Variants
| Variant | When to choose |
|---|---|
| Contiguous KV Cache Allocator | Choose (with buffers pre-allocated for maximum length) only when sequence lengths are bounded and predictable, accepting higher initial memory use. |
| Paged KV Cache Allocator | Choose by default for production serving, especially workloads with high variance in generation length; benefit is minimal for fixed-length outputs. |
Relationships
is configured by structural
is invoked by dependency
writes dependency
Design guidance
- SHOULD be treated as the primary throughput lever because KV cache, not weights, dominates memory under concurrency.
Quantitative guidance
As stated by the sources; verify before use.
- Static allocation on A100 40GB caps concurrency at 12-14 concurrent 2048-token requests (Ch7.4).
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/