Model Serving · Software component

KV Cache Allocator

Software componentModel ServingModelsVariation point (abstract)arc:KVCacheAllocator

An abstract inference-engine component that allocates and releases accelerator memory for each request's key-value cache.

Responsibility. Allocates KV-cache memory to in-flight requests.

is invoked bywritesis configured byis specialized byis specialized byKV Cache Manager: is invoked byKV Cache ManagerKV Cache Store: writesKV Cache StoreEngine Build Configuration: is configured byEngine Build ConfigurationPaged KV Cache Allocator: is specialized byPaged KV Cache AllocatorContiguous KV Cache Allocator: is specialized byContiguous KV Cache Allo…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Contiguous KV Cache AllocatorChoose (with buffers pre-allocated for maximum length) only when sequence lengths are bounded and predictable, accepting higher initial memory use.
Paged KV Cache AllocatorChoose by default for production serving, especially workloads with high variance in generation length; benefit is minimal for fixed-length outputs.

Relationships

is configured by structural

is invoked by dependency

writes dependency

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Sources

  1. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  2. Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
  3. Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/