Model Serving · Data artifact
Engine Build Configuration
Data artifactModel ServingModelsarc:EngineBuildConfig
A build-time configuration selecting engine precision, attention kernel plugins, paged KV-cache use and batch limits when compiling a model into an inference engine.
Responsibility. Fixes the optimizations compiled into an inference engine.
Also known as: Engine build flags, world_size / tp_size / pp_size build parameters, Build flags, Shape profile, Paged KV-cache build parameters
Relationships
configures structural
Design guidance
- MUST set min/opt/max shape profiles to the real workload range; engines tuned for batch 1 perform poorly at batch 32.
- SHOULD default to 64-token KV-cache blocks, balancing internal fragmentation against indirection overhead.
- SHOULD size the total token budget from GPU memory left after weights and activations divided by per-token KV-cache cost.
- SHOULD always enable input-padding removal; it is independent of cache strategy and has no trade-off.
- SHOULD size workspace memory to enable fusion and autotuning without starving request memory (about 8 GB for 7B; 12-16 GB for larger models).
Quantitative guidance
As stated by the sources; verify before use.
- 16-token blocks keep fragmentation <1% but cost 5-10% on attention-bound workloads; 128-token blocks gain 3-5% but waste ~60% for 50-token sequences (Ch7.4).
- 64-token blocks give 20-30% internal fragmentation for median 80-120-token generations (Ch7.4).
- A100 40GB: 24 GB left for KV cache at 0.00048 MB/token (Llama 2 7B) gives ~40,000 64-token blocks (~2.56M tokens); INT8 KV cache doubles this (Ch7.4).
- Fused multi-head attention adds 10-20% speedup; removing input padding adds 10-15% (Ch7.4).
- Typical workspace memory 1-4 GB (Ref7.01).
Classification
- Patterns
- min/opt/max shape profilesPaged KV-cache block size (tokens_per_block)Total token budget (max_num_tokens)Fused multi-head attention flagInput-padding removalTensor parallelism
- Technologies
- TensorRT-LLM GPT Attention Plugin
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html