Model Serving · Data artifact
Balanced Batching Config
Data artifactModel ServingModelsarc:BalancedBatchingConfig
An inference serving configuration with moderate batch-size ranges, millisecond-scale batching timeouts and multiple instances per GPU, allowing opportunistic batching without pathological latency at low traffic.
Responsibility. Balances latency and throughput for mixed workloads.
Also known as: Mixed-workload batching configuration
Variant of Inference Serving Configuration abstract
When to choose. Choose for mixed or semi-interactive workloads, such as customer service chatbots where users tolerate 300-500 ms latency and traffic allows some batching.
Relationships
configures structural
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- Batch size 1-8, 5 ms dynamic batching timeout, 2 inference instances per GPU (Ch4.2).
- 10ms max queue delay, preferred sizes [8,16,32], 2 priority levels balances interactive apps requiring <100ms response (Ch7.1A).
- Balanced/mixed-workload recipe: maximum batch 32, 10 ms maximum queue delay, 2 priority levels with REJECT timeout action (Ref7.02).
Classification
- Patterns
- Dynamic batching
- Technologies
- NVIDIA Triton Inference Server
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Latency degradation during low trafficUnder-utilisation during peaks
Sources
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html