Model Serving · Data artifact
Latency-Oriented Batching Config
Data artifactModel ServingModelsarc:LatencyOrientedBatchingConfig
An inference serving configuration with small maximum batch sizes and short batch timeouts to minimise per-request latency.
Responsibility. Configures batching for interactive response latency.
Also known as: sentiment_realtime configuration
Variant of Inference Serving Configuration abstract
When to choose. Choose for interactive query-time workloads such as visual question answering where sub-second responses are required.
Relationships
configures structural
alternative to variability
Design guidance
- SHOULD avoid queueing delays for interactive applications targeting sub-200ms P95 latency.
- SHOULD reduce prefill batch size for latency-critical (time-to-first-token) workloads.
Quantitative guidance
As stated by the sources; verify before use.
- Max batch size 4-8 with 50 ms batch timeout for sub-second VQA (Ch2.7).
- Min batch 1, max batch 2, dynamic batching timeout < 100 microseconds; ~50 RPS vs ~200 RPS throughput-optimised (Ch4.2).
- Real-time BERT config: max_batch_size 32, preferred [8,16], max delay 2ms, 2 GPU instances; P50 45ms / P95 85ms / P99 120ms at ~40 req/s (Ch4.5).
- Interactive requirement: max 300ms latency (Ch4.5).
- Triton dynamic_batching: max_queue_delay_microseconds 100000, max_batch_size 32, preserve_ordering true (Ch4.7).
- 1ms delay, max batch 8, preferred [1,2,4,8]: P50 18-25ms, P99 40-60ms at 2-3x lower throughput (Ch7.1A).
- Users perceive <200ms as instantaneous, 200-1000ms as responsive, >1000ms as sluggish (Ch7.2).
- Low-latency interactive recipe: maximum batch 8, 1 ms maximum queue delay (Ref7.02).
Classification
- Patterns
- Dynamic batching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note