Model Serving · Software component
Dynamic Batch Scheduler
Software componentModel ServingModelsarc:DynamicBatchScheduler
A batch scheduler that queues requests per model and dispatches a batch when a preferred batch size is reached or a maximum queue delay expires.
Responsibility. Assembles request-level batches under a bounded queue delay for one model.
Also known as: Triton dynamic batcher, Request-level dynamic batching, Dynamic batching, Dynamic batcher
Variant of Inference Batch Scheduler abstract
When to choose. Choose for models on backends that accept batched tensors (TensorFlow, PyTorch, ONNX Runtime, TensorRT); do not use for engines with internal continuous batching.
Relationships
deployed on structural
is configured by structural
invokes dependency
alternative to variability
excludes variability
Design guidance
- SHOULD set preferred batch sizes aligned to accelerator matrix-unit dimensions (e.g., multiples of 32).
- SHOULD use short maximum queue delays (1-2ms) for latency-sensitive paths and longer delays (10-20ms or more) for throughput-oriented batch workloads.
- SHOULD bound queue depth and apply queue timeouts so overload yields rejection and client retry rather than server memory exhaustion.
- SHOULD tune the timeout as the primary latency-throughput operating point (short for responsiveness, long for throughput).
- SHOULD aggregate sporadic embedding requests into batches transparently to keep GPU utilization high.
- SHOULD set max queue delay as the explicit latency-throughput trade-off: 1-5ms for <50ms median latency, 50-100ms for 80-90% GPU utilization.
- SHOULD prefer GPU-optimal preferred batch sizes (e.g., 8, 16, 32).
- SHOULD start with a conservative maximum batch size (32 or 64) and raise it according to GPU memory and latency requirements.
- SHOULD configure a preferred batch size only if it performs significantly better than other batch sizes.
- SHOULD tune the maximum queue delay to trade latency for batch efficiency, raising it when production batches stay small and lowering it when queue delay exceeds the computation benefit.
- MAY use ragged batching to process variable-length inputs without explicit padding.
- SHOULD cap maximum batch size so GPU memory stays below saturation; memory-bound inference degrades latency via host-memory swapping over PCIe.
Quantitative guidance
As stated by the sources; verify before use.
- Typical balance: max_queue_delay 5ms; overnight 5-user traffic still gains 6-8x throughput for +5ms latency (Ch4.5).
- 200 concurrent users: 128-request batches assembled within 2-3ms (Ch4.5).
- Worked config: max_batch_size 32, max_queue_delay 100 ms, preserve_ordering; worst case 100 ms + 150 ms = 250 ms vs 1,000 ms P95 SLA; ~85% GPU utilization at peak, ~40% off-peak (Ch4.7).
- Timeouts of 10 ms favour responsiveness; 500 ms favours throughput (Ch4.7).
- A 7B forward pass on one query takes 15ms using 8-12% of cores; batching 32 reaches 85-90% utilization but adds 480ms to early arrivals (Ch7.1A).
- Enabling dynamic batching yields 3-5x throughput; defaults are 100us max queue delay, sizes 1-32, single-priority FIFO (Ch7.1A).
- Preferred batch sizes give an 8-12% throughput gain (Ch7.1A).
- Multiples of 32 are cited as optimizing tensor-core utilization for FP16/INT8 (Ref7.02).
- Batch-formation probability modelled as P = 1 - exp(-(arrival_rate x max_delay)) (Ref7.02).
- Reducing max batch size 32->16 dropped GPU memory 95%->50% (A100 40 GB, 38 GB used, 45% compute) and restored P95 from 4.5 s to 1.3 s within 15 minutes of a 2-minute change (Ch8.1).
Classification
- Patterns
- Dynamic batchingPriority queuingBackpressureSize-or-timeout batchingDelayed batching (maximum queue delay)Preferred batch sizesRagged batching (no padding of variable-length inputs)
- Technologies
- NVIDIA Triton Inference ServerNVIDIA Triton Inference Server (dynamic_batching)
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Unbounded batch-formation waitGPU memory exhaustion from oversized batchesPadding waste on variable-length inputs
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note