Model Serving · Software component
Adaptive Batch Size Controller
Software componentModel ServingModelsarc:AdaptiveBatchSizeController
A control component that adjusts inference batch size at runtime from observed queue depth and latency, enlarging batches during surges and shrinking them to hold latency objectives.
Responsibility. Adapts batching to variable demand while meeting latency SLOs.
Also known as: Adaptive batching algorithm, Memory-aware batch size manager
Relationships
writes dependency
- Inference Serving Configuration abstract Ch4.2 Ch7.1A
is constrained by control
monitors assurance
- GPU Node abstract Ch7.1A
- Inference Server Ch4.2
Design guidance
- SHOULD reduce batch size when observed tail latency exceeds target and restore it once latency recovers.
Quantitative guidance
As stated by the sources; verify before use.
- Example: halve batch size when p95 latency exceeds 200 ms (Ch4.2).
- Speculation that forces batch 32 to 12 yields 2.5x x 0.375 = 0.94x net degradation (Ch7.1A).
Classification
- Patterns
- Feedback controlSLO-driven batching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Flexibility (ISO/IEC 25010)
- Risks mitigated
- SLA violations under unexpected traffic with static batchingQueue growth during surges
Sources
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.