Model Serving · Software component
Static Batch Scheduler
Software componentModel ServingModelsarc:StaticBatchScheduler
A batch scheduler that waits until a fixed number of requests accumulate before executing, maximizing GPU saturation at the cost of unbounded wait for early arrivals.
Responsibility. Executes only full, fixed-size batches.
Also known as: Static batching
Variant of Inference Batch Scheduler abstract
When to choose. Choose for offline batch workloads (overnight reports, bulk document processing, scheduled evaluation) where no user waits.
Relationships
alternative to variability
Design guidance
- SHOULD NOT be used for interactive agents because low traffic makes first-arriving requests wait seconds or minutes.
Quantitative guidance
As stated by the sources; verify before use.
- Batch 32 at 50 req/s: 640 ms average formation wait + 150 ms compute = 790 ms, with unbounded waits during traffic transitions (Ch4.7).
Classification
- Patterns
- Fixed-size batching
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/