Model Serving · Software component
In-Flight Batch Scheduler
Software componentModel ServingModelsarc:ContinuousBatchScheduler
A batch scheduler inside an LLM generation engine that admits new sequences and retires finished ones at each generation iteration instead of waiting for a full batch.
Responsibility. Continuously re-forms the active generation batch as sequences arrive and complete.
Also known as: Continuous batching, In-flight batching, Inflight batching, Iteration-level scheduling, In-Flight Batch Scheduler, In-flight request substitution, Sequence batching, Iterative sequence batching
Variant of Inference Batch Scheduler abstract
When to choose. Choose for autoregressive LLM generation where sequences finish at different times and latency and throughput must be balanced without artificial batch-assembly delays.
Relationships
deployed on structural
invokes dependency
alternative to variability
excludes variability
Design guidance
- SHOULD use in-flight batching for high-throughput inference with variable request lengths.
- MUST track per-request generation state, variable sequence lengths and dynamic memory allocation, which justifies its complexity mainly for high-throughput production.
- MUST run on a backend that supports stateful per-iteration sequence control and KV-cache state; plain compiled-plan backends lack continuous batching.
- SHOULD be used for LLM token generation so slots freed by completed requests are refilled by waiting requests without synchronizing the batch.
Quantitative guidance
As stated by the sources; verify before use.
- Orca-style continuous batching improves throughput 2-3x over dynamic batching for long-sequence generation (Ch4.7).
- Utilization rises from 37.5% (static) to 90-95% for variable-length workloads; 5-10x throughput reported with length variation >0.8 (Ch7.1A).
- Batch-4 continuous batching matches batch-16-32 dynamic batching throughput with lower median latency (Ch7.1A).
- 5-10x throughput improvement on real workloads (Ref7.05).
- In a 70B worked example, adding in-flight batching raised throughput from 25 to 100 tokens/s (4x) while per-token latency rose from 40 to 50 ms (Ref7.15).
Classification
- Patterns
- Iteration-level scheduling
- Technologies
- vLLMTensorRT-LLMNVIDIA NIMNVIDIA TensorRT-LLM
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Idle batch slots from early-finishing sequences
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note