Model Serving · Software component

In-Flight Batch Scheduler

Software componentModel ServingModelsarc:ContinuousBatchScheduler

A batch scheduler inside an LLM generation engine that admits new sequences and retires finished ones at each generation iteration instead of waiting for a full batch.

Responsibility. Continuously re-forms the active generation batch as sequences arrive and complete.

Also known as: Continuous batching, In-flight batching, Inflight batching, Iteration-level scheduling, In-Flight Batch Scheduler, In-flight request substitution, Sequence batching, Iterative sequence batching

Variant of Inference Batch Scheduler abstract

When to choose. Choose for autoregressive LLM generation where sequences finish at different times and latency and throughput must be balanced without artificial batch-assembly delays.

alternative to; excludesdeployed oninvokesdeployed onspecializesexcludesalternative toinvokesalternative toDynamic Batch Scheduler: alternative to; excludesDynamic Batch SchedulerLLM Inference Service: deployed onLLM Inference ServiceKV Cache Manager: invokesKV Cache ManagerLLM Generation Backend: deployed onLLM Generation BackendInference Batch Scheduler: specializesInference Batch SchedulerTensor Framework Backend: excludesTensor Framework BackendSequence Batch Scheduler: alternative toSequence Batch SchedulerPaged KV Cache Allocator: invokesPaged KV Cache AllocatorStatic Batch Scheduler: alternative toStatic Batch Scheduler
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

invokes dependency

alternative to variability

excludes variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Iteration-level scheduling
Technologies
vLLMTensorRT-LLMNVIDIA NIMNVIDIA TensorRT-LLM
Quality attributes
Performance efficiency (ISO/IEC 25010)
Risks mitigated
Idle batch slots from early-finishing sequences

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
  3. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  4. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  5. Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
  6. Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
  7. Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
  8. Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
  9. Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  10. Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note