Model Serving · Software component
Sequence Batch Scheduler
Software componentModel ServingModelsarc:SequenceBatchScheduler
A batch scheduler for stateful models that binds each request sequence, identified by a correlation ID, to a slot on one model instance while dynamically batching across concurrent sequences.
Responsibility. Routes every request of a correlated sequence to the model instance slot holding that sequence's state.
Also known as: Sequence batcher, Stateful model batcher
Variant of Inference Batch Scheduler abstract
When to choose. Choose when the served model is stateful (recurrent networks, language models with hidden state, streaming speech recognition, conversational models carrying context) so every request of a sequence must reach the same model instance.
Relationships
deployed on structural
is configured by structural
- Inference Serving Configuration abstract Ref7.02
routes to dynamic
alternative to variability
Design guidance
- MUST tie every request of a stateful sequence together with a correlation ID and signal sequence start and end so completed sequences free their slot.
- SHOULD size the number of concurrent sequence slots to what each model instance can maintain.
Quantitative guidance
As stated by the sources; verify before use.
- Example configuration: 4 sequence slots with a 10 ms maximum queue delay (Ref7.02).
- Rated high throughput, medium latency and high complexity, versus highest throughput and low complexity for the dynamic batcher (Ref7.02 strategy table).
Classification
- Patterns
- Sequence batchingCorrelation-ID request affinitySequence start/end control signals
- Technologies
- NVIDIA Triton Inference Server
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Loss of per-sequence model state when a sequence's requests are spread across instances
Sources
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note