Model Serving · Data artifact
Inference Queue Policy
Data artifactModel ServingModelsarc:PriorityTieredBatchingConfig
A per-model configuration of request priority levels, queue timeout actions and maximum queue depth that provides admission control and backpressure for an inference server.
Responsibility. Bounds and prioritises the inference request queue.
Also known as: default_queue_policy, Queue properties, Mixed workload batching configuration, Queue policy, Priority queue policy, Inference Queue Policy
Variant of Inference Serving Configuration abstract
When to choose. Choose when serving heterogeneous users with tiered SLAs (premium <50ms, standard <200ms, best-effort batch) on shared infrastructure.
Relationships
configures structural
alternative to variability
Design guidance
- SHOULD use priority levels to give high-value interactive requests preferential batching over bulk analytics.
- MUST set a maximum queue size and a timeout action (e.g., reject) for batch workloads.
- SHOULD reject requests exceeding a queue timeout and bound queue size to prevent unbounded latency and memory exhaustion during spikes.
- SHOULD define priority levels so high-priority (SLA-bound) requests bypass lower-priority batch traffic in mixed workloads.
- SHOULD set per-priority queue timeouts with an explicit timeout action (reject, defer, or allow delayed scheduling).
Quantitative guidance
As stated by the sources; verify before use.
- Batch example: max_queue_size 5000, default timeout 10s with REJECT action (Ch4.5).
- Example policy: REJECT after 30ms queue time, max_queue_size 128 (Ch7.1A).
- Under high load premium keeps P99 45-60ms while standard degrades to P99 180-250ms and best-effort tolerates seconds (Ch7.1A).
- Example: 2 priority levels; default policy 30 ms timeout with REJECT; per-priority policy 1 ms timeout with DEFER (Ref7.02).
Classification
- Patterns
- Priority schedulingTiered SLAs
- Technologies
- NVIDIA Triton Inference Server
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Fairness (NIST AI RMF: fair, harmful bias managed)
- Risks mitigated
- Memory exhaustion from unbounded request accumulationInfinite queuing when traffic exceeds capacityHigh-value requests starved by bulk analytics
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note