Model Serving · Data artifact
Inference Serving Configuration
Data artifactModel ServingModelsVariation point (abstract)arc:InferenceServingConfig
A configuration artifact fixing the served model, sampling temperature, maximum output tokens, engine optimisation, and batching and streaming modes for an inference endpoint.
Responsibility. Sets model and runtime parameters of an inference endpoint.
Also known as: NIM configuration, nim_config, Per-node LLM configuration, Sampling Configuration, Inference Parameter Settings, Dynamic batching configuration, Instance and concurrency settings, Triton model configuration (config.pbtxt), Per-model serving configuration, Model configuration file (model.pbtxt)
Variants
| Variant | When to choose |
|---|---|
| Balanced Batching Config | Choose for mixed or semi-interactive workloads, such as customer service chatbots where users tolerate 300-500 ms latency and traffic allows some batching. |
| Inference Queue Policy | Choose when serving heterogeneous users with tiered SLAs (premium <50ms, standard <200ms, best-effort batch) on shared infrastructure. |
| Latency-Oriented Batching Config | Choose for interactive query-time workloads such as visual question answering where sub-second responses are required. |
| Latency-Tuned Decoding Configuration | Choose for interactive applications (chat, code completion, real-time moderation) with short, factual responses where deterministic, concise output is acceptable. |
| Model Deployment Profile abstract | — |
| Throughput-Oriented Batching Config | Choose for offline RAG preprocessing (bulk captioning or extraction) where images processed per second matters more than per-image latency. |
Relationships
configures structural
is written by dependency
is constrained by control
is produced by lifecycle
Design guidance
- SHOULD disable batching and streaming for real-time routing endpoints to prioritise latency over throughput.
- SHOULD cap maximum output tokens for routing calls, which need only a category name.
- MAY differ per workflow node, e.g., temperature 0 for analysis nodes and higher temperature for creative generation.
- SHOULD set temperature to 0 for agents that select and call tools, reserving higher temperatures for creative generation.
- SHOULD set batching parameters by workload: throughput-oriented for offline preprocessing, latency-oriented for interactive queries.
- MAY bind model instances to NUMA nodes and CPU cores on multi-socket hosts.
- SHOULD maintain separate temperature settings per agent sub-task rather than one global value.
- SHOULD use near-zero temperature for factual QA and deterministic tool selection.
- MAY combine higher temperature with a top-p bound to keep creativity while preventing degenerate outputs.
- SHOULD use low temperature (0.0-0.3) for production agents requiring deterministic, consistent behaviour.
- MUST explicitly prioritise latency or throughput from business requirements; no configuration minimises latency and maximises throughput simultaneously.
- SHOULD tune batch size, dynamic batching timeout and instance count to reach the chosen point on the latency-throughput Pareto frontier.
- SHOULD define per-model tensor shapes, types, backend, batching parameters, version policy and instance counts explicitly for every served model.
- SHOULD template common serving settings once and specialise per model by parameter substitution to reduce duplication and configuration drift.
- MUST omit the batch dimension from tensor shapes when server-side batching is enabled, since the server adds it automatically.
Quantitative guidance
As stated by the sources; verify before use.
- Model instance count and dynamic batching are the primary Triton optimisation settings (Ref2.01).
- Low temperature 0.0-0.3 for factual/structured tasks; 0.7-1.0 flattens distribution; creative writing 0.7-0.9; tool-calling agents 0.2-0.4 (Ch3.4).
- Recommended sub-task temperatures: tool selection 0.0, reasoning narration 0.3, user-facing responses 0.5-0.8 (Ch3.4).
- Top-p threshold typically 0.9-0.95; e.g., temperature 0.7 with top-p 0.9 for a research-paper generation agent (Ch3.4).
- Financial advisory example: temperature 0.9 yields inconsistent recommendations; 0.2 ensures consistent advice (Ch3.4).
- Higher temperatures (0.7-1.0) increase diversity at the cost of consistency; conservative 0.2-0.3 with explicit prompt constraints yields consistency (Ch3.5).
- Batch 16 takes ~1.5x the time of batch 1 for ~10x throughput; batch 4 may reach ~80% of max throughput with ~40% latency increase (Ch4.2).
- Typical production dynamic batching: max batch 8-32 requests with 5-10 ms timeout windows (Ch4.4).
- A deployment serving 20 models requires managing 20+ configuration files (Ch4.5).
- Production teams often keep 3-4 configurations per model: ultra-low-latency, balanced, high-throughput and experimental (Ch4.5).
Classification
- Patterns
- Latency-over-throughput tuningPer-sub-task temperature settingsNucleus (top-p) samplingTop-k sampling
- Technologies
- NVIDIA NIM
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Inconsistent answers to repeated identical queriesDegenerate or off-topic outputs from high-temperature samplingHallucinated references
Sources
- Ch1.5B: T. Nguyen, "Stateful Orchestration - Worked Examples," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.5B. ISBN: 9798244538229.
- Ch2.2: T. Nguyen, "LangGraph," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.2. ISBN: 9798244538229.
- Ch2.3: T. Nguyen, "LangChain Sequential Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.3. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ref2.01: NVIDIA, "Optimization," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 26, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/optimization.html
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.08: NVIDIA, "NVIDIA Deep Learning Triton Inference Server Documentation," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/