Model Serving · Data artifact

Inference Serving Configuration

Data artifactModel ServingModelsVariation point (abstract)arc:InferenceServingConfig

A configuration artifact fixing the served model, sampling temperature, maximum output tokens, engine optimisation, and batching and streaming modes for an inference endpoint.

Responsibility. Sets model and runtime parameters of an inference endpoint.

Also known as: NIM configuration, nim_config, Per-node LLM configuration, Sampling Configuration, Inference Parameter Settings, Dynamic batching configuration, Instance and concurrency settings, Triton model configuration (config.pbtxt), Per-model serving configuration, Model configuration file (model.pbtxt)

configuresconfiguresconfiguresis produced byis constrained byconfiguresis specialized byis specialized byconfiguresis specialized byis specialized byis specialized byis written byis specialized byis produced byis written byLLM Inference Service: configuresLLM Inference ServiceInference Server: configuresInference ServerDynamic Batch Scheduler: configuresDynamic Batch SchedulerAgent Hyperparameter Optimizer: is produced byAgent Hyperparameter Opt…Service Level Objective Specification: is constrained byService Level Objective …Tensor Framework Backend: configuresTensor Framework BackendThroughput-Oriented Batching Config: is specialized byThroughput-Oriented Batc…Latency-Oriented Batching Config: is specialized byLatency-Oriented Batchin…Sequence Batch Scheduler: configuresSequence Batch SchedulerBalanced Batching Config: is specialized byBalanced Batching ConfigModel Deployment Profile: is specialized byModel Deployment ProfileInference Queue Policy: is specialized byInference Queue PolicyAdaptive Batch Size Controller: is written byAdaptive Batch Size Cont…Latency-Tuned Decoding Configuration: is specialized byLatency-Tuned Decoding C…Serving Configuration Optimizer: is produced byServing Configuration Op…Serving Parameter Tuner: is written byServing Parameter Tuner
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Balanced Batching ConfigChoose for mixed or semi-interactive workloads, such as customer service chatbots where users tolerate 300-500 ms latency and traffic allows some batching.
Inference Queue PolicyChoose when serving heterogeneous users with tiered SLAs (premium <50ms, standard <200ms, best-effort batch) on shared infrastructure.
Latency-Oriented Batching ConfigChoose for interactive query-time workloads such as visual question answering where sub-second responses are required.
Latency-Tuned Decoding ConfigurationChoose for interactive applications (chat, code completion, real-time moderation) with short, factual responses where deterministic, concise output is acceptable.
Model Deployment Profile abstract—
Throughput-Oriented Batching ConfigChoose for offline RAG preprocessing (bulk captioning or extraction) where images processed per second matters more than per-image latency.

Relationships

configures structural

is written by dependency

is constrained by control

is produced by lifecycle

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Latency-over-throughput tuningPer-sub-task temperature settingsNucleus (top-p) samplingTop-k sampling
Technologies
NVIDIA NIM
Quality attributes
Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Inconsistent answers to repeated identical queriesDegenerate or off-topic outputs from high-temperature samplingHallucinated references

Sources

  1. Ch1.5B: T. Nguyen, "Stateful Orchestration - Worked Examples," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.5B. ISBN: 9798244538229.
  2. Ch2.2: T. Nguyen, "LangGraph," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.2. ISBN: 9798244538229.
  3. Ch2.3: T. Nguyen, "LangChain Sequential Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.3. ISBN: 9798244538229.
  4. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
  5. Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
  6. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  7. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  8. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  9. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  10. Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
  11. Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
  12. Ref2.01: NVIDIA, "Optimization," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 26, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/optimization.html
  13. Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
  14. Ref7.08: NVIDIA, "NVIDIA Deep Learning Triton Inference Server Documentation," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/