Model Serving · Software component

Inference Backend

Software componentModel ServingModelsVariation point (abstract)arc:InferenceBackend

A pluggable execution module, loaded dynamically by an inference server according to model configuration, that runs a model in one specific framework runtime.

Responsibility. Executes inference for models of one framework behind the server's backend contract.

Also known as: Triton backend

deployed onis specialized byis specialized byis specialized byis specialized byis orchestrated byInference Server: deployed onInference ServerLLM Generation Backend: is specialized byLLM Generation BackendPre-compiled Engine Backend: is specialized byPre-compiled Engine Back…Tensor Framework Backend: is specialized byTensor Framework BackendPortable LLM Runtime Backend: is specialized byPortable LLM Runtime Bac…Model Ensemble Orchestrator: is orchestrated byModel Ensemble Orchestra…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
LLM Generation BackendChoose for large language model generation, where the engine's continuous or in-flight batching outperforms request-level batching.
Portable LLM Runtime BackendChoose for uncommon or heterogeneous GPUs (RTX 4090, A10G, V100, T4), edge and mixed research clusters, or development prioritizing iteration speed.
Pre-compiled Engine BackendChoose when pre-compiled engines exist for the deployed GPU (e.g., A100 40/80GB, H100 80GB, L40S 48GB).
Tensor Framework BackendChoose for non-autoregressive models (classifiers, encoders, vision, tree-based models, custom Python logic) that benefit from server-side dynamic batching; OpenVINO for Intel CPU or edge CPU-only infrastructure.

Relationships

deployed on structural

is orchestrated by control

Design guidance

Classification

Patterns
Plugin architecture
Technologies
NVIDIA Triton Inference Server
Quality attributes
Maintainability (ISO/IEC 25010)Compatibility (ISO/IEC 25010)

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  3. Ref7.08: NVIDIA, "NVIDIA Deep Learning Triton Inference Server Documentation," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/