Model Serving · Software component

Tensor Framework Backend

Software componentModel ServingModelsarc:TensorFrameworkBackend

An inference backend that executes a fixed-graph model (TensorFlow, TorchScript, ONNX, TensorRT, OpenVINO, tree ensembles, Python) on batched input tensors supplied by the server.

Responsibility. Runs one framework's model on server-assembled tensor batches.

Also known as: TensorFlow backend, PyTorch (TorchScript) backend, ONNX Runtime backend, TensorRT backend, Python backend, OpenVINO backend, FIL backend

Variant of Inference Backend abstract

When to choose. Choose for non-autoregressive models (classifiers, encoders, vision, tree-based models, custom Python logic) that benefit from server-side dynamic batching; OpenVINO for Intel CPU or edge CPU-only infrastructure.

deployed onhostsis configured byis invoked byis target of alternativeTois target of excludesspecializesis routed to byInference Server: deployed onInference ServerOptimized Inference Engine: hostsOptimized Inference EngineInference Serving Configuration: is configured byInference Serving Config…Dynamic Batch Scheduler: is invoked byDynamic Batch SchedulerLLM Generation Backend: is target of alternativeToLLM Generation BackendIn-Flight Batch Scheduler: is target of excludesIn-Flight Batch SchedulerInference Backend: specializesInference BackendSequence Batch Scheduler: is routed to bySequence Batch Scheduler
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

hosts structural

is configured by structural

is invoked by dependency

is routed to by dynamic

alternative to variability

excludes variability

Classification

Technologies
TensorFlowPyTorch TorchScriptONNX RuntimeTensorRTOpenVINOForest Inference Library (XGBoost, LightGBM)Pythontensorrt_planTensorRT backendONNX Runtime backendPyTorch backendTensorFlow backendPython backend

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  3. Ref7.08: NVIDIA, "NVIDIA Deep Learning Triton Inference Server Documentation," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/
  4. Ref7.14: "NVIDIA Agentic AI Platform Ecosystem Integration," unpublished reference note (14-NVIDIA-Ecosystem-Integration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note