Model Serving · Interface

Tensor Inference API

InterfaceModel ServingModelsarc:TensorInferenceAPI

A framework-agnostic HTTP/gRPC inference interface through which clients submit named input tensors to a specific model and version and receive output tensors.

Responsibility. Provides uniform request/response access to any served model regardless of its framework.

Also known as: Triton HTTP/gRPC endpoint, /v2/models/{model}/infer

is exposed byis invoked byis routed to byis invoked byis invoked byInference Server: is exposed byInference ServerTool Integration Adapter: is invoked byTool Integration AdapterModel Router: is routed to byModel RouterInference Performance Analyzer: is invoked byInference Performance An…Request Batcher: is invoked byRequest Batcher
Direct neighbourhood (hover for relationship types)

Relationships

is exposed by structural

is invoked by dependency

is routed to by dynamic

Design guidance

Classification

Technologies
NVIDIA Triton Inference ServerHTTPgRPC
Quality attributes
Compatibility (ISO/IEC 25010)

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.