Model Serving · Interface
Tensor Inference API
InterfaceModel ServingModelsarc:TensorInferenceAPI
A framework-agnostic HTTP/gRPC inference interface through which clients submit named input tensors to a specific model and version and receive output tensors.
Responsibility. Provides uniform request/response access to any served model regardless of its framework.
Also known as: Triton HTTP/gRPC endpoint, /v2/models/{model}/infer
Relationships
is exposed by structural
is invoked by dependency
is routed to by dynamic
Design guidance
- SHOULD expose model metadata and configuration endpoints so clients and operators can verify which models and configurations are loaded.
Classification
- Technologies
- NVIDIA Triton Inference ServerHTTPgRPC
- Quality attributes
- Compatibility (ISO/IEC 25010)
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.