Model Serving · Interface
Self-Hosted Inference Endpoint
InterfaceModel ServingModelsarc:SelfHostedInferenceEndpoint
A model inference endpoint served from the organisation's own GPU infrastructure that remains compatible with common function-calling request and response formats.
Responsibility. Provides function-calling inference within the organisation's own infrastructure boundary.
Also known as: On-premises inference endpoint, NIM endpoint
Variant of Native Function-Calling API abstract
When to choose. Choose when data residency requirements, security policies, or cost considerations prevent using cloud-hosted LLM APIs, or for high-throughput tool-heavy agents.
Relationships
is exposed by structural
alternative to variability
Design guidance
- SHOULD maintain API compatibility with common function-calling formats so agent code changes only the endpoint.
Quantitative guidance
As stated by the sources; verify before use.
- On-premises NIM deployment eliminates cloud round-trips, reducing TTFT by 200-400 ms (Ch2.9).
Classification
- Technologies
- NVIDIA NIM
- Quality attributes
- Privacy (NIST AI RMF: privacy-enhanced)Performance efficiency (ISO/IEC 25010)Cost efficiency
Sources
- Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/