Model Serving · Interface

Hosted Provider Inference API

InterfaceModel ServingModelsarc:HostedProviderInferenceAPI

A cloud-hosted model provider's API offering function calling under provider-specific request, response, and schema-dialect conventions.

Responsibility. Provides function-calling inference as an external managed service.

Also known as: Cloud LLM API

Variant of Native Function-Calling API abstract

When to choose. Choose when data residency, security policy, and cost considerations permit sending prompts to a cloud-hosted LLM provider.

is invoked byis routed to byspecializesalternative toAnswer Synthesizer: is invoked byAnswer SynthesizerModel Router: is routed to byModel RouterNative Function-Calling API: specializesNative Function-Calling …Self-Hosted Inference Endpoint: alternative toSelf-Hosted Inference En…
Direct neighbourhood (hover for relationship types)

Relationships

is invoked by dependency

is routed to by dynamic

alternative to variability

Design guidance

Classification

Technologies
OpenAI APIAnthropic Claude APIGoogle Gemini APIAnthropic API
Quality attributes
Maintainability (ISO/IEC 25010)

Sources

  1. Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
  2. Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
  3. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
  4. Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
  5. Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/