Model Serving · Interface

Native Function-Calling API

InterfaceModel ServingModelsVariation point (abstract)arc:ModelInferenceAPI

A model-serving interface that accepts tool definitions as JSON schemas and returns selected tool calls as guaranteed-valid structured JSON objects instead of free text requiring parsing.

Responsibility. Returns schema-conformant structured tool calls from a function-calling-tuned model.

Also known as: OpenAI function-calling API, Structured tool calling, LLM API endpoint, Function calling API, Chat completions endpoint, Function Calling API

is exposed byis invoked byis invoked byis configured byis invoked byis invoked byis invoked byis specialized byis specialized byLLM Inference Service: is exposed byLLM Inference ServiceReasoning Engine: is invoked byReasoning EngineAnswer Synthesizer: is invoked byAnswer SynthesizerTool Schema: is configured byTool SchemaDirect Tool-Calling Controller: is invoked byDirect Tool-Calling Cont…Request Batcher: is invoked byRequest BatcherPrompt Chain Orchestrator: is invoked byPrompt Chain OrchestratorHosted Provider Inference API: is specialized byHosted Provider Inferenc…Self-Hosted Inference Endpoint: is specialized bySelf-Hosted Inference En…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Hosted Provider Inference APIChoose when data residency, security policy, and cost considerations permit sending prompts to a cloud-hosted LLM provider.
Self-Hosted Inference EndpointChoose when data residency requirements, security policies, or cost considerations prevent using cloud-hosted LLM APIs, or for high-throughput tool-heavy agents.

Relationships

is configured by structural

is exposed by structural

is invoked by dependency

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Native function calling
Technologies
OpenAI function callingOpenAI function calling (functions/function_call)Anthropic tool use (tools/tool_use)Google Gemini function calling
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Compatibility (ISO/IEC 25010)Flexibility (ISO/IEC 25010)
Risks mitigated
Tool-call parse errorsConfusion between similar tools

Sources

  1. Ch2.3: T. Nguyen, "LangChain Sequential Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.3. ISBN: 9798244538229.
  2. Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
  3. Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.