Model Serving · Interface
OpenAI-Compatible Inference API
InterfaceModel ServingModelsarc:OpenAICompatibleInferenceAPI
A standardised HTTP inference API following OpenAI request formats, through which text, vision and embedding models are called uniformly regardless of the underlying model.
Responsibility. Provides a uniform, model-agnostic endpoint for inference requests.
Also known as: Unified model API, NIM OpenAI-compatible endpoint, /v1/chat/completions, Triton OpenAI-compatible API mode, OpenAI-compatible embeddings endpoint, Chat Completion API, Embedding API
Relationships
is exposed by structural
is cached by dependency
is invoked by dependency
- Agent Controller abstract Ch7.3
- Answer Synthesizer abstract Ch2.7 Ref2.07
- Concurrent Inference Request Dispatcher Ch7.2
- Function-Calling Controller Ref7.13
- Guardrail Orchestrator Ch7.1B
- Image-to-Text Grounder abstract Ch2.7
- Joint Multimodal Embedding Service Ch2.7
- LLM Provider Adapter Ch4.5
- Model Health Prober Ref7.16
- Multimodal Answer Synthesizer Ch2.7
- ReAct Agent Controller Ref7.07 Ref7.13
- Retry Handler Ch2.7
- VLM-based Image Type Classifier Ch2.7
has access controlled by control
is guarded by control
Design guidance
- SHOULD be used so applications can move from hosted commercial APIs to on-premises endpoints without code refactoring.
- MUST be protected by API-key authentication when using hosted endpoints.
- SHOULD expose OpenAI-compatible endpoints so existing agent frameworks switch from a hosted provider to self-hosted serving by changing only the endpoint URL.
- SHOULD protect the inference endpoint with strong API keys, TLS encryption and rate limiting.
Classification
- Patterns
- Drop-in endpoint replacement
- Technologies
- NVIDIA NIMOpenAI Python libraryLangChainLlamaIndexNVIDIA Triton Inference ServervLLMAsyncOpenAI clientOpenAI Python client
- Quality attributes
- Compatibility (ISO/IEC 25010)Flexibility (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Vendor lock-inRefactoring on model migrationAgent code changes when switching model providers
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch6.1: T. Nguyen, "Embeddings and RAG Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.1. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.3: T. Nguyen, "NeMo Agent Toolkit Profiling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.3. ISBN: 9798244538229.
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.13: NVIDIA, "Llama Nemotron," NVIDIA NeMo Framework User Guide, v25.09. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo-framework/user-guide/25.09/llms/llama_nemotron.html
- Ref7.14: "NVIDIA Agentic AI Platform Ecosystem Integration," unpublished reference note (14-NVIDIA-Ecosystem-Integration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note