Model Serving · Data artifact
Inference Service Image
Data artifactModel ServingModelsVariation point (abstract)arc:InferenceServiceImage
A container image packaging an inference runtime, engine-selection logic and a standard API so the same image deploys identically across cloud, data center and workstation GPUs.
Responsibility. Packages a portable, ready-to-run inference service.
Also known as: Inference microservice container, NIM container
Variants
| Variant | When to choose |
|---|---|
| Model-Specific Inference Image | Choose for production with SLA or compliance requirements (SOC 2, HIPAA, FedRAMP) and mission-critical services where 10-15% latency gains justify reduced flexibility. |
| Multi-Model Inference Image | Choose for research, experimentation, custom fine-tuned model pipelines, multi-model task switching and prototyping where experimentation velocity outweighs production stability. |
Relationships
configures structural
Design guidance
- SHOULD carry environment-specific settings (GPU counts, memory limits, network policies) in orchestration manifests rather than code.
Classification
- Technologies
- NVIDIA NIMDocker
- Quality attributes
- Flexibility (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
Sources
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.