Model Serving · Data artifact

Inference Service Image

Data artifactModel ServingModelsVariation point (abstract)arc:InferenceServiceImage

A container image packaging an inference runtime, engine-selection logic and a standard API so the same image deploys identically across cloud, data center and workstation GPUs.

Responsibility. Packages a portable, ready-to-run inference service.

Also known as: Inference microservice container, NIM container

configuresis specialized byis specialized byInference Server: configuresInference ServerModel-Specific Inference Image: is specialized byModel-Specific Inference…Multi-Model Inference Image: is specialized byMulti-Model Inference Im…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Model-Specific Inference ImageChoose for production with SLA or compliance requirements (SOC 2, HIPAA, FedRAMP) and mission-critical services where 10-15% latency gains justify reduced flexibility.
Multi-Model Inference ImageChoose for research, experimentation, custom fine-tuned model pipelines, multi-model task switching and prototyping where experimentation velocity outweighs production stability.

Relationships

configures structural

Design guidance

Classification

Technologies
NVIDIA NIMDocker
Quality attributes
Flexibility (ISO/IEC 25010)Maintainability (ISO/IEC 25010)

Sources

  1. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.