Infrastructure · Software component

Inference Service Operator

Software componentInfrastructureInfrastructurearc:InferenceServiceOperator

A container-orchestrator extension that reconciles declarative inference-service resources into running model-serving deployments, handling image pulls, GPU allocation, health checks, service exposure and scaling.

Responsibility. Manages the lifecycle of inference microservice deployments from declared desired state.

Also known as: NVIDIA NIM Operator, Kubernetes operator for inference, NIM Operator

scales; orchestratesdeployed onis configured byreadsLLM Inference Service: scales; orchestratesLLM Inference ServiceContainer Orchestrator: deployed onContainer OrchestratorDeployment Manifest: is configured byDeployment ManifestContainer and Model Artifact Registry: readsContainer and Model Arti…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is configured by structural

reads dependency

orchestrates control

scales control

Design guidance

Classification

Patterns
Kubernetes operatorDeclarative desired-state reconciliationTraffic splittingCanary deploymentAutoscaling
Technologies
NVIDIA NIM OperatorKubernetes custom resourcesKServe
Quality attributes
Maintainability (ISO/IEC 25010)

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/