Infrastructure · Software component
Inference Service Operator
Software componentInfrastructureInfrastructurearc:InferenceServiceOperator
A container-orchestrator extension that reconciles declarative inference-service resources into running model-serving deployments, handling image pulls, GPU allocation, health checks, service exposure and scaling.
Responsibility. Manages the lifecycle of inference microservice deployments from declared desired state.
Also known as: NVIDIA NIM Operator, Kubernetes operator for inference, NIM Operator
Relationships
deployed on structural
is configured by structural
reads dependency
orchestrates control
scales control
Design guidance
- SHOULD manage inference deployments declaratively through manifests rather than manual container management.
- MUST NOT be assumed to remove the need for underlying orchestrator knowledge (scheduling, resource limits, networking, persistent volumes, secrets).
Classification
- Patterns
- Kubernetes operatorDeclarative desired-state reconciliationTraffic splittingCanary deploymentAutoscaling
- Technologies
- NVIDIA NIM OperatorKubernetes custom resourcesKServe
- Quality attributes
- Maintainability (ISO/IEC 25010)