Infrastructure · Software component
Accelerator Operator
Software componentInfrastructureInfrastructurearc:AcceleratorOperator
A cluster controller that automates deployment, upgrade and lifecycle of the per-node accelerator software stack (driver, container runtime hooks, device plugin, telemetry exporter, node labeling) as managed operands.
Responsibility. Manages the lifecycle of accelerator software on cluster nodes.
Also known as: GPU Operator, NVIDIA GPU Operator
Relationships
deployed on structural
orchestrates control
Design guidance
- SHOULD be installed before deploying GPU workloads and validated with a sample GPU pod.
- SHOULD run with RBAC and restricted privileges, with drivers regularly updated for security patches.
Classification
- Patterns
- Kubernetes operator patternDaemonSet node agentsContainerized driver deployment
- Technologies
- NVIDIA GPU OperatorHelm
- Quality attributes
- Maintainability (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Manual per-node driver installation drift
Sources
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.
- Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html