Infrastructure · Software component

Accelerator Operator

Software componentInfrastructureInfrastructurearc:AcceleratorOperator

A cluster controller that automates deployment, upgrade and lifecycle of the per-node accelerator software stack (driver, container runtime hooks, device plugin, telemetry exporter, node labeling) as managed operands.

Responsibility. Manages the lifecycle of accelerator software on cluster nodes.

Also known as: GPU Operator, NVIDIA GPU Operator

deployed onorchestratesorchestratesorchestratesorchestratesorchestratesContainer Orchestrator: deployed onContainer OrchestratorGPU Telemetry Exporter: orchestratesGPU Telemetry ExporterGPU Partition Manager: orchestratesGPU Partition ManagerAccelerator Container Runtime: orchestratesAccelerator Container Ru…GPU Device Plugin: orchestratesGPU Device PluginNode Capability Labeler: orchestratesNode Capability Labeler
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

orchestrates control

Design guidance

Classification

Patterns
Kubernetes operator patternDaemonSet node agentsContainerized driver deployment
Technologies
NVIDIA GPU OperatorHelm
Quality attributes
Maintainability (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Manual per-node driver installation drift

Sources

  1. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.
  2. Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html