Infrastructure · Infrastructure resource

GPU Partition

Infrastructure resourceInfrastructureInfrastructurearc:GPUPartition

A hardware-isolated slice of a physical GPU with dedicated memory and compute that the orchestrator treats as an independent GPU.

Responsibility. Provides isolated, right-sized accelerator capacity to one workload on shared hardware.

Also known as: MIG instance, Multi-Instance GPU, MIG slice, GPU Instance (GI), MIG device

Variant of GPU Allocation Unit abstract

When to choose. Choose for multi-tenant SaaS with SLA guarantees, mission-critical agents, regulatory isolation needs, or production agents that must not be slowed by co-located batch jobs.

hostshostsis monitored byhostsis monitored byis configured byis scaled byis target of alternativeTois configured byalternative tois configured byspecializesLLM Inference Service: hostsLLM Inference ServiceInference Server: hostsInference ServerPlatform Operator: is monitored byPlatform OperatorINT8 Quantized Engine: hostsINT8 Quantized EngineGPU Telemetry Exporter: is monitored byGPU Telemetry ExporterGPU Partition Layout: is configured byGPU Partition LayoutGPU Partition Manager: is scaled byGPU Partition ManagerDedicated GPU Device: is target of alternativeToDedicated GPU DeviceMixed GPU Partition Layout: is configured byMixed GPU Partition LayoutTime-Sliced GPU Share: alternative toTime-Sliced GPU ShareUniform GPU Partition Layout: is configured byUniform GPU Partition La…GPU Allocation Unit: specializesGPU Allocation Unit
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

is configured by structural

is scaled by control

is monitored by assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Hardware-enforced partitioningProfiles 1g.10gb / 2g.20gb / 3g.39gb / 7g.79gb (A100-80GB)Partition right-sizingDedicated L2 cache banks, memory controllers and DRAM paths
Technologies
NVIDIA Multi-Instance GPU (MIG)A100H100A30H100 NVL
Quality attributes
Performance efficiency (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)
Risks mitigated
Underutilised full GPUs for small modelsResource contention between co-located workloadsNoisy-neighbour latency interferenceCross-tenant memory corruptionMemory-exhaustion cascades across co-located processes

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
  3. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.