Infrastructure · Infrastructure resource

Dedicated GPU Device

Infrastructure resourceInfrastructureInfrastructurearc:DedicatedGPUDevice

A whole, unpartitioned physical GPU allocated exclusively to one workload, giving it all compute, memory and memory bandwidth.

Responsibility. Provides the full capacity of one GPU to a single workload.

Also known as: Full GPU allocation, MIG strategy none

Variant of GPU Allocation Unit abstract

When to choose. Choose for large models (70B+), high-concurrency single-tenant deployments, or latency-critical workloads needing full memory bandwidth and large continuous batches.

hostsalternative toalternative tois configured byspecializesInference Server: hostsInference ServerGPU Partition: alternative toGPU PartitionTime-Sliced GPU Share: alternative toTime-Sliced GPU ShareUnpartitioned GPU Layout: is configured byUnpartitioned GPU LayoutGPU Allocation Unit: specializesGPU Allocation Unit
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

is configured by structural

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Classification

Quality attributes
Performance efficiency (ISO/IEC 25010)

Sources

  1. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.