Infrastructure · Infrastructure resource
Dedicated GPU Device
Infrastructure resourceInfrastructureInfrastructurearc:DedicatedGPUDevice
A whole, unpartitioned physical GPU allocated exclusively to one workload, giving it all compute, memory and memory bandwidth.
Responsibility. Provides the full capacity of one GPU to a single workload.
Also known as: Full GPU allocation, MIG strategy none
Variant of GPU Allocation Unit abstract
When to choose. Choose for large models (70B+), high-concurrency single-tenant deployments, or latency-critical workloads needing full memory bandwidth and large continuous batches.
Relationships
hosts structural
is configured by structural
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- One GPU per tenant for 50 tenants needs 50 A100s ($750,000) at ~12% utilisation each (Ch7.6).
- A single agent serving 10,000 concurrent users with sub-100ms SLA needs the full A100 for 32-64-request batches (Ch7.6).
Classification
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
Sources
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.