Infrastructure · Infrastructure resource
GPU Allocation Unit
Infrastructure resourceInfrastructureInfrastructureVariation point (abstract)arc:GPUAllocationUnit
An abstract schedulable unit of accelerator capacity (whole GPU, hardware partition, or time-shared slot) onto which the orchestrator places a single workload.
Responsibility. Supplies a workload with accelerator compute and memory under a defined sharing and isolation model.
Also known as: GPU sharing mode, GPU resource
Variants
| Variant | When to choose |
|---|---|
| Dedicated GPU Device | Choose for large models (70B+), high-concurrency single-tenant deployments, or latency-critical workloads needing full memory bandwidth and large continuous batches. |
| GPU Partition | Choose for multi-tenant SaaS with SLA guarantees, mission-critical agents, regulatory isolation needs, or production agents that must not be slowed by co-located batch jobs. |
| Time-Sliced GPU Share | Choose for internal development, trusted co-located workloads, long-running batch jobs, or cost-constrained deployments that tolerate latency variance and cannot afford partitioning overhead. |
Quantitative guidance
As stated by the sources; verify before use.
- A 7B LLM agent uses only 10-15% of an A100's compute and 8-12 GB of its 80 GB memory, leaving 85-90% idle (Ch7.6).
Classification
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)Cost efficiency
- Risks mitigated
- GPU underutilisation by small models
Sources
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.