Infrastructure · Infrastructure resource

GPU Allocation Unit

Infrastructure resourceInfrastructureInfrastructureVariation point (abstract)arc:GPUAllocationUnit

An abstract schedulable unit of accelerator capacity (whole GPU, hardware partition, or time-shared slot) onto which the orchestrator places a single workload.

Responsibility. Supplies a workload with accelerator compute and memory under a defined sharing and isolation model.

Also known as: GPU sharing mode, GPU resource

is specialized byis specialized byis specialized byGPU Partition: is specialized byGPU PartitionDedicated GPU Device: is specialized byDedicated GPU DeviceTime-Sliced GPU Share: is specialized byTime-Sliced GPU Share
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Dedicated GPU DeviceChoose for large models (70B+), high-concurrency single-tenant deployments, or latency-critical workloads needing full memory bandwidth and large continuous batches.
GPU PartitionChoose for multi-tenant SaaS with SLA guarantees, mission-critical agents, regulatory isolation needs, or production agents that must not be slowed by co-located batch jobs.
Time-Sliced GPU ShareChoose for internal development, trusted co-located workloads, long-running batch jobs, or cost-constrained deployments that tolerate latency variance and cannot afford partitioning overhead.

Quantitative guidance

As stated by the sources; verify before use.

Classification

Quality attributes
Performance efficiency (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)Cost efficiency
Risks mitigated
GPU underutilisation by small models

Sources

  1. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.