Infrastructure · Infrastructure resource

Time-Sliced GPU Share

Infrastructure resourceInfrastructureInfrastructurearc:TimeSlicedGPUShare

A share of a whole GPU in which the driver schedules co-located processes in round-robin time slices over a single shared memory pool, without hardware isolation.

Responsibility. Lets several workloads share one GPU flexibly, including bursting into memory left idle by others.

Also known as: GPU time-slicing, Software GPU sharing

Variant of GPU Allocation Unit abstract

When to choose. Choose for internal development, trusted co-located workloads, long-running batch jobs, or cost-constrained deployments that tolerate latency variance and cannot afford partitioning overhead.

hostsis target of alternativeTois target of alternativeTospecializesInference Server: hostsInference ServerGPU Partition: is target of alternativeToGPU PartitionDedicated GPU Device: is target of alternativeToDedicated GPU DeviceGPU Allocation Unit: specializesGPU Allocation Unit
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Technologies
CUDA driver time-slicingNVIDIA vGPU
Quality attributes
Flexibility (ISO/IEC 25010)Cost efficiency

Sources

  1. Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.