Infrastructure · Infrastructure resource
Time-Sliced GPU Share
Infrastructure resourceInfrastructureInfrastructurearc:TimeSlicedGPUShare
A share of a whole GPU in which the driver schedules co-located processes in round-robin time slices over a single shared memory pool, without hardware isolation.
Responsibility. Lets several workloads share one GPU flexibly, including bursting into memory left idle by others.
Also known as: GPU time-slicing, Software GPU sharing
Variant of GPU Allocation Unit abstract
When to choose. Choose for internal development, trusted co-located workloads, long-running batch jobs, or cost-constrained deployments that tolerate latency variance and cannot afford partitioning overhead.
Relationships
hosts structural
alternative to variability
Design guidance
- SHOULD NOT host SLA-critical, multi-tenant or latency-sensitive agent inference.
Quantitative guidance
As stated by the sources; verify before use.
- Time slices are typically 50-200 ms; a long co-tenant kernel can block others for 500 ms+ (Ch7.6).
- 4 processes on A100-80GB (7B, batch 1): P50 120 ms, P99 450 ms, +/-250 ms vs MIG 1g.10gb P50 75 ms, P99 95 ms, +/-15 ms (Ch7.6).
Classification
- Technologies
- CUDA driver time-slicingNVIDIA vGPU
- Quality attributes
- Flexibility (ISO/IEC 25010)Cost efficiency
Sources
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.