Infrastructure · Data artifact
GPU Partition Layout
Data artifactInfrastructureInfrastructureVariation point (abstract)arc:GPUPartitionLayout
A declarative allocation of a physical GPU into isolated instances with fixed memory and compute fractions assigned per workload.
Responsibility. Specifies how GPU capacity is split among co-located workloads.
Also known as: MIG configuration, MIG partitioning strategy, mig-parted config
Variants
| Variant | When to choose |
|---|---|
| Mixed GPU Partition Layout | Choose for multi-tier offerings and diverse model sizes (7B-70B) when the team accepts extra scheduling and configuration complexity for better utilisation. |
| Uniform GPU Partition Layout | Choose for homogeneous workloads and SaaS platforms with many similar-sized tenants where operational simplicity is the priority. |
| Unpartitioned GPU Layout | Choose for 70B+ models, high-concurrency single-tenant or latency-critical deployments needing a full GPU. |
Relationships
configures structural
is read by dependency
Design guidance
- MUST apply partition layout changes only in maintenance windows, because reconfiguration resets the GPU and terminates all running workloads.
Classification
- Technologies
- NVIDIA Multi-Instance GPU (MIG)NVIDIA Fleet Commandnvidia-mig-manager ConfigMap
- Risks mitigated
- Latency SLA misses from GPU contention
Sources
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.