Infrastructure · Infrastructure resource
GPU Partition
Infrastructure resourceInfrastructureInfrastructurearc:GPUPartition
A hardware-isolated slice of a physical GPU with dedicated memory and compute that the orchestrator treats as an independent GPU.
Responsibility. Provides isolated, right-sized accelerator capacity to one workload on shared hardware.
Also known as: MIG instance, Multi-Instance GPU, MIG slice, GPU Instance (GI), MIG device
Variant of GPU Allocation Unit abstract
When to choose. Choose for multi-tenant SaaS with SLA guarantees, mission-critical agents, regulatory isolation needs, or production agents that must not be slowed by co-located batch jobs.
Relationships
hosts structural
is configured by structural
is scaled by control
is monitored by assurance
alternative to variability
Design guidance
- SHOULD partition GPUs when serving many small models that would underutilise full GPUs.
- SHOULD give latency-sensitive workloads a dedicated partition so co-located throughput workloads cannot starve them.
- MUST NOT be assumed on GPUs without partitioning support (e.g., Ada Lovelace L40S); such platforms must use time-slicing or vGPU.
- SHOULD be sized for model weights plus KV-cache and batching headroom, not weights alone.
- SHOULD be slightly oversized for high-throughput deployments where larger batches improve cost per token.
- MUST account for 5-10% partitioning overhead in latency-critical capacity planning.
Quantitative guidance
As stated by the sources; verify before use.
- 10 small models needing 10 GPUs fit on 2 GPUs with 5 MIG instances each (Ch4.5).
- A100 partitions into up to seven instances; retail example 4/7 vision, 1/7 conversational agent, 2/7 forecasting (Ch4.6).
- Up to 7 isolated instances per GPU; A100-80GB has 8 memory slices (10 GB) and 7 compute slices; Hopper offers 8 SM slices (Ch7.6).
- 7B INT8 (3.5 GB) fits 1g.10gb; 13B (6.5 GB) needs 2g.20gb; 70B quantized (35 GB) needs 3g.39gb or larger (Ch7.6).
- A 50-tenant platform drops from 50 to 7 A100s (-86% capex) at 1g partitions; utilisation 65-75% vs 12%; P99 280 ms vs 500 ms SLA; $450 vs $3,000 per tenant-year (Ch7.6).
- 7B on 1g.10gb OOMs at batch 32 while 2g.20gb reaches 380 tokens/s (35% higher) (Ch7.6).
- 7B reaches 320-330 tokens/s on 1g.10gb vs 350 on a dedicated A100 (5-10% overhead) (Ch7.6).
Classification
- Patterns
- Hardware-enforced partitioningProfiles 1g.10gb / 2g.20gb / 3g.39gb / 7g.79gb (A100-80GB)Partition right-sizingDedicated L2 cache banks, memory controllers and DRAM paths
- Technologies
- NVIDIA Multi-Instance GPU (MIG)A100H100A30H100 NVL
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)
- Risks mitigated
- Underutilised full GPUs for small modelsResource contention between co-located workloadsNoisy-neighbour latency interferenceCross-tenant memory corruptionMemory-exhaustion cascades across co-located processes
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch7.6: T. Nguyen, "Multi-Instance GPU," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.6. ISBN: 9798244538229.