Infrastructure · Infrastructure resource
GPU Interconnect
Infrastructure resourceInfrastructureInfrastructureVariation point (abstract)arc:GPUInterconnect
The intra-node data path linking GPUs for peer memory access and collective communication, whose bandwidth and latency bound multi-GPU parallel efficiency.
Responsibility. Carries inter-GPU traffic within a node.
Variants
| Variant | When to choose |
|---|---|
| GPU Switch Fabric | Choose for 8+ GPU nodes and high-batch tensor-parallel inference where all-reduce bandwidth dominates. |
| PCIe GPU Bus | Choose for smaller or lower-utilization deployments and data-parallel workloads where interconnect bandwidth is not the bottleneck. |
| Peer-to-Peer GPU Link | Choose for tensor-parallel inference or training of >40B models with sustained utilization (100+ GPU-hours/month). |
Relationships
hosts structural
Design guidance
- SHOULD NOT be treated as pooled memory: a coherent interconnect enables peer access but each GPU keeps its own capacity.
- SHOULD NOT be sized on nominal aggregate bandwidth; point-to-point transfers achieve only 22-28% of it.
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.