Infrastructure · Infrastructure resource

GPU Interconnect

Infrastructure resourceInfrastructureInfrastructureVariation point (abstract)arc:GPUInterconnect

The intra-node data path linking GPUs for peer memory access and collective communication, whose bandwidth and latency bound multi-GPU parallel efficiency.

Responsibility. Carries inter-GPU traffic within a node.

hostshostsis specialized byis specialized byis specialized byTensor Parallel Executor: hostsTensor Parallel ExecutorCollective Communication Library: hostsCollective Communication…GPU Switch Fabric: is specialized byGPU Switch FabricPCIe GPU Bus: is specialized byPCIe GPU BusPeer-to-Peer GPU Link: is specialized byPeer-to-Peer GPU Link
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
GPU Switch FabricChoose for 8+ GPU nodes and high-batch tensor-parallel inference where all-reduce bandwidth dominates.
PCIe GPU BusChoose for smaller or lower-utilization deployments and data-parallel workloads where interconnect bandwidth is not the bottleneck.
Peer-to-Peer GPU LinkChoose for tensor-parallel inference or training of >40B models with sustained utilization (100+ GPU-hours/month).

Relationships

hosts structural

Design guidance

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.