Infrastructure · Infrastructure resource
Peer-to-Peer GPU Link
Infrastructure resourceInfrastructureInfrastructurearc:PeerToPeerGPULink
A short-reach, cache-coherent point-to-point link between GPUs offering far higher bandwidth than a general-purpose bus and unified virtual addressing for peer memory access.
Responsibility. Provides high-bandwidth coherent GPU-to-GPU links.
Also known as: NVLink
Variant of GPU Interconnect abstract
When to choose. Choose for tensor-parallel inference or training of >40B models with sustained utilization (100+ GPU-hours/month).
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- 600-1800 GB/s aggregate; NVLink 4.0 at 50 GT/s; H100 P2P achieves 200-250 GB/s of 900 GB/s nominal (Ch7.1A).
- H100 SXM $1.95/h vs $1.37/h PCIe (42% premium) yet 70B training 8% cheaper and inference cost per token halves ($0.024 to $0.011) (Ch7.1A).
Classification
- Technologies
- NVIDIA NVLink 3.0/4.0
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.