Infrastructure · Software component
Collective Communication Library
Software componentInfrastructureInfrastructurearc:CollectiveCommunicationLibrary
A GPU communication library executing collective operations (all-reduce, all-gather, reduce-scatter) with topology-aware algorithms chosen for the detected interconnect.
Responsibility. Synchronizes tensors across GPUs efficiently.
Relationships
deployed on structural
is invoked by dependency
- Distributed Model Executor abstract Ch7.1A
Quantitative guidance
As stated by the sources; verify before use.
- 8-GPU all-reduce with NVSwitch approaches 600-700 GB/s; small 4KB transfers achieve only 15-20 GB/s (Ch7.1A).
Classification
- Patterns
- Ring all-reduceTree all-reduceDouble-binary-treeDirect peer-to-peer all-reduce on fully connected topologies
- Technologies
- NVIDIA NCCL
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.