Infrastructure · Software component
Distributed Model Executor
Software componentInfrastructureInfrastructureVariation point (abstract)arc:DistributedModelExecutor
An execution component that partitions a model's training or inference computation across multiple GPUs according to a parallelism strategy and synchronizes partial results through collective communication.
Responsibility. Runs one model across many GPUs.
Also known as: Multi-GPU parallelism strategy, Model parallel executor
Variants
| Variant | When to choose |
|---|---|
| Data Parallel Executor | Choose for models under ~13B that fit on one GPU; scales nearly linearly. |
| Fully Sharded Data Parallel Executor | Choose for training models whose weights plus optimizer state exceed GPU memory (e.g., 70B with 280 GB state on 4x80GB). |
| Pipeline Parallel Executor | Choose for cross-node scaling over slower interconnects where memory efficiency matters more than per-sample latency. |
| Tensor Parallel Executor | Choose for memory-constrained models (30-40B+) and latency-sensitive inference, confined to a single-node high-bandwidth GPU interconnect domain. |
Relationships
deployed on structural
hosts structural
invokes dependency
Design guidance
- SHOULD NOT assume N GPUs give Nx speedup; serial all-reduce consumes 5-15% of time even at 900 GB/s (Amdahl's law).
Quantitative guidance
As stated by the sources; verify before use.
- 175B model with DP=32, TP=4, PP=1 achieves 120-way parallelism at 75-85% efficiency (Ch7.1A).
Classification
- Patterns
- 3D parallelism (DP across nodes, TP within nodes, PP for massive models)
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.