Infrastructure · Software component
Data Parallel Executor
Software componentInfrastructureInfrastructurearc:DataParallelExecutor
A distributed executor that replicates the full model on each GPU, processes different batches independently and synchronizes gradients with one all-reduce per step.
Responsibility. Scales throughput by replicating the model across GPUs.
Variant of Distributed Model Executor abstract
When to choose. Choose for models under ~13B that fit on one GPU; scales nearly linearly.
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- 5-8% overhead; NVLink gives only 1.1-1.3x benefit over PCIe (Ch7.1A).
Classification
- Patterns
- Data parallelism
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.