Infrastructure · Software component

Data Parallel Executor

Software componentInfrastructureInfrastructurearc:DataParallelExecutor

A distributed executor that replicates the full model on each GPU, processes different batches independently and synchronizes gradients with one all-reduce per step.

Responsibility. Scales throughput by replicating the model across GPUs.

Variant of Distributed Model Executor abstract

When to choose. Choose for models under ~13B that fit on one GPU; scales nearly linearly.

specializesalternative toalternative toalternative toDistributed Model Executor: specializesDistributed Model ExecutorTensor Parallel Executor: alternative toTensor Parallel ExecutorPipeline Parallel Executor: alternative toPipeline Parallel ExecutorFully Sharded Data Parallel Executor: alternative toFully Sharded Data Paral…
Direct neighbourhood (hover for relationship types)

Relationships

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Data parallelism

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.