Infrastructure · Software component

Distributed Model Executor

Software componentInfrastructureInfrastructureVariation point (abstract)arc:DistributedModelExecutor

An execution component that partitions a model's training or inference computation across multiple GPUs according to a parallelism strategy and synchronizes partial results through collective communication.

Responsibility. Runs one model across many GPUs.

Also known as: Multi-GPU parallelism strategy, Model parallel executor

deployed onhostsis specialized byis specialized byis specialized byinvokesis specialized byGPU Node: deployed onGPU NodeRLHF Policy Optimizer: hostsRLHF Policy OptimizerTensor Parallel Executor: is specialized byTensor Parallel ExecutorData Parallel Executor: is specialized byData Parallel ExecutorPipeline Parallel Executor: is specialized byPipeline Parallel ExecutorCollective Communication Library: invokesCollective Communication…Fully Sharded Data Parallel Executor: is specialized byFully Sharded Data Paral…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Data Parallel ExecutorChoose for models under ~13B that fit on one GPU; scales nearly linearly.
Fully Sharded Data Parallel ExecutorChoose for training models whose weights plus optimizer state exceed GPU memory (e.g., 70B with 280 GB state on 4x80GB).
Pipeline Parallel ExecutorChoose for cross-node scaling over slower interconnects where memory efficiency matters more than per-sample latency.
Tensor Parallel ExecutorChoose for memory-constrained models (30-40B+) and latency-sensitive inference, confined to a single-node high-bandwidth GPU interconnect domain.

Relationships

deployed on structural

hosts structural

invokes dependency

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
3D parallelism (DP across nodes, TP within nodes, PP for massive models)

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  2. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.