Infrastructure · Software component

Pipeline Parallel Executor

Software componentInfrastructureInfrastructurearc:PipelineParallelExecutor

A distributed executor that partitions a model vertically into sequential layer stages on different GPUs, using micro-batching to overlap stages and reduce pipeline bubbles.

Responsibility. Distributes model layers across GPUs or nodes as sequential stages.

Variant of Distributed Model Executor abstract

When to choose. Choose for cross-node scaling over slower interconnects where memory efficiency matters more than per-sample latency.

specializesalternative tois target of alternativeTodeployed onDistributed Model Executor: specializesDistributed Model ExecutorTensor Parallel Executor: alternative toTensor Parallel ExecutorData Parallel Executor: is target of alternativeToData Parallel ExecutorInter-Node Network Fabric: deployed onInter-Node Network Fabric
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Pipeline parallelismMicro-batching

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.