Infrastructure · Software component
Pipeline Parallel Executor
Software componentInfrastructureInfrastructurearc:PipelineParallelExecutor
A distributed executor that partitions a model vertically into sequential layer stages on different GPUs, using micro-batching to overlap stages and reduce pipeline bubbles.
Responsibility. Distributes model layers across GPUs or nodes as sequential stages.
Variant of Distributed Model Executor abstract
When to choose. Choose for cross-node scaling over slower interconnects where memory efficiency matters more than per-sample latency.
Relationships
deployed on structural
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- Per-sample latency increases 4x for a 4-stage pipeline; 2-GPU PP achieves 1.4-1.6x rather than 2x (Ch7.1A).
Classification
- Patterns
- Pipeline parallelismMicro-batching
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.