Infrastructure · Software component
Fully Sharded Data Parallel Executor
Software componentInfrastructureInfrastructurearc:FullyShardedDataParallelExecutor
A distributed training executor that shards parameters, gradients and optimizer states across GPUs, all-gathering layers for compute and reduce-scattering gradients.
Responsibility. Trains models whose full training state exceeds single-GPU memory.
Also known as: FSDP
Variant of Distributed Model Executor abstract
When to choose. Choose for training models whose weights plus optimizer state exceed GPU memory (e.g., 70B with 280 GB state on 4x80GB).
Relationships
trains lifecycle
- Foundation LLM abstract Ch7.1A
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- NVLink reduces FSDP communication overhead 5-8x, making 200B parameters feasible (Ch7.1A).
Classification
- Patterns
- Fully sharded data parallelism
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.