Infrastructure · Software component

Tensor Parallel Executor

Software componentInfrastructureInfrastructurearc:TensorParallelExecutor

A distributed executor that shards each layer's weight matrices across GPUs, computes partial results in parallel and combines them with all-reduce after every attention and feed-forward block.

Responsibility. Fits and accelerates a large model by splitting each layer across GPUs.

Variant of Distributed Model Executor abstract

When to choose. Choose for memory-constrained models (30-40B+) and latency-sensitive inference, confined to a single-node high-bandwidth GPU interconnect domain.

hostshostsspecializesis configured bydeployed onis target of alternativeTois target of alternativeToFoundation LLM: hostsFoundation LLMOptimized Inference Engine: hostsOptimized Inference EngineDistributed Model Executor: specializesDistributed Model ExecutorEngine Build Configuration: is configured byEngine Build ConfigurationGPU Interconnect: deployed onGPU InterconnectData Parallel Executor: is target of alternativeToData Parallel ExecutorPipeline Parallel Executor: is target of alternativeToPipeline Parallel Executor
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

hosts structural

is configured by structural

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Tensor parallelismHorizontal layer shardingTime-sliced multi-model sharing across GPUs

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.