Infrastructure · Software component
Tensor Parallel Executor
Software componentInfrastructureInfrastructurearc:TensorParallelExecutor
A distributed executor that shards each layer's weight matrices across GPUs, computes partial results in parallel and combines them with all-reduce after every attention and feed-forward block.
Responsibility. Fits and accelerates a large model by splitting each layer across GPUs.
Variant of Distributed Model Executor abstract
When to choose. Choose for memory-constrained models (30-40B+) and latency-sensitive inference, confined to a single-node high-bandwidth GPU interconnect domain.
Relationships
deployed on structural
- GPU Interconnect abstract Ch7.1A
hosts structural
- Foundation LLM abstract Ch7.1A
- Optimized Inference Engine abstract Ch7.1A
is configured by structural
alternative to variability
Design guidance
- SHOULD be confined to a single-node NVLink domain, with pipeline parallelism used across nodes.
Quantitative guidance
As stated by the sources; verify before use.
- 70B split 4 ways achieves 3.4-3.7x speedup (85-92% efficiency) on NVLink; cross-node InfiniBand reduces efficiency to 65-80% (Ch7.1A).
- ~160 all-reduces per forward pass for 80 layers; per-token latency grows only 5-10% (Ch7.1A).
- Llama 70B TP=2 on H100: 142 tokens/s/GPU at batch 32 with NVSwitch vs 112 without and 68 with PCIe (Ch7.1A).
- Sharing 8 GPUs via TP with time-slicing raises utilization from 60-70% to 85-90% (Ch7.1A).
Classification
- Patterns
- Tensor parallelismHorizontal layer shardingTime-sliced multi-model sharing across GPUs
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.