Model Serving · Model asset
TF32 Inference Engine
Model assetModel ServingModelsarc:TF32InferenceEngine
A model execution configuration using the TensorFloat-32 format (FP32 exponent range with a 10-bit mantissa) that gives near-FP32 accuracy with FP16-like speed and no quantization workflow.
Responsibility. Accelerates inference on Ampere-or-newer GPUs without calibration or explicit quantization.
Also known as: TensorFloat-32
Variant of Optimized Inference Engine abstract
When to choose. Choose when targeting Ampere-or-newer hardware exclusively and near-FP32 accuracy is required without a quantization workflow.
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- <0.1% accuracy loss with 4-6x speedup over FP32 on compute capability 8.0+ (Ch7.4, Ref7.01).
Classification
- Technologies
- NVIDIA Ampere Tensor Cores
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
Sources
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html