Model Serving · Model asset
FP8 Quantized Engine
Model assetModel ServingModelsarc:FP8InferenceEngine
A compiled model engine using 8-bit floating-point precision for higher throughput with minimal accuracy loss on supporting hardware.
Responsibility. Serves the model at FP8 precision for higher throughput.
Also known as: FP8 quantized engine
Variant of Optimized Inference Engine abstract
When to choose. Choose on GPUs with FP8 support when roughly doubled throughput with minimal accuracy loss is acceptable for the model.
Relationships
deployed on structural
is produced by lifecycle
alternative to variability
Design guidance
- MUST NOT be deployed on Ampere, Volta or earlier GPUs; those require INT8 instead.
Quantitative guidance
As stated by the sources; verify before use.
- Doubles throughput on H100 GPUs with minimal accuracy loss for many models (Ch4.4).
- FP8: 2.5x speedup, 50% memory saving, ~0-2% accuracy loss (Ref4.01, Ref4.06).
- 7B on H100: 250-300 tokens/s at batch 8 vs 100-120 for FP16 (2.5-3x), <2% loss on MMLU/GSM8K, 50% memory reduction (Ch7.4).
- 8-16x speedup over FP32 (Ch7.4, Ref7.01); 2-3x faster computation with <2% accuracy loss (Ref7.05).
- In a 70B worked example, FP8 quantization raised throughput from 100 to 120 tokens/s, cut memory from 35 to 18 GB per GPU and halved hourly cost ($10 to $5) (Ref7.15).
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note