Model Serving · Model asset
FP4 Inference Engine
Model assetModel ServingModelsarc:FP4InferenceEngine
An optimized inference engine using 4-bit floating-point representations for maximum compression.
Responsibility. Serves a model at 4-bit floating-point precision.
Also known as: FP4 quantized engine
Variant of Optimized Inference Engine abstract
When to choose. Choose for maximum compression and throughput in use cases that accept ~3-7% accuracy loss.
Relationships
alternative to variability
Quantitative guidance
As stated by the sources; verify before use.
- FP4: 4x speedup, 75% memory saving, ~3-7% accuracy loss (Ref4.01, Ref4.06).
Sources
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM