Model Serving · Model asset

FP4 Inference Engine

Model assetModel ServingModelsarc:FP4InferenceEngine

An optimized inference engine using 4-bit floating-point representations for maximum compression.

Responsibility. Serves a model at 4-bit floating-point precision.

Also known as: FP4 quantized engine

Variant of Optimized Inference Engine abstract

When to choose. Choose for maximum compression and throughput in use cases that accept ~3-7% accuracy loss.

specializesalternative toalternative tois target of alternativeToalternative toOptimized Inference Engine: specializesOptimized Inference EngineINT8 Quantized Engine: alternative toINT8 Quantized EngineFP8 Quantized Engine: alternative toFP8 Quantized EngineFP16 Inference Engine: is target of alternativeToFP16 Inference EngineINT4 Quantized Engine: alternative toINT4 Quantized Engine
Direct neighbourhood (hover for relationship types)

Relationships

alternative to variability

Quantitative guidance

As stated by the sources; verify before use.

Sources

  1. Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM