Model Serving · Model asset
Quantized Model Checkpoint
Model assetModel ServingModelsarc:QuantizedModelCheckpoint
A framework-neutral model graph whose weights are stored at reduced integer or floating-point precision together with embedded quantization metadata (scaling factors, zero points), not yet compiled for a target GPU.
Responsibility. Holds the calibrated low-precision model and its quantization parameters as input for engine compilation.
Also known as: Quantized ONNX, INT8 ONNX model
Relationships
is optimized by lifecycle
is produced by lifecycle
Quantitative guidance
As stated by the sources; verify before use.
- A 14 GB FP16 7B model becomes approximately 3.5 GB in INT8 (4x compression), speeding model loading and cutting storage cost (Ch7.4).
Classification
- Technologies
- NVIDIA TensorRT Model Optimizer (ModelOpt)ONNX
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
Sources
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html