Model Serving · Software component

Quantization Calibrator

Software componentModel ServingModelsarc:QuantizationCalibrator

A model optimisation component that runs representative inputs through a network to collect per-tensor activation statistics and computes quantization thresholds (scaling factors) minimising information loss.

Responsibility. Derives the scaling factors used to map full-precision activations to low-precision integers.

Also known as: Entropy calibrator, INT8 calibration, Entropy calibration, Calibration step

is invoked byreadsModel Quantizer: is invoked byModel QuantizerQuantization Calibration Dataset: readsQuantization Calibration…
Direct neighbourhood (hover for relationship types)

Relationships

is invoked by dependency

reads dependency

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Entropy calibration (KL-divergence minimisation)Entropy (KL-divergence) calibrationMin-max calibrationPercentile calibrationE4M3 FP8 calibration
Technologies
TensorRT-LLMTensorRTNVIDIA TensorRT Model Optimizer (ModelOpt)
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Accuracy loss from mis-set quantization thresholdsINT8 slower than FP16 from poor quantize-dequantize placement'No scaling factors detected' calibration failuresSuboptimal scaling factors amplifying quantization error

Sources

  1. Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
  2. Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.