Model Serving · Software component
Model Quantizer
Software componentModel ServingModelsarc:ModelQuantizer
A model optimisation component that re-represents weights and activations at lower numeric precision to cut memory use and raise throughput without retraining.
Responsibility. Reduces model numeric precision for efficient serving.
Also known as: FP8 quantization, Model representation optimization, Precision reduction, INT8 quantization, SmoothQuant, AWQ, FP8 precision conversion, ModelOptimizer quantization, Post-training quantizer
Relationships
invokes dependency
is invoked by dependency
reads dependency
optimizes lifecycle
produces lifecycle
Design guidance
- MAY apply aggressive quantization only where the application tolerates small (2-3%) accuracy degradation.
- SHOULD be applied where infrastructure cost dominates (on-premise, edge) to gain efficiency unavailable through prompt optimisation.
- SHOULD apply quantization for memory-bandwidth-bound or memory-capacity-bound inference, not when CPU preprocessing limits throughput.
- SHOULD use per-channel or per-token scaling for decoder-only LLMs whose activation ranges vary widely, accepting slight overhead to preserve generation quality.
- MUST measure accuracy impact of every precision reduction against a full-precision baseline before deployment.
- SHOULD try FP8 first and move to 4-bit formats only when accuracy tolerance allows.
- MUST validate quantized quality on the specific domain, since medical or legal tasks may degrade more than general benchmarks.
- SHOULD quantize only operations that tolerate lower precision and keep precision-sensitive attention and GEMM intermediates at FP16.
- SHOULD handle attention activation outliers with per-channel or per-token scaling rather than naive per-tensor scaling.
- MAY extend quantization to KV-cache entries to double concurrent-request capacity within a fixed memory budget.
- SHOULD start from FP16, then test INT8 with calibration, evaluate accuracy impact, and use FP8 where hardware allows (Ref7.01).
- MUST validate the accuracy impact before adopting a reduced precision.
Quantitative guidance
As stated by the sources; verify before use.
- FP8 quantization halves memory versus FP16 and roughly doubles throughput with negligible accuracy loss on instruction-following tasks (Ch2.7).
- FP32 (4 bytes/param) to INT8 (1 byte) or INT4 (0.5 byte) compresses models 4-8x; GPTQ/AWQ achieve 3-4x latency reduction with < 1-2% accuracy loss (Ch3.4).
- INT8 instead of FP32 uses 75% less memory and runs faster (Ch3.10).
- 7B model: 28 GB FP32 -> 7 GB INT8 (4x) -> 3.5 GB 4-bit (8x) -> ~1.75 GB 2-bit (16x); INT8 ops 3-4x faster on CPUs (Ch4.3).
- QAT loses ~1-2% accuracy vs 2-5% for naive PTQ (Ch4.3).
- Worked example INT8 PTQ: 3.7x smaller, 3.4x faster, 3.6x less memory, 2.2% accuracy loss (Ch4.3).
- FP16 -> FP8 halves weight memory 16.8 -> 8.4 GB with typically < 2% accuracy loss on reasoning benchmarks; FP8/INT4 halve or quarter bytes moved per operation (Ch4.2).
- Well-calibrated INT8 preserves 97-99% of FP32 accuracy for encoder-only models like BERT (Ch4.6).
- A 1% token-probability error can cascade into nonsensical output after 50-100 generation steps in decoder-only models (Ch4.6).
- FP8 vs BF16: 1.5-2x throughput, 50% memory reduction, <2% accuracy impact with proper tuning (Ref4.02).
- FP16 to INT8 cuts memory 50-75% (14GB to 7GB) and raises throughput 2-3x (Ch7.2).
- INT8 reduces memory 75% vs FP32 with 4-8x throughput and 0.5-1% typical accuracy loss (Ch7.4, Ref7.01).
- FP8 KV cache shrinks a 7B model's per-request KV cache from 2 GB (FP16) to 1 GB (Ch7.4).
- INT4/INT3 give extreme compression (75% memory reduction) with more accuracy loss and need careful calibration (Ref7.05).
- Quantization yields 50-75% memory reduction for Nemotron models (Ref7.12); INT8 and FP8 available for Llama Nemotron (Ref7.13).
- INT8/FP8 models reduce memory 50-75% with minimal quality loss (Ref7.04).
- Quantization (INT8/INT4) reduces memory 50-75%; mixed precision (FP16) 30-50% (Ref8.05).
Classification
- Patterns
- Post-training quantizationCalibration-based weight quantizationMixed-precision preservation of critical weightsAdaptive mixed-precision inference4-bit base-model quantization for QLoRAPost-training quantization (PTQ)Quantization-aware training (QAT)INT8/INT4/2-bit precisionFP8 weight quantizationINT4 quantizationPer-tensor scalingPer-channel scalingPer-token scalingPer-channel quantizationPer-token quantizationFP8 KV-cache quantizationMixed precision (FP16 attention/GEMM plugins with FP8 weights)
- Technologies
- TensorRT-LLMGPTQAWQNVIDIA NIMNVIDIA TensorRT Model Optimizer (ModelOpt)TensorRT
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Cost efficiencyFunctional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- GPU memory exhaustionAccuracy degradation from naive precision reductionAutoregressive error cascade in decoder-only modelsMemory-bandwidth-bound decode latencyKV-cache memory exhaustionOver-quantization quality loss
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref4.02: E. Potyraj, "Measure and Improve AI Workload Performance with NVIDIA DGX Cloud Benchmarking," NVIDIA Technical Blog, Mar. 18, 2025. [Online]. Available: https://developer.nvidia.com/blog/measure-and-improve-ai-workload-performance-with-nvidia-dgx-cloud-benchmarking/
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.13: NVIDIA, "Llama Nemotron," NVIDIA NeMo Framework User Guide, v25.09. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo-framework/user-guide/25.09/llms/llama_nemotron.html
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note