Model Serving · Model asset
INT4 Quantized Engine
Model assetModel ServingModelsarc:INT4InferenceEngine
A compiled model engine with 4-bit weights, optionally using activation-aware or mixed-precision schemes that keep influential weights at higher precision.
Responsibility. Serves the model at 4-bit weight precision for maximum efficiency.
Also known as: INT4 AWQ engine
Variant of Optimized Inference Engine abstract
When to choose. Choose only when 2-5% accuracy loss is acceptable and maximum memory and cost reduction is needed.
Relationships
alternative to variability
Design guidance
- SHOULD be reserved for cost-sensitive applications; avoid aggressive quantization when response quality is critical.
Quantitative guidance
As stated by the sources; verify before use.
- About 8x memory reduction with 3-5% accuracy loss; AWQ/GPTQ achieve 4-6x memory reduction with ~2-3% degradation (Ch4.4).
- Moderation example: INT4 fell to 94.4% (-2.4%), below the 95% threshold, and was rejected (Ch4.4).
- INT4 AWQ: 3-4x speedup, 75% memory saving, ~2-5% accuracy loss (Ref4.01, Ref4.06).
- INT4/INT3 give 75% memory reduction with more accuracy loss than FP8 (Ref7.05).
Classification
- Patterns
- Activation-aware weight quantization (AWQ)GPTQMixed-precision quantization
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref7.05: S. Verma and N. Vaidya, "Mastering LLM Techniques: Inference Optimization," NVIDIA Technical Blog, Nov. 17, 2023. [Online]. Available: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note