Model Serving · Model asset
INT8 Quantized Engine
Model assetModel ServingModelsarc:INT8InferenceEngine
A compiled model engine with 8-bit integer weights and activations whose quantization parameters come from calibration on representative inputs.
Responsibility. Serves the model at calibrated INT8 precision.
Also known as: INT8 calibrated engine, INT8 SmoothQuant engine
Variant of Optimized Inference Engine abstract
When to choose. Choose when 1-2% accuracy reduction is tolerable (e.g., conversational or human-reviewed moderation workloads) in exchange for roughly halved infrastructure cost; requires representative calibration.
Relationships
deployed on structural
is configured by structural
is evaluated by assurance
is produced by lifecycle
alternative to variability
Design guidance
- SHOULD calibrate on data statistically similar to production inputs; MAY maintain separate calibrated builds per workload domain.
Quantitative guidance
As stated by the sources; verify before use.
- Well-calibrated INT8 typically within 1-2% of FP16 on MMLU (Ch4.4).
- Moderation example: 2 GPUs, 324 req/s (162.4 req/s/GPU, 3.59x per-GPU), 95.9% accuracy (-0.9%), $12,000/month, $288,000 annual savings (Ch4.4).
- GPT-2: 11ms/token, 2.2GB (68% reduction), 91 tok/s (3.8x), 98.5% of FP32 accuracy on held-out set, 8-12 concurrent users (Ch4.6).
- INT8 SmoothQuant: 2x throughput, 50% memory saving, ~1-3% accuracy loss (Ref4.06).
- INT8: 2x batch size + 20% kernel speedup = 2.4x throughput, 50% cost savings, typically <2% accuracy loss; 20-30% latency improvement (Ch7.2).
- 7B on A100: 160-180 tokens/s at batch 8 vs 100 for FP16 (60-80% gain), 40-50% lower per-token latency (Ch7.4).
- MMLU 68.2% (FP16) vs 67.5% (INT8): <1 point degradation when well calibrated (Ch7.4).
- Throughput +60% vs FP16, <1% accuracy loss, 75% memory reduction (Ref7.01).
Classification
- Patterns
- Post-training calibrationPer-domain calibrated builds
- Risks mitigated
- Excess GPU cost
Sources
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref4.01: NVIDIA, "TensorRT-LLM," GitHub repository. Accessed: Sep. 27, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note