Model Serving · Data artifact
Memory-Optimized Deployment Profile
Data artifactModel ServingModelsarc:MemoryOptimizedDeploymentProfile
A quantized model deployment profile that reduces memory footprint so a model fits smaller GPUs or edge devices, possibly at some cost in throughput and latency.
Responsibility. Configures an inference microservice to fit constrained GPU memory.
Also known as: Quantized profile, Edge profile
Variant of Model Deployment Profile abstract
When to choose. Choose when available GPU memory is insufficient for full-precision weights plus activations, such as smaller GPUs or edge devices.
Relationships
alternative to variability
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.