Model Serving · Data artifact
Model Deployment Profile
Data artifactModel ServingModelsVariation point (abstract)arc:ModelDeploymentProfile
A pre-validated, hardware-specific bundle of serving decisions for one model (precision format, multi-GPU layout, batch sizes) selected when an inference microservice is deployed.
Responsibility. Encapsulates the quantization, parallelism and batching choices that fit a given GPU target and latency-throughput objective.
Also known as: NIM profile, Optimization profile, NIM_MODEL_PROFILE
Variant of Inference Serving Configuration abstract
Variants
| Variant | When to choose |
|---|---|
| Latency-Optimized Deployment Profile | Choose for latency-critical interactive applications where time-to-first-token SLOs dominate. |
| Memory-Optimized Deployment Profile | Choose when available GPU memory is insufficient for full-precision weights plus activations, such as smaller GPUs or edge devices. |
| Throughput-Optimized Deployment Profile | Choose when maximising queries per second matters more than minimum time-to-first-token and multiple GPUs are available. |
Relationships
configures structural
Design guidance
- MUST select a profile compatible with the target GPU type, GPU count and quantization support rather than relying on the default profile.
- SHOULD treat profile selection as a latency-throughput-memory trade-off decision recorded per deployment.
- SHOULD be selected automatically from the detected GPU configuration so no manual serving configuration is needed.
Classification
- Technologies
- NVIDIA NIM
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Deployment failure when default profile exceeds available GPU resourcesSeverely suboptimal performance when a conservative profile runs on powerful hardware