Model Serving · Software component
Dynamic Model Loader
Software componentModel ServingModelsarc:DynamicModelLoader
An inference-server component that loads, unloads and switches model versions from a model repository at runtime without restarting the server or disrupting in-flight requests.
Responsibility. Applies model version changes to a running inference server.
Also known as: Triton dynamic model loading, Model version policy
Relationships
deployed on structural
is configured by structural
reads dependency
is guarded by control
Classification
- Technologies
- NVIDIA Triton Inference Server
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Downtime for model updates
Sources
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.