Model Serving · Data store

Model Repository

Data storeModel ServingModelsarc:ModelRepository

A versioned directory store of model artifacts and their serving configurations, mounted by an inference server from persistent or object storage.

Responsibility. Holds deployable model versions separately from inference server code.

Also known as: Triton model repository, Shared model-weight storage, NGC model catalog

is read byis read byis written byis read byis read byis read byLLM Inference Service: is read byLLM Inference ServiceInference Server: is read byInference ServerEngine Builder: is written byEngine BuilderModel Integrity Validator: is read byModel Integrity ValidatorDynamic Model Loader: is read byDynamic Model LoaderInference Engine Selector: is read byInference Engine Selector
Direct neighbourhood (hover for relationship types)

Relationships

is read by dependency

is written by dependency

Classification

Technologies
Kubernetes PersistentVolumeAmazon S3Google Cloud StorageKubernetes ConfigMap
Quality attributes
Maintainability (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Downtime for model updates

Sources

  1. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  2. Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
  3. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  4. Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
  5. Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
  6. Ref7.08: NVIDIA, "NVIDIA Deep Learning Triton Inference Server Documentation," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/