Model Serving · Data store

Model Artifact Cache

Data storeModel ServingModelsarc:ModelArtifactCache

A shared store of model weights and adapters mounted by inference replicas so models load locally instead of being re-downloaded by each replica.

Responsibility. Serves model artifacts to replicas without redundant downloads.

Also known as: Model cache volume, Model cache directory

is read byis read byis read bydeployed onLLM Inference Service: is read byLLM Inference ServiceWorker Agent: is read byWorker AgentModel Integrity Validator: is read byModel Integrity ValidatorPersistent Volume: deployed onPersistent Volume
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is read by dependency

Design guidance

Classification

Quality attributes
Performance efficiency (ISO/IEC 25010)Cost efficiency
Risks mitigated
Redundant model downloads

Sources

  1. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  2. Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/