Infrastructure · Data artifact

Deployment Manifest

Data artifactInfrastructureInfrastructurearc:DeploymentManifest

A declarative workload specification defining container image, labels, resource requests/limits, metrics annotations and health probe timings for replicas.

Responsibility. Declares how replica containers are packaged, resourced and probed.

Also known as: Kubernetes Deployment spec, Probe configuration, Deployment resource, Workload specification, Serverless function configuration (memory, timeout), Kubernetes Deployment for NIM, Docker Compose file

configures; constrainsconfigures; constrainsconfiguresconfiguresconfiguresconfiguresconfiguresis read byconfiguresconfiguresconfiguresconfiguresis configured byAgent Controller: configures; constrainsAgent ControllerSelf-Managed Vector Index Store: configures; constrainsSelf-Managed Vector Inde…LLM Inference Service: configuresLLM Inference ServiceInference Server: configuresInference ServerRollout Manager: configuresRollout ManagerEvent-Triggered Agent: configuresEvent-Triggered AgentContainer Health Prober: configuresContainer Health ProberGitOps Reconciler: is read byGitOps ReconcilerRolling Update Controller: configuresRolling Update ControllerWorkload Controller: configuresWorkload ControllerInference Service Operator: configuresInference Service OperatorAPI Key Authenticator: configuresAPI Key AuthenticatorEnvironment Overlay: is configured byEnvironment Overlay
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

is configured by structural

is read by dependency

constrains control

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Requests equal limits (guaranteed QoS)Separate liveness/readiness probe timingInfrastructure as codeResource requests and limits
Technologies
Kubernetes DeploymentConfigMapSecretServiceIngressHorizontalPodAutoscalerPodDisruptionBudgetNetworkPolicyDocker Compose
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)
Risks mitigated
Noisy neighboursPods killed during model loadingScheduling failuresOut-of-memory kills from uninformed schedulingNoisy-neighbour resource starvationOver-provisioning cost waste

Sources

  1. Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
  2. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  3. Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
  4. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  5. Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
  6. Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
  7. Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
  8. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  9. Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html