Infrastructure · Data artifact
Deployment Manifest
Data artifactInfrastructureInfrastructurearc:DeploymentManifest
A declarative workload specification defining container image, labels, resource requests/limits, metrics annotations and health probe timings for replicas.
Responsibility. Declares how replica containers are packaged, resourced and probed.
Also known as: Kubernetes Deployment spec, Probe configuration, Deployment resource, Workload specification, Serverless function configuration (memory, timeout), Kubernetes Deployment for NIM, Docker Compose file
Relationships
configures structural
- API Key Authenticator Ch6.2B
- Agent Controller abstract Ch4.7
- Container Health Prober Ch1.8 Ch4.7
- Event-Triggered Agent Ch4.2
- Inference Server Ch7.1B
- Inference Service Operator Ch4.5
- LLM Inference Service Ch1.8 Ch4.5
- Rollout Manager abstract Ch4.4
- Rolling Update Controller Ch4.3
- Self-Managed Vector Index Store Ch6.2B
- Workload Controller abstract Ch4.3
is configured by structural
is read by dependency
constrains control
Design guidance
- SHOULD set GPU requests equal to limits so pods receive guaranteed resources.
- SHOULD delay liveness checks beyond model load time while starting readiness checks early.
- MUST define both resource requests (for scheduling) and limits (for isolation) for every agent container.
- SHOULD size requests near the 75th and limits near the 95th percentile of utilization measured under realistic staging load, plus safety margin.
- SHOULD validate manifests (dry-run, schema validation) and keep them under version control to enable rollback.
- SHOULD size function memory to the workload, since serverless platforms allocate CPU proportionally to memory (1 GB+ for CPU-intensive initialisation).
- MUST declare resource requests (for scheduling) and limits (to contain leaks or runaway consumption) per replica.
- SHOULD set explicit GPU resource limits and node affinity for GPU workloads.
Quantitative guidance
As stated by the sources; verify before use.
- Liveness initialDelaySeconds 300; readiness initialDelaySeconds 30 (Ch1.8).
- Over-provisioned clusters observed at ~15% utilization; 8 GB requests for a 500 MB agent multiply cost ~10x (Ch4.3).
- Example coordinator: requests 0.5 CPU/1 GiB, limits 1 CPU/2 GiB; GPU worker: requests 2 CPU/8 GiB/1 GPU, limits 4 CPU/16 GiB (Ch4.3).
- Worked example memory: 1024 MB session analyser, 2048 MB pattern matcher (large indices), 3008 MB recommendation generator (Ch4.2).
- A function with 3 GB memory initialises faster than the same code at 512 MB (Ch4.2).
- Example: 4 replicas of Llama 3.1 70B, 2 A100 80GB GPUs each; liveness probe initial delay 120s/period 30s, readiness 90s/10s (Ch4.5).
- Worked manifest: requests 2 CPU / 8 GiB, limits 4 CPU / 16 GiB; liveness /health initialDelay 30 s period 10 s; readiness /ready initialDelay 10 s period 5 s (Ch4.7).
Classification
- Patterns
- Requests equal limits (guaranteed QoS)Separate liveness/readiness probe timingInfrastructure as codeResource requests and limits
- Technologies
- Kubernetes DeploymentConfigMapSecretServiceIngressHorizontalPodAutoscalerPodDisruptionBudgetNetworkPolicyDocker Compose
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Noisy neighboursPods killed during model loadingScheduling failuresOut-of-memory kills from uninformed schedulingNoisy-neighbour resource starvationOver-provisioning cost waste
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html