Infrastructure · Infrastructure resource
Container Orchestrator
Infrastructure resourceInfrastructureInfrastructurearc:ContainerOrchestrator
A cluster platform that schedules containers across nodes, maintains desired replica counts and restarts or replaces failed pods.
Responsibility. Schedules and maintains containerized replicas across cluster nodes.
Also known as: Kubernetes cluster, Microservices deployment platform, Kubernetes
Variant of Agent Hosting Platform abstract
When to choose. Choose when consistent latency matters more than cost (24/7 warm instances, no cold starts), Kubernetes already runs other services to amortise overhead, the team has strong Kubernetes expertise, continuous background processing is needed, or debugging requires persistent logs and long-term metric retention.
Relationships
hosts structural
- Accelerator Operator Ref4.07
- Agent Controller abstract Ch1.8 Ch4.1 +1
- Answer Synthesizer abstract Ch4.2
- Autoscaler abstract Ch1.8 Ch4.7
- Container Health Prober Ch1.8 Ch4.2
- Distributed Vector Index Store Ch6.2A
- Escalation Handler Ch4.2
- Event Broker abstract Ch4.3
- GitOps Reconciler Ch4.4
- Gossip Membership Service Ch6.2B
- Guardrail Orchestrator Ref7.03
- High-Performance Reverse Proxy Ch4.1
- Inference Server Ch2.7 Ch4.5 +7
- Inference Service Operator Ch4.5
- Intent Router Ch4.2
- Knowledge Retrieval Agent Ch4.2
- Layer-4 Load Balancer Ch4.7
- Layer-7 Load Balancer Ch4.2
- LLM Inference Service Ch1.8 Ch2.6 +3
- Load Balancer abstract Ch1.8
- Metric-Driven Autoscaler Ch4.2 Ch4.3 +1
- Metrics Collector Ref4.05
- Metrics Dashboard abstract Ref4.05
- Rollout Manager abstract Ch4.1 Ch4.4
- RAG Query Orchestrator Ch6.5
- Rolling Update Controller Ch4.2 Ch4.3
- Self-Managed Vector Index Store Ch6.2B
- Service Mesh Proxy Ch4.3
- Shard Query Router Ch6.2B
- Shard Replication Manager Ch6.2B
- Workload Controller abstract Ch4.3
receives data from dynamic
is constrained by control
is monitored by assurance
alternative to variability
Design guidance
- SHOULD start with a single cluster, fixed replicas, basic resource limits and standard services, adding autoscaling, service mesh and multi-region only when operational pain or requirements justify them.
- SHOULD spread replicas of a service across multiple physical nodes so a node failure does not eliminate all capacity.
- SHOULD NOT adopt microservices unless components have fundamentally different scaling requirements, multiple teams must deploy independently, or traffic volume justifies the operational overhead.
- MUST avoid a distributed monolith: services sharing database tables, synchronous call chains or coordinated deployments are not independently deployable.
- SHOULD consolidate over-fragmented services back into larger cohesive services when coordination overhead exceeds capability work.
Quantitative guidance
As stated by the sources; verify before use.
- Worked example: 3 x t3.medium (2 vCPU, 4 GB, $0.096/h) at peak, 2 off-peak = $172.80/month + $50 LB/monitoring = $222.80/month; managed EKS/GKE adds $100-200/month (Ch4.2).
- Peak sizing: 8,750 sessions/h = 2.4 sessions/s x 5 s = 12 concurrent threads -> 3 instances x 4 threads (Ch4.2).
- Container rebuild + Kubernetes deployment adds 5-8 minutes per deployment cycle; instances idle ~50% of off-peak time (Ch4.2).
Classification
- Patterns
- Declarative desired stateReconciliation loopSelf-healingMicroservicesMulti-instance redundancy across physical nodesRolling updateIndependent per-service scalingRolling updates (zero-downtime)Resource requests/limits schedulingGPU resource requests (nvidia.com/gpu)Node affinity for GPU nodesNamespace resource quotasRestart policy unless-stopped
- Technologies
- KubernetesDockerNVIDIA NGCAmazon EKSGoogle GKEcontainerdCRI-O
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Security (ISO/IEC 25010 | NIST AI RMF: secure and resilient)
- Risks mitigated
- Single node hardware failure eliminating all capacitySlow manual response to production failures
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.6: T. Nguyen, "Tool Integration and Function Calling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.6. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch6.2A: T. Nguyen, "Vector Database Selection," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2A. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ref4.07: NVIDIA, "About the NVIDIA GPU Operator," NVIDIA GPU Operator Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html
- Ref4.08: NVIDIA, "Agentic AI in the Factory," NVIDIA Enterprise AI Factory Design Guide White Paper. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/ai-enterprise/planning-resource/ai-factory-white-paper/latest/agentic-ai-in-the-factory.html
- Ref7.03: NVIDIA, "Overview," NVIDIA NeMo Guardrails Library Developer Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.12: "Advanced Nemotron Deployment Patterns," unpublished reference note (12-Nemotron-Advanced-Deployment.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.14: "NVIDIA Agentic AI Platform Ecosystem Integration," unpublished reference note (14-NVIDIA-Ecosystem-Integration.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note