Infrastructure · Software component
Container Health Prober
Software componentInfrastructureInfrastructurearc:ContainerHealthProber
A node-level agent that periodically probes replica liveness and readiness endpoints to trigger restarts or removal from service rotation.
Responsibility. Probes replica health endpoints to detect dead or unready instances.
Also known as: Kubelet probes, Health checker, Kubernetes liveness/readiness probes, Application health check with restart
Relationships
deployed on structural
is configured by structural
invokes dependency
writes dependency
sends data to dynamic
triggers dynamic
- Workload Controller abstract Ch4.3
overrides control
monitors assurance
Design guidance
- SHOULD tune period and failure thresholds to balance detection latency against false positives.
- SHOULD use application-level checks (e.g., a small inference) where process liveness does not imply serving ability.
- SHOULD set liveness initial delays long enough for model loading to avoid spurious restarts.
- SHOULD restart containers failing health checks and remove unhealthy instances from the load-balancer pool until they recover.
- SHOULD admit replicas to traffic only after they are fully initialized and responsive.
Quantitative guidance
As stated by the sources; verify before use.
- 5 s / 2-failure probes detect in ~10 s; 30 s / 3-failure in ~90 s (Ch1.8).
- Example: liveness every 10 s with 3 failures triggering restart; 60 s initial delay for an 8B-parameter model worker (Ch4.3).
- Example probes: liveness on /v1/health (30 s initial delay, 10 s period); readiness on /v1/models (10 s initial delay, 5 s period) (Ref7.04).
Classification
- Patterns
- Liveness probeReadiness probeTCP / HTTP / application-level health checks
- Technologies
- Kubernetes kubelet probes
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Restart thrashing during model loadTraffic to failed instances
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ref4.08: NVIDIA, "Agentic AI in the Factory," NVIDIA Enterprise AI Factory Design Guide White Paper. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/ai-enterprise/planning-resource/ai-factory-white-paper/latest/agentic-ai-in-the-factory.html
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note