Infrastructure · Software component
Cluster Consensus Coordinator
Software componentInfrastructureInfrastructurearc:ClusterConsensusCoordinator
A replicated coordination component that maintains consistent cluster state by majority quorum and elects a leader node, redistributing workloads when the leader or a node fails.
Responsibility. Provides quorum-based leader election and automatic failover for a small on-site cluster.
Also known as: Edge high-availability controller, Leader election
Relationships
deployed on structural
receives data from dynamic
orchestrates control
Design guidance
- MUST deploy at least three nodes per site for automatic failover, since a two-node cluster loses quorum on any single failure.
Quantitative guidance
As stated by the sources; verify before use.
- Leader re-election typically within 2-5 seconds; HA flagship stores achieve 99.9%+ uptime vs 99.5%+ non-HA (Ch4.6).
Classification
- Patterns
- Raft consensusLeader electionController-plus-worker nodes
- Technologies
- NVIDIA Fleet Commandetcd
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Loss of automatic failover in two-node configurations lacking quorumSite outage from single hardware failure
Sources
- Ch4.6: T. Nguyen, "TensorRT-LLM and NVIDIA Fleet Command," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.6. ISBN: 9798244538229.
- Ch6.2A: T. Nguyen, "Vector Database Selection," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2A. ISBN: 9798244538229.