Cognition · Model asset
Value Network
Model assetCognitionCognition & Memoryarc:ValueNetwork
Neural network weights trained, typically through self-play, to predict the eventual outcome (e.g., win probability) from an intermediate state.
Responsibility. Predicts outcome value of a state without simulation to terminal.
Also known as: Critic network, Centralized critic, Twin critics, Critic Value Model, Value function model (PPO critic)
Relationships
deployed on structural
is trained by lifecycle
Classification
- Patterns
- Actor-criticCentralized critic over global state and joint actionsTwin critics with minimum (TD3)
- Risks mitigated
- High variance of pure policy-gradient updatesMulti-agent credit assignment errors
Sources
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.