Cognition · Model asset

Value Network

Model assetCognitionCognition & Memoryarc:ValueNetwork

Neural network weights trained, typically through self-play, to predict the eventual outcome (e.g., win probability) from an intermediate state.

Responsibility. Predicts outcome value of a state without simulation to terminal.

Also known as: Critic network, Centralized critic, Twin critics, Critic Value Model, Value function model (PPO critic)

deployed onis trained byis trained byis trained bydeployed onRLHF Policy Optimizer: deployed onRLHF Policy OptimizerReinforcement Learning Policy Learner: is trained byReinforcement Learning P…Policy/Value Network Trainer: is trained byPolicy/Value Network Tra…Centralized-Training Decentralized-Execution Learner: is trained byCentralized-Training Dec…Value Network Evaluator: deployed onValue Network Evaluator
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

is trained by lifecycle

Classification

Patterns
Actor-criticCentralized critic over global state and joint actionsTwin critics with minimum (TD3)
Risks mitigated
High variance of pure policy-gradient updatesMulti-agent credit assignment errors

Sources

  1. Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
  2. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
  3. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.