Model Adaptation · Software component

Policy/Value Network Trainer

Software componentModel AdaptationModelsarc:PolicyValueNetworkTrainer

A training component that fits a policy network to expert demonstrations and a value network to self-play game outcomes for use in neural-guided search.

Responsibility. Trains the neural networks that guide tree search.

Also known as: Self-play trainer

deployed oninvokestrainstrainsreadsGPU Node: deployed onGPU NodeEnvironment Simulator: invokesEnvironment SimulatorPolicy Network: trainsPolicy NetworkValue Network: trainsValue NetworkExpert Demonstration Dataset: readsExpert Demonstration Dat…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

invokes dependency

reads dependency

trains lifecycle

Classification

Patterns
Supervised learning from expert gamesSelf-play

Sources

  1. Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.