Model Adaptation · Software component

Policy Learner

Software componentModel AdaptationModelsVariation point (abstract)arc:PolicyLearner

An abstract model-adaptation component that updates a learned policy model from interaction experience or expert demonstrations.

Responsibility. Updates the learned policy model from experience or demonstrations.

Also known as: Learning algorithm, Policy trainer

is specialized byis specialized byis specialized bytrainsReinforcement Learning Policy Learner: is specialized byReinforcement Learning P…Multi-Agent Policy Learner: is specialized byMulti-Agent Policy LearnerImitation Learner: is specialized byImitation LearnerLearned Decision Policy: trainsLearned Decision Policy
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Imitation Learner abstractChoose when expert demonstrations are available and autonomous exploration is too slow, costly, or unsafe.
Multi-Agent Policy Learner abstract—
Reinforcement Learning Policy LearnerChoose when a reward signal is available and the agent can safely explore (e.g., in simulation), with ample data and compute.

Relationships

trains lifecycle

Design guidance

Classification

Patterns
Offline trainingOnline adaptationTransfer learning from simulation to real world

Sources

  1. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.