Model Adaptation · Software component

Imitation Learner

Software componentModel AdaptationModelsVariation point (abstract)arc:ImitationLearner

An abstract policy learner that derives a policy from expert demonstrations rather than autonomous trial-and-error exploration.

Responsibility. Learns a policy from expert demonstrations.

Also known as: Learning from demonstration

Variant of Policy Learner abstract

When to choose. Choose when expert demonstrations are available and autonomous exploration is too slow, costly, or unsafe.

alternative tois specialized byis specialized byis specialized byspecializesReinforcement Learning Policy Learner: alternative toReinforcement Learning P…DAgger Trainer: is specialized byDAgger TrainerInverse RL Reward Learner: is specialized byInverse RL Reward LearnerBehavior Cloning Trainer: is specialized byBehavior Cloning TrainerPolicy Learner: specializesPolicy Learner
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Behavior Cloning TrainerChoose when demonstrations comprehensively cover deployment situations, the expert is consistent, and decision sequences are short enough that errors do not compound.
DAgger TrainerChoose when experts can provide repeated labels and the learned policy can be safely executed during training to expose its failure modes.
Inverse RL Reward LearnerChoose when demonstrations come from multiple experts with different strategies, deployment differs qualitatively from demonstrations, or the reward is needed for explanation or policy evaluation.

Relationships

alternative to variability

Design guidance

Sources

  1. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.