Model Adaptation · Software component
Imitation Learner
Software componentModel AdaptationModelsVariation point (abstract)arc:ImitationLearner
An abstract policy learner that derives a policy from expert demonstrations rather than autonomous trial-and-error exploration.
Responsibility. Learns a policy from expert demonstrations.
Also known as: Learning from demonstration
Variant of Policy Learner abstract
When to choose. Choose when expert demonstrations are available and autonomous exploration is too slow, costly, or unsafe.
Variants
| Variant | When to choose |
|---|---|
| Behavior Cloning Trainer | Choose when demonstrations comprehensively cover deployment situations, the expert is consistent, and decision sequences are short enough that errors do not compound. |
| DAgger Trainer | Choose when experts can provide repeated labels and the learned policy can be safely executed during training to expose its failure modes. |
| Inverse RL Reward Learner | Choose when demonstrations come from multiple experts with different strategies, deployment differs qualitatively from demonstrations, or the reward is needed for explanation or policy evaluation. |
Relationships
alternative to variability
Design guidance
- SHOULD be used to initialize policies with competent behaviour before reinforcement fine-tuning where exploration is costly or unsafe.
Sources
- Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.