Model Adaptation · Data artifact
Preference Optimization Config
Data artifactModel AdaptationModelsarc:PreferenceOptimizationConfig
A configuration artifact fixing preference-optimization hyperparameters such as the KL-divergence penalty coefficient, PPO update settings and training duration or early-stopping criteria.
Responsibility. Sets the balance between reward maximization and staying close to the reference model.
Also known as: KL penalty coefficient, RLHF hyperparameters
Relationships
configures structural
- Preference Optimizer abstract Ch10.3
Design guidance
- MUST tune the KL coefficient deliberately: too small lets the policy drift into reward hacking; too large leaves the policy barely changed from the reference model.
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.