Model Adaptation · Software component
Preference Optimizer
Software componentModel AdaptationModelsVariation point (abstract)arc:PreferenceOptimizer
An abstract model-adaptation component that updates a supervised-fine-tuned policy model so its outputs better match human preferences expressed as comparative judgments.
Responsibility. Aligns a policy model to learned preferences.
Also known as: Preference alignment, RLHF/DPO alignment training
Variants
| Variant | When to choose |
|---|---|
| Direct Preference Optimizer | Choose first for simpler preference-learning tasks where direct optimization suffices and lower compute is desired. |
| RLHF Policy Optimizer | Choose for complex multi-dimensional preferences where explicit reward modeling provides interpretability benefits. |
Relationships
is configured by structural
is triggered by dynamic
- Fine-Tuning Pipeline abstract Ch3.5 Ch10.3
is orchestrated by control
trains lifecycle
Design guidance
- SHOULD be applied only after supervised fine-tuning establishes baseline task competence.
- SHOULD NOT be assumed to solve value alignment comprehensively; it aligns only to annotators' expressed preferences.
- SHOULD be combined with constitutional principles, content filtering, human oversight and continuous monitoring rather than used as a standalone alignment mechanism.
- SHOULD be customized per domain (preference elicitation, annotator expertise, evaluation metrics); language-model RLHF techniques do not transfer directly to robotics or other domains.
Quantitative guidance
As stated by the sources; verify before use.
- RLHF-tuned customer-service agents reported 35-45% higher customer satisfaction than accuracy- or speed-optimized agents (Ch3.5).
- Preference-aligned recommendation reported 15-25% user retention improvements over engagement-only optimization (Ch3.5).
Classification
- Patterns
- Reinforcement learning from human feedbackPreference learningRLHF three-phase pipeline (SFT -> reward modeling -> RL)Preference-based alignment
- Technologies
- NVIDIA NeMo Customizer
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Behaviour that resists explicit demonstration or rule specificationReward specification intractability
Sources
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.