Model Adaptation · Model asset
Reward Model
Model assetModel AdaptationModelsarc:RewardModel
A neural network, initialized from a pre-trained language model with a scalar output head, that scores a prompt-response pair by predicted human preference.
Responsibility. Predicts a scalar preference score for a response.
Also known as: Preference model, IRL-inferred reward function, AI-feedback-trained reward model, Fairness reward model, Value model
Relationships
deployed on structural
sends data to dynamic
is evaluated by assurance
is monitored by assurance
is trained by lifecycle
Design guidance
- SHOULD be sized empirically on held-out data rather than assuming bigger is better.
- SHOULD be periodically retrained because frozen reward models drift out of alignment and are exploited by reward hacking.
- SHOULD be treated as an imperfect proxy: combine multiple reward functions and keep human oversight rather than treating any single reward as ground truth.
- MUST be treated as an imperfect, biased proxy for human preferences rather than a ground-truth objective; high held-out accuracy does not guarantee faithful preference capture.
- SHOULD be retrained regularly on production data incorporating newly discovered edge cases so it evolves alongside the policy.
- SHOULD itself be governed (meta-oversight): monitor its outputs for unexpected discrimination or policy violations.
Quantitative guidance
As stated by the sources; verify before use.
- Reward models are typically 10-20% of the policy's scale, e.g., 7-14B for a 70B policy (Ch3.5).
- Larger reward models (10-52B) capture subtler preferences; smaller ones (1-10B) are cheaper but may miss signals (Ch3.5).
- Trained on tens of thousands to hundreds of thousands of pairwise comparisons, each providing a single bit of information (Ch10.3).
Classification
- Patterns
- Scalar reward headReward model ensemblesContext-aware reward modelsBradley-Terry pairwise preference modelReward model ensemblePreference model pretraining
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Reward hackingSpecification gaming (Goodhart's Law)
Sources
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ch9.6: T. Nguyen, "Value Alignment Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.6. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.