Model Adaptation · Model asset

Reward Model

Model assetModel AdaptationModelsarc:RewardModel

A neural network, initialized from a pre-trained language model with a scalar output head, that scores a prompt-response pair by predicted human preference.

Responsibility. Predicts a scalar preference score for a response.

Also known as: Preference model, IRL-inferred reward function, AI-feedback-trained reward model, Fairness reward model, Value model

is evaluated byis monitored byis evaluated bysends data tois evaluated bydeployed onis trained bydeployed onis trained byis evaluated bydeployed ondeployed ondeployed onEvaluation Harness: is evaluated byEvaluation HarnessOnline Evaluator: is monitored byOnline EvaluatorBias Evaluator: is evaluated byBias EvaluatorRLHF Policy Optimizer: sends data toRLHF Policy OptimizerRed Team Tester: is evaluated byRed Team TesterReward Scorer: deployed onReward ScorerInverse RL Reward Learner: is trained byInverse RL Reward LearnerPreference Reward Scorer: deployed onPreference Reward ScorerReward Model Trainer: is trained byReward Model TrainerAdversarial Robustness Evaluator: is evaluated byAdversarial Robustness E…Constitutional Reward Scorer: deployed onConstitutional Reward Sc…Ensemble Reward Scorer: deployed onEnsemble Reward ScorerSegment-Routed Reward Scorer: deployed onSegment-Routed Reward Sc…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

sends data to dynamic

is evaluated by assurance

is monitored by assurance

is trained by lifecycle

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Scalar reward headReward model ensemblesContext-aware reward modelsBradley-Terry pairwise preference modelReward model ensemblePreference model pretraining
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Reward hackingSpecification gaming (Goodhart's Law)

Sources

  1. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  2. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
  3. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  4. Ch9.6: T. Nguyen, "Value Alignment Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.6. ISBN: 9798244538229.
  5. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
  6. Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.