Model Adaptation · Software component

Reward Scorer

Software componentModel AdaptationModelsVariation point (abstract)arc:RewardScorer

An abstract component that computes the scalar reward signal for a prompt-response pair used to guide reinforcement-learning policy optimization.

Responsibility. Supplies reward signals to policy optimization.

Also known as: Reward function, Reward signal source

is invoked byhostsis specialized byis specialized byis specialized byis specialized byis specialized byis specialized byRLHF Policy Optimizer: is invoked byRLHF Policy OptimizerReward Model: hostsReward ModelPreference Reward Scorer: is specialized byPreference Reward ScorerConstitutional Reward Scorer: is specialized byConstitutional Reward Sc…Ensemble Reward Scorer: is specialized byEnsemble Reward ScorerSegment-Routed Reward Scorer: is specialized bySegment-Routed Reward Sc…Composite Reward Scorer: is specialized byComposite Reward ScorerFairness-Constrained Reward Scorer: is specialized byFairness-Constrained Rew…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Composite Reward ScorerChoose when human preference must be grounded by verifiable correctness, e.g., to prevent optimizing toward fluent hallucinations or annotator bias.
Constitutional Reward ScorerChoose when harmlessness rewards must scale without human labels and reflect explicit, inspectable principles rather than implicit annotator preferences.
Ensemble Reward ScorerChoose when reward hacking via single-pattern exploitation is a concern and the cost of training multiple reward models is acceptable.
Fairness-Constrained Reward ScorerChoose for high-stakes decisions such as lending where the policy must satisfy a chosen fairness metric (e.g., equalized odds) while keeping legitimate business criteria.
Preference Reward ScorerChoose when desired behaviour is a subjective judgment (e.g., tone, trade-offs) that rigid rules cannot capture and annotator bias is controlled.
Segment-Routed Reward ScorerChoose when user segments legitimately hold divergent preferences and sufficient preference data can be collected for each segment.

Relationships

hosts structural

is invoked by dependency

Design guidance

Classification

Patterns
Hybrid reward functions
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Reward hackingPrinciple gaming (safety and helpfulness principle hacking)

Sources

  1. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  2. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  3. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.