Model Adaptation · Software component

Preference Reward Scorer

Software componentModel AdaptationModelsarc:PreferenceRewardScorer

A reward scorer whose reward is solely the learned reward model's prediction of human preference.

Responsibility. Rewards responses by predicted human preference alone.

Also known as: Traditional RLHF reward

Variant of Reward Scorer abstract

When to choose. Choose when desired behaviour is a subjective judgment (e.g., tone, trade-offs) that rigid rules cannot capture and annotator bias is controlled.

hostsspecializesis target of alternativeTois target of alternativeToalternative tois target of alternativeToReward Model: hostsReward ModelReward Scorer: specializesReward ScorerConstitutional Reward Scorer: is target of alternativeToConstitutional Reward Sc…Ensemble Reward Scorer: is target of alternativeToEnsemble Reward ScorerSegment-Routed Reward Scorer: alternative toSegment-Routed Reward Sc…Composite Reward Scorer: is target of alternativeToComposite Reward Scorer
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

alternative to variability

Design guidance

Sources

  1. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  2. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  3. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.