Model Adaptation · Software component
Ensemble Reward Scorer
Software componentModel AdaptationModelsarc:EnsembleRewardScorer
A reward scorer that combines scores from multiple independent reward models trained on different data or with different architectures into one reward signal.
Responsibility. Produces a reward that a policy can only exploit by fooling several reward models simultaneously.
Also known as: Ensemble reward models
Variant of Reward Scorer abstract
When to choose. Choose when reward hacking via single-pattern exploitation is a concern and the cost of training multiple reward models is acceptable.
Relationships
hosts structural
alternative to variability
Design guidance
- SHOULD NOT be assumed to eliminate reward hacking; pair with behavioural monitoring and adversarial testing.
Classification
- Risks mitigated
- Reward hackingSpecification gaming
Sources
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.