Model Adaptation · Software component
Reward Shaper
Software componentModel AdaptationModelsarc:RewardShaper
A training-time component that adds intermediate shaping rewards to a sparse reward signal to accelerate learning while preserving the optimal policy.
Responsibility. Densifies sparse rewards with potential-based shaping terms during learning.
Also known as: Reward shaping
Relationships
reads dependency
sends data to dynamic
Design guidance
- SHOULD use potential-based shaping so the optimal policy under shaped rewards equals that under the original rewards.
Classification
- Patterns
- Potential-based reward shaping F(s,a,s') = gamma*Phi(s') - Phi(s)
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Slow learning in sparse-reward environmentsReward hacking of shaped rewards
Sources
- Ch5.10: T. Nguyen, "Utility-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.10. ISBN: 9798244538229.