Cognition · Data artifact
Reward Function Specification
Data artifactCognitionCognition & Memoryarc:RewardFunctionSpecification
A specification of the immediate payoff received on each state transition, together with the discount factor, from which long-term utility is derived.
Responsibility. Defines the reward signal that simulations produce and search optimizes.
Also known as: Terminal reward signal, Reward shaping, Reward function, Reward structure, Shaped reward, Potential function
Relationships
configures structural
is configured by structural
is read by dependency
is written by dependency
is produced by lifecycle
Design guidance
- SHOULD directly reflect task objectives (e.g., +100 at goal, 0 otherwise) rather than intermediate progress.
- If shaping is necessary, SHOULD reward progress only when strictly closer to the goal and cap total shaped reward.
- SHOULD use potential-based shaping so dense guidance does not change which policy is optimal.
- SHOULD include large penalties for actions the safety rules would override, so the learned policy internalizes constraints.
Quantitative guidance
As stated by the sources; verify before use.
- Grid world reward: +100 goal, -100 trap, -1 per move (Ch5.12).
Classification
- Patterns
- Binary rewardsScalar rewardsVector (multi-objective) rewards with scalarizationWeighted-sum scalarizationPareto optimizationDiscounted rewardsReward engineeringPotential-based reward shaping (gamma*Phi(s') - Phi(s))Sparse vs. dense rewardConstraint-violation penalties
- Risks mitigated
- Reward misalignment / perverse incentives from shapingReward hackingSparse-reward learning failureNaive shaping changing the optimal policy
Sources
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch5.10: T. Nguyen, "Utility-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.10. ISBN: 9798244538229.
- Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
- Ch5.13: T. Nguyen, "Hybrid Decision Systems Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.13. ISBN: 9798244538229.