Cognition · Data artifact

Reward Function Specification

Data artifactCognitionCognition & Memoryarc:RewardFunctionSpecification

A specification of the immediate payoff received on each state transition, together with the discount factor, from which long-term utility is derived.

Responsibility. Defines the reward signal that simulations produce and search optimizes.

Also known as: Terminal reward signal, Reward shaping, Reward function, Reward structure, Shaped reward, Potential function

configures; is read byis written by; is produced byconfiguresis configured byconfiguresconfiguresconfiguresconfiguresis read byReinforcement Learning Policy Learner: configures; is read byReinforcement Learning P…Inverse Reward Learner: is written by; is produced byInverse Reward LearnerEnvironment Simulator: configuresEnvironment SimulatorOperational Norm Set: is configured byOperational Norm SetMonte Carlo Planner: configuresMonte Carlo PlannerRollout Simulator: configuresRollout SimulatorMDP Policy Solver: configuresMDP Policy SolverPOMDP Policy Solver: configuresPOMDP Policy SolverReward Shaper: is read byReward Shaper
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

is configured by structural

is read by dependency

is written by dependency

is produced by lifecycle

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Binary rewardsScalar rewardsVector (multi-objective) rewards with scalarizationWeighted-sum scalarizationPareto optimizationDiscounted rewardsReward engineeringPotential-based reward shaping (gamma*Phi(s') - Phi(s))Sparse vs. dense rewardConstraint-violation penalties
Risks mitigated
Reward misalignment / perverse incentives from shapingReward hackingSparse-reward learning failureNaive shaping changing the optimal policy

Sources

  1. Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
  2. Ch5.10: T. Nguyen, "Utility-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.10. ISBN: 9798244538229.
  3. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.
  4. Ch5.13: T. Nguyen, "Hybrid Decision Systems Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.13. ISBN: 9798244538229.