Cognition · Model asset

Learned Decision Policy

Model assetCognitionCognition & MemoryVariation point (abstract)arc:LearnedPolicyModel

Learned parameters mapping states to actions, produced by reinforcement learning to maximize expected long-term reward.

Responsibility. Encodes a learned state-to-action policy.

Also known as: Policy, Q-function, Action-value function

is trained bydeployed onis specialized byis trained byis specialized byReinforcement Learning Policy Learner: is trained byReinforcement Learning P…Learned-Policy Decision Engine: deployed onLearned-Policy Decision …Policy Network: is specialized byPolicy NetworkPolicy Learner: is trained byPolicy LearnerTabular Value Function: is specialized byTabular Value Function
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Policy NetworkChoose when states are high-dimensional (images, sensors) or actions are continuous, so tables cannot be stored or visited.
Tabular Value FunctionChoose when the state-action space is small and discrete enough to enumerate and visit.

Relationships

deployed on structural

is trained by lifecycle

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Q(s,a) action-value functionV(s) state-value functionBellman equation

Sources

  1. Ch5.10: T. Nguyen, "Utility-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.10. ISBN: 9798244538229.
  2. Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.