Cognition · Model asset
Learned Decision Policy
Model assetCognitionCognition & MemoryVariation point (abstract)arc:LearnedPolicyModel
Learned parameters mapping states to actions, produced by reinforcement learning to maximize expected long-term reward.
Responsibility. Encodes a learned state-to-action policy.
Also known as: Policy, Q-function, Action-value function
Variants
| Variant | When to choose |
|---|---|
| Policy Network | Choose when states are high-dimensional (images, sensors) or actions are continuous, so tables cannot be stored or visited. |
| Tabular Value Function | Choose when the state-action space is small and discrete enough to enumerate and visit. |
Relationships
deployed on structural
is trained by lifecycle
Quantitative guidance
As stated by the sources; verify before use.
- With gamma = 0.9 a reward ten steps ahead counts 0.9^10 ~ 0.35 of an immediate reward (Ch5.12).
Classification
- Patterns
- Q(s,a) action-value functionV(s) state-value functionBellman equation
Sources
- Ch5.10: T. Nguyen, "Utility-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.10. ISBN: 9798244538229.
- Ch5.12: T. Nguyen, "Learning-Based Decision Making Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.12. ISBN: 9798244538229.