Part 5 — Advanced Reasoning & Decision Making

13 chapters · 70.4 study hours allocated in the Study Plan · 13 slide decks · 82 videos · 8 code example files

On this page
  1. Chapters
  2. Chapter summaries
    1. 5.1. Chain-of-Thought Reasoning
    2. 5.2. Tree-of-Thought (ToT)
    3. 5.3. Self-Consistency
    4. 5.4. Hierarchical Planning
    5. 5.5. Monte Carlo Tree Search (MCTS)
    6. 5.6. A* Search
    7. 5.7. Episodic Memory
    8. 5.8. Semantic Memory
    9. 5.9. Working Memory
    10. 5.10. Utility-Based Decision Making
    11. 5.11. Rule-Based Decision Making
    12. 5.12. Learning-Based Decision Making
    13. 5.13. Hybrid Decision Systems
    14. Additional worked examples
    15. Labs

Chapters

Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.

Ch. Title Hours Slides Quiz Videos Figures Code H-rated for
5.1 Chain-of-Thought Reasoning 3.2 PDF Quiz 7 12 — NV AWS GCP MS
5.2 Tree-of-Thought (ToT) 6.2 PDF Quiz 7 6 3 NV AWS GCP MS
5.3 Self-Consistency 3.9 PDF Quiz 7 8 — NV AWS GCP MS
5.4 Hierarchical Planning 4.4 PDF Quiz 2 6 — NV AWS GCP MS
5.5 Monte Carlo Tree Search (MCTS) 4.8 PDF Quiz 7 9 3 NV AWS DBX MS
5.6 A* Search 11.2 PDF Quiz 7 8 — NV AWS MS
5.7 Episodic Memory 3.7 PDF — 7 6 — NV AWS GCP MS
5.8 Semantic Memory 2.4 PDF Quiz 7 6 1 NV AWS GCP MS
5.9 Working Memory 4.9 PDF Quiz 6 11 — NV AWS MS
5.10 Utility-Based Decision Making 5.5 PDF Quiz 7 7 — NV AWS GCP MS
5.11 Rule-Based Decision Making 4.2 PDF — 0 5 — NV AWS DBX GCP MS
5.12 Learning-Based Decision Making 6.3 PDF Quiz 9 10 1 NV AWS DBX MS
5.13 Hybrid Decision Systems 9.7 PDF Quiz 9 6 — NV AWS DBX MS

Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_05_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.

† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).

‡ A combined deck that covers more than one chapter.

A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.

Chapter summaries

Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.

5.1. Chain-of-Thought Reasoning

This chapter covers Chain-of-Thought (CoT) as the foundational reasoning technique enabling agent memory, planning, and multi-agent coordination through structured intermediate reasoning steps. It explores CoT implementation approaches, integration with agent memory systems, task decomposition, grounding in observable reality, and advanced reasoning architectures.

Videos (7)

5.2. Tree-of-Thought (ToT)

Tree-of-Thought addresses a fundamental challenge in agent planning: how to intelligently explore decision spaces when linear reasoning produces irreversible choices and poor early decisions create cascading failures. ToT transforms reasoning into structured exploration through four integrated components—thought decomposition that identifies meaningful intermediate steps, candidate generation that explores alternatives, formal evaluation that prunes unproductive branches, and systematic search that enables backtracking—enabling agents to achieve 74% accuracy on Game of 24 versus Chain-of-Thought’s 4%.

Videos (7)
Code examples (3 files)

5.3. Self-Consistency

Self-Consistency fundamentally transforms how language models approach complex reasoning by separating generation of reasoning chains from selection of final answers, addressing Chain-of-Thought’s vulnerability to greedy decoding. By sampling multiple independent reasoning paths using stochastic decoding and aggregating through majority voting, Self-Consistency leverages the convergence property that correct answers emerge consistently across diverse solution strategies while errors scatter across samples, enabling 74% accuracy on GSM8K versus 58% baseline.

Videos (7)
Attention Is All You Need · Yannic Kilcher

5.4. Hierarchical Planning

Hierarchical planning addresses the impracticality of flat task sequences for complex problems by introducing multiple abstraction layers—strategic goals decompose into tactical phases, which decompose into operational actions. This multi-level organization mirrors human cognition, suppressing irrelevant details at higher levels while preserving decision quality, and transforms intractable problems with hundreds of interdependent tasks into manageable hierarchical structures where complex goal decomposition reduces search space from factorial to polynomial.

Videos (2)
3.13. AIPLAN - HTN Planning · Open Education Edinburgh

5.5. Monte Carlo Tree Search (MCTS)

Monte Carlo Tree Search addresses a fundamental challenge in agent planning: how to intelligently explore exponentially large action spaces without exhaustive enumeration. By iteratively building search trees through simulation—selecting promising branches, expanding to unexplored frontiers, simulating complete episodes, and backpropagating results—MCTS concentrates computational effort where it matters most. The algorithm combines chain-of-thought reasoning through iterative tree expansion, planning strategies for sequential decision-making, working memory storing visit counts and rewards, and stateful orchestration across simulation cycles.

Videos (7)
Code examples (3 files)

A* Search represents a fundamental breakthrough in intelligent pathfinding, combining actual costs already incurred with informed estimates of remaining distances to balance optimization with efficiency. By expanding nodes in order of lowest f-value where f(n) = g(n) + h(n), the algorithm integrates actual path costs (g(n)) with heuristic estimates of distance to goal (h(n)), enabling guaranteed optimal solutions through admissible heuristics while achieving computational efficiency through goal-directed guidance. A* powers applications from video game pathfinding to robot navigation to logistics optimization.

Videos (7)

5.7. Episodic Memory

Episodic memory stores specific past experiences with temporal and personal context, distinguishing it from semantic memory which stores generalized knowledge. It enables personalization by maintaining awareness of individual customer situations, preferences, and interaction history across multiple sessions through encoding, consolidation, and retrieval mechanisms.

Videos (7)

5.8. Semantic Memory

Semantic memory stores generalized facts, concepts, rules, and relationships independent of personal context or learning episodes, bridging the gap between parametric knowledge frozen at training time and dynamic external knowledge. It enables agents to access dynamic, current, domain-specific information through vector databases and knowledge graphs.

Videos (7)
Cosine Similarity, Clearly Explained!!! · StatQuest with Josh Starmer
Code examples (1 files)

5.9. Working Memory

Working memory is temporary storage and processing mechanism enabling real-time reasoning, context maintenance, and immediate decision-making within a single conversation episode. It implements the bounded context window fundamental to LLMs, managing competing demands including system prompts, conversation history, retrieved documents, and reasoning traces within token budgets.

Videos (6)
RAG From Scratch · LangChain

5.10. Utility-Based Decision Making

Utility-based decision making provides a mathematical framework evaluating actions by their expected utility rather than binary success criteria, enabling systematic reasoning about complex trade-offs. It uses expected utility theory to select actions maximizing average desirability across all possible futures while accounting for risk attitudes and probabilistic uncertainty.

Videos (7)
Introduction to Consumer Choice · Marginal Revolution University

5.11. Rule-Based Decision Making

Rule-based decision making represents an AI approach where agents make decisions by applying explicitly programmed conditional rules to current information, creating transparent and auditable reasoning chains. It enables deterministic behavior with complete explainability through inference traces, making it ideal for regulatory compliance and safety-critical applications.

5.12. Learning-Based Decision Making

This chapter establishes the fundamental paradigm shift from engineering explicit decision rules to cultivating intelligence through learning from consequences. Learning-based agents discover effective strategies through trial-and-error interaction with feedback signals that guide autonomous pattern discovery, enabling discovery of non-obvious strategies and continuous adaptation to novel situations that static rule-based systems cannot handle.

Videos (9)

Proximal Policy Optimization (PPO) Explained · the uploader has disabled embedding

Code examples (1 files)

5.13. Hybrid Decision Systems

Hybrid decision systems integrate multiple paradigms—utility-based optimization, rule-based logic, and learning-based policies—despite their fundamentally different internal representations. This chapter shows how to connect heterogeneous components through integration architectures, manage information loss at paradigm boundaries, and apply systematic decision frameworks for choosing appropriate paradigm combinations based on problem characteristics.

Videos (9)
Project Chimera - Demo · Aytuğ Akarlar
SAT Net · Neuro Symbolic

Stanford CS229 Machine Learning · the uploader has disabled embedding

Additional worked examples

From more_examples/part_05/:

Labs

No lab or legacy example exists for this Part yet. See Labs and Contributing.