Part 3 — Evaluation & Optimization

12 chapters · 63.8 study hours allocated in the Study Plan · 11 slide decks · 35 videos · 35 code example files

On this page
  1. Chapters
  2. Chapter summaries
    1. 3.1A. Implement Evaluation Pipelines and Task Benchmarks - Introduction, Motivation, and Core Concepts
    2. 3.1B. Implement Evaluation Pipelines and Task Benchmarks - Custom Metrics and CI/CD Integration
    3. 3.1C. Implement Evaluation Pipelines and Task Benchmarks - Independent Practice and Comprehensive System Design
    4. 3.2. Compare Agent Performance Across Tasks and Datasets - Multi-Benchmark Evaluation and Statistical Rigor
    5. 3.3. Web Navigation and Interaction Benchmarks - Web Agent Evaluation and Multi-Hop Question Answering
    6. 3.4. Tune Parameters
    7. 3.5. Prompt Optimization, Few-Shot Learning, Fine-Tuning with Agent Trajectories, and Reward Modeling
    8. 3.6. Trace Analysis and Execution Debugging
    9. 3.7. Tool Auditing and Validation
    10. 3.8. Action Accuracy
    11. 3.9. Reasoning Quality
    12. 3.10. Efficiency Metrics
    13. Other code examples
    14. Additional worked examples
    15. Labs

Chapters

Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.

Ch. Title Hours Slides Quiz Videos Figures Code H-rated for
3.1A Implement Evaluation Pipelines and Task Benchmarks - Introduction, Motivation, and Core Concepts 3.0 PDF Quiz 0 12 — NV AWS GCP MS
3.1B Implement Evaluation Pipelines and Task Benchmarks - Custom Metrics and CI/CD Integration 1.4 PDF Quiz 3 12 — NV AWS GCP MS
3.1C Implement Evaluation Pipelines and Task Benchmarks - Independent Practice and Comprehensive System Design § — — Quiz 2 — — NV AWS GCP MS
3.2 Compare Agent Performance Across Tasks and Datasets - Multi-Benchmark Evaluation and Statistical Rigor 3.9 PDF Quiz 2 12 3 NV AWS GCP MS
3.3 Web Navigation and Interaction Benchmarks - Web Agent Evaluation and Multi-Hop Question Answering 8.1 PDF Quiz 4 12 — NV AWS GCP MS
3.4 Tune Parameters 3.8 PDF Quiz 1 12 — NV AWS GCP MS
3.5 Prompt Optimization, Few-Shot Learning, Fine-Tuning with Agent Trajectories, and Reward Modeling 6.4 PDF Quiz 5 11 — NV AWS GCP MS
3.6 Trace Analysis and Execution Debugging 9.2 PDF — 0 12 — NV AWS GCP MS
3.7 Tool Auditing and Validation 5.3 PDF Quiz 4 10 — NV AWS GCP MS
3.8 Action Accuracy 4.5 PDF Quiz 3 12 13 NV AWS GCP
3.9 Reasoning Quality 8.0 PDF Quiz 5 12 1 NV AWS GCP MS
3.10 Efficiency Metrics 10.2 PDF Quiz 6 12 — NV AWS GCP

Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_03_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.

† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).

‡ A combined deck that covers more than one chapter.

A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.

§ Has no section of its own in Study_Plan.md; the title comes from the Study Plan’s table of contents or a cross-reference there, or (9.16, 9.17) from the quiz list.

Chapter summaries

Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.

3.1A. Implement Evaluation Pipelines and Task Benchmarks - Introduction, Motivation, and Core Concepts

This foundational chapter establishes the vocabulary, conceptual frameworks, and architectural principles for systematic agent evaluation. It introduces the evaluation pyramid (unit tests, offline evaluation, staging, A/B testing), distinguishes offline evaluation as prediction from online evaluation as measurement, and demonstrates how continuous evaluation prevents regression from future changes that break previously functional capabilities.

3.1B. Implement Evaluation Pipelines and Task Benchmarks - Custom Metrics and CI/CD Integration

This guided practice chapter extends the foundational evaluation pipeline concepts from 3.1A with practical implementation of custom domain-specific metrics and continuous integration infrastructure. It demonstrates how to measure business value beyond generic accuracy and latency through keyword matching, LLM-as-judge scoring, and rule-based validation, then integrates these custom metrics into GitHub Actions workflows for automated quality assurance.

Videos (3)

3.1C. Implement Evaluation Pipelines and Task Benchmarks - Independent Practice and Comprehensive System Design

No summary in the Study Plan for this chapter.

Videos (2)

3.2. Compare Agent Performance Across Tasks and Datasets - Multi-Benchmark Evaluation and Statistical Rigor

This chapter extends single-metric evaluation to comprehensive multi-benchmark assessment, revealing capability distributions hidden by aggregate scoring. It addresses the critical dangers of benchmarking misconceptions, teaches controlled comparison methodology preventing confounding variables, and demonstrates how cross-dataset generalization testing exposes brittleness versus robust reasoning. Special focus on AgentBench’s eight-environment framework and the continuous feedback loop connecting offline evaluation to production deployment.

Videos (2)
Code examples (3 files)

3.3. Web Navigation and Interaction Benchmarks - Web Agent Evaluation and Multi-Hop Question Answering

This chapter specializes evaluation methodologies for web navigation agents and multi-hop reasoning tasks, addressing the unique challenges of evaluating agents in dynamic, interactive environments. It covers web agent benchmarks (Mind2Web, WebArena, Web Bench), metrics capturing critical intermediate actions, and the critical gap between offline static benchmarks and online production reality where agents encounter CAPTCHA, dynamic pricing, and authentication. Special emphasis on multi-hop question answering failure modes, handling non-determinism, and domain-specific benchmarking necessity.

Videos (4)
Intro to AI Safety, Remastered · Robert Miles AI Safety

3.4. Tune Parameters

Systematic parameter tuning requires understanding how configuration changes affect multiple performance dimensions simultaneously through accuracy-latency-cost trade-off spaces. This chapter establishes multi-objective optimization frameworks and Pareto frontier analysis for production agent deployment decisions, preventing single-metric optimization pathologies.

Videos (1)

3.5. Prompt Optimization, Few-Shot Learning, Fine-Tuning with Agent Trajectories, and Reward Modeling

Prompt optimization represents systematic engineering of agent instructions through measurement-driven refinement where minor phrasing changes dramatically shift accuracy (8-15 points). This chapter covers prompt optimization, few-shot learning leveraging 2-5 demonstration examples, trajectory-based fine-tuning combining human expertise with LLM-generated variants, and reward modeling through reinforcement learning from human feedback (RLHF).

Videos (5)

3.6. Trace Analysis and Execution Debugging

Trace analysis transforms opaque agent failures into actionable debugging insights through systematic instrumentation, visualization, and forensic analysis. This chapter establishes comprehensive frameworks for debugging non-deterministic probabilistic reasoning where traditional software debugging approaches prove inadequate, making invisible decision processes observable through detailed execution traces.

3.7. Tool Auditing and Validation

Tool auditing exposes what agents actually do through function calls and API invocations, complementing reasoning inspection with comprehensive monitoring of tool selection, parameter generation, and execution sequencing. This chapter establishes formal tool contracts, validation frameworks distinguishing syntactic from semantic errors, recovery mechanisms, and production monitoring patterns.

Videos (4)

3.8. Action Accuracy

Action accuracy represents granular evaluation of discrete decisions, tool selections, parameters, and execution steps complementing task-level metrics that only measure final outcomes. This chapter establishes frameworks distinguishing tool selection accuracy from parameter accuracy, trajectory evaluation metrics, and production monitoring patterns revealing hidden problems in agent behavior.

Videos (3)
Code examples (13 files)

3.9. Reasoning Quality

Reasoning quality evaluation examines how agents navigate decision-making paths through logical coherence, chain validity, transparency, and informativeness—dimensions distinct from task success rates. The chapter establishes frameworks for assessing multi-dimensional reasoning quality through chain-of-thought decomposition, formal logic verification, and evaluation metrics that enable production systems to catch flawed reasoning before critical failures.

Videos (5)
How I use LLMs · Andrej Karpathy
Code examples (1 files)

3.10. Efficiency Metrics

Efficiency metrics measure how effectively agents utilize computational resources, API calls, and execution steps, translating technical optimization into business-critical metrics. The chapter demonstrates that substantial efficiency improvements exist without accuracy sacrifices through systematic measurement of token consumption, step reduction, and cost attribution—critical for production viability.

Videos (6)

Other code examples

These files are named for a chapter number that is not in the current chapter list (an older numbering), so they are not attached to a chapter above.

18 files

Additional worked examples

From more_examples/part_03/:

Labs

No lab or legacy example exists for this Part yet. See Labs and Contributing.