Part 3 — Evaluation & Optimization
12 chapters · 63.8 study hours allocated in the Study Plan · 11 slide decks · 35 videos · 35 code example files
On this page
- Chapters
- Chapter summaries
- 3.1A. Implement Evaluation Pipelines and Task Benchmarks - Introduction, Motivation, and Core Concepts
- 3.1B. Implement Evaluation Pipelines and Task Benchmarks - Custom Metrics and CI/CD Integration
- 3.1C. Implement Evaluation Pipelines and Task Benchmarks - Independent Practice and Comprehensive System Design
- 3.2. Compare Agent Performance Across Tasks and Datasets - Multi-Benchmark Evaluation and Statistical Rigor
- 3.3. Web Navigation and Interaction Benchmarks - Web Agent Evaluation and Multi-Hop Question Answering
- 3.4. Tune Parameters
- 3.5. Prompt Optimization, Few-Shot Learning, Fine-Tuning with Agent Trajectories, and Reward Modeling
- 3.6. Trace Analysis and Execution Debugging
- 3.7. Tool Auditing and Validation
- 3.8. Action Accuracy
- 3.9. Reasoning Quality
- 3.10. Efficiency Metrics
- Other code examples
- Additional worked examples
- Labs
Chapters
Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.
Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_03_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.
† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).
‡ A combined deck that covers more than one chapter.
A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.
§ Has no section of its own in Study_Plan.md; the title comes from the Study Plan’s table of contents or a cross-reference there, or (9.16, 9.17) from the quiz list.
Chapter summaries
Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.
3.1A. Implement Evaluation Pipelines and Task Benchmarks - Introduction, Motivation, and Core Concepts
This foundational chapter establishes the vocabulary, conceptual frameworks, and architectural principles for systematic agent evaluation. It introduces the evaluation pyramid (unit tests, offline evaluation, staging, A/B testing), distinguishes offline evaluation as prediction from online evaluation as measurement, and demonstrates how continuous evaluation prevents regression from future changes that break previously functional capabilities.
3.1B. Implement Evaluation Pipelines and Task Benchmarks - Custom Metrics and CI/CD Integration
This guided practice chapter extends the foundational evaluation pipeline concepts from 3.1A with practical implementation of custom domain-specific metrics and continuous integration infrastructure. It demonstrates how to measure business value beyond generic accuracy and latency through keyword matching, LLM-as-judge scoring, and rule-based validation, then integrates these custom metrics into GitHub Actions workflows for automated quality assurance.
Videos (3)
3.1C. Implement Evaluation Pipelines and Task Benchmarks - Independent Practice and Comprehensive System Design
No summary in the Study Plan for this chapter.
Videos (2)
3.2. Compare Agent Performance Across Tasks and Datasets - Multi-Benchmark Evaluation and Statistical Rigor
This chapter extends single-metric evaluation to comprehensive multi-benchmark assessment, revealing capability distributions hidden by aggregate scoring. It addresses the critical dangers of benchmarking misconceptions, teaches controlled comparison methodology preventing confounding variables, and demonstrates how cross-dataset generalization testing exposes brittleness versus robust reasoning. Special focus on AgentBench’s eight-environment framework and the continuous feedback loop connecting offline evaluation to production deployment.
Videos (2)
Code examples (3 files)
3.3. Web Navigation and Interaction Benchmarks - Web Agent Evaluation and Multi-Hop Question Answering
This chapter specializes evaluation methodologies for web navigation agents and multi-hop reasoning tasks, addressing the unique challenges of evaluating agents in dynamic, interactive environments. It covers web agent benchmarks (Mind2Web, WebArena, Web Bench), metrics capturing critical intermediate actions, and the critical gap between offline static benchmarks and online production reality where agents encounter CAPTCHA, dynamic pricing, and authentication. Special emphasis on multi-hop question answering failure modes, handling non-determinism, and domain-specific benchmarking necessity.
Videos (4)
3.4. Tune Parameters
Systematic parameter tuning requires understanding how configuration changes affect multiple performance dimensions simultaneously through accuracy-latency-cost trade-off spaces. This chapter establishes multi-objective optimization frameworks and Pareto frontier analysis for production agent deployment decisions, preventing single-metric optimization pathologies.
Videos (1)
3.5. Prompt Optimization, Few-Shot Learning, Fine-Tuning with Agent Trajectories, and Reward Modeling
Prompt optimization represents systematic engineering of agent instructions through measurement-driven refinement where minor phrasing changes dramatically shift accuracy (8-15 points). This chapter covers prompt optimization, few-shot learning leveraging 2-5 demonstration examples, trajectory-based fine-tuning combining human expertise with LLM-generated variants, and reward modeling through reinforcement learning from human feedback (RLHF).
Videos (5)
3.6. Trace Analysis and Execution Debugging
Trace analysis transforms opaque agent failures into actionable debugging insights through systematic instrumentation, visualization, and forensic analysis. This chapter establishes comprehensive frameworks for debugging non-deterministic probabilistic reasoning where traditional software debugging approaches prove inadequate, making invisible decision processes observable through detailed execution traces.
3.7. Tool Auditing and Validation
Tool auditing exposes what agents actually do through function calls and API invocations, complementing reasoning inspection with comprehensive monitoring of tool selection, parameter generation, and execution sequencing. This chapter establishes formal tool contracts, validation frameworks distinguishing syntactic from semantic errors, recovery mechanisms, and production monitoring patterns.
Videos (4)
3.8. Action Accuracy
Action accuracy represents granular evaluation of discrete decisions, tool selections, parameters, and execution steps complementing task-level metrics that only measure final outcomes. This chapter establishes frameworks distinguishing tool selection accuracy from parameter accuracy, trajectory evaluation metrics, and production monitoring patterns revealing hidden problems in agent behavior.
Videos (3)
Code examples (13 files)
comprehensive_parameter_validation_code_05_comprehensive_parameter_validation.pyerror_recovery_evaluation_code_11_error_recovery_evaluation.pyhealthcare_action_safety_validation_code_13_healthcare_action_safety_validation.pyhealthcare_sequence_correctness_code_12_healthcare_sequence_correctness.pyparameter_correctness_calculation_code_08_parameter_correctness_calculation.pyparameter_source_enumeration_code_03_parameter_source_enumeration.pyparameter_validation_schema_code_02_parameter_validation_schema.pyreference_trajectory_definition_code_10_reference_trajectory_definition.pysecurity_constraints_validation_code_04_security_constraints_validation.pystep_utility_calculation_code_06_step_utility_calculation.pytest_case_refund_workflow_code_01_test_case_refund_workflow.pytool_execution_success_rate_code_09_tool_execution_success_rate.pytool_selection_accuracy_code_07_tool_selection_accuracy.py
3.9. Reasoning Quality
Reasoning quality evaluation examines how agents navigate decision-making paths through logical coherence, chain validity, transparency, and informativeness—dimensions distinct from task success rates. The chapter establishes frameworks for assessing multi-dimensional reasoning quality through chain-of-thought decomposition, formal logic verification, and evaluation metrics that enable production systems to catch flawed reasoning before critical failures.
Videos (5)
Code examples (1 files)
3.10. Efficiency Metrics
Efficiency metrics measure how effectively agents utilize computational resources, API calls, and execution steps, translating technical optimization into business-critical metrics. The chapter demonstrates that substantial efficiency improvements exist without accuracy sacrifices through systematic measurement of token consumption, step reduction, and cost attribution—critical for production viability.
Videos (6)
Other code examples
These files are named for a chapter number that is not in the current chapter list (an older numbering), so they are not attached to a chapter above.
18 files
Part_03_Chapter_3.1_01_load_and_validate_test_dataset.pyPart_03_Chapter_3.1_02_evaluate_agent_against_test_cases.pyPart_03_Chapter_3.1_03_compute_evaluation_metrics.pyPart_03_Chapter_3.1_04_log_results_to_mlflow.pyPart_03_Chapter_3.1_05_compare_to_baseline_with_statistical_testing.pyPart_03_Chapter_3.1_06_calculate_empathy_score.pyPart_03_Chapter_3.1_07_calculate_policy_compliance.pyPart_03_Chapter_3.1_08_calculate_tool_efficiency.pyPart_03_Chapter_3.1_09_custom_evaluation_pipeline_class.pyPart_03_Chapter_3.1_10_evaluate_agent_with_custom_metrics_integration.pyPart_03_Chapter_3.1_11_log_evaluation_with_custom_metrics.pyPart_03_Chapter_3.1_12_github_actions_workflow.yamlPart_03_Chapter_3.1_13_evaluation_execution_step.yamlPart_03_Chapter_3.1_14_run_evaluation_script.pyPart_03_Chapter_3.1_15_comment_formatting_function.jsPart_03_Chapter_3.1_16_quality_gate_check.pyPart_03_Chapter_3.1_17_threshold_configuration.pyPart_03_Chapter_3.1_18_travel_agent_evaluator.py
Additional worked examples
From more_examples/part_03/:
Labs
No lab or legacy example exists for this Part yet. See Labs and Contributing.
