Observability & Evaluation · Software component
Task Success Evaluator
Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:TaskSuccessEvaluator
An abstract evaluation component that judges whether, and how well, an agent run accomplished its task, producing success or partial-credit scores.
Responsibility. Decides task success for an agent run.
Also known as: Task completion scorer, Success rate evaluation, Goal achievement rate, Task completion rate tracker, Goal achievement evaluator
Variant of Response Scorer abstract
Variants
| Variant | When to choose |
|---|---|
| Milestone Evaluator | Choose for complex tasks admitting partial credit, where locating the failing intermediate step guides optimization. |
| State Outcome Scorer | Choose for transactional domains where success has clear state-verifiable criteria (e.g., refund issued, inventory updated). |
| Trajectory Matching Evaluator | Choose only when a task has a single valid execution path; the text warns it penalizes valid alternative routes of stochastic agents. |
Relationships
is invoked by dependency
reads dependency
emits telemetry to dynamic
evaluates assurance
- Agent Controller abstract Ch3.3
- MCTS Planner Ch5.5
Design guidance
- SHOULD validate outcomes rather than match trajectories when tasks admit multiple valid execution paths.
- SHOULD measure success distributions across repeated trials rather than point estimates.
- SHOULD score tool selection and parameter accuracy, since conceptually correct calls can fail on parameter naming mismatches.
- SHOULD log task success separately from shaped reward; rising reward with stagnant success indicates misalignment.
- MUST define successful completion as outcome success combined with quality/satisfaction assessment, not technical closure.
- SHOULD classify outcomes as successful, partial or failed completion: rising partials indicate capability gaps (missing tools/knowledge); rising failures indicate upstream task classification or routing problems.
- SHOULD NOT be maximised in isolation; completion must be read alongside satisfaction to expose premature closure.
Quantitative guidance
As stated by the sources; verify before use.
- Illustrative: 80% pass@1 may coexist with 95% pass@3 but only 50% pass^3 (Ch3.3).
- Targets: successful >70%, partial 10-20% acceptable, failed <10% (Ch8.4).
- Achievable completion by domain: objective retrieval 80-95%; customer service 70-85%; complex multi-step reasoning such as software debugging 20-40% (Ch8.4).
- Production success-rate target >95%; also track steps to completion and retry rate (Ref8.03).
Classification
- Patterns
- Success Rate (SR) / Task Goal Completionpass@k and pass^k across independent trialsAction advancement scoringTool selection and parameter accuracy
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- False confidence about production readiness
Sources
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch8.4: T. Nguyen, "Success Metrics and Multi-Dimensional Measurement," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.4. ISBN: 9798244538229.
- Ref1.01: NVIDIA, "NVIDIA NeMo Agent Toolkit overview," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 26, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html
- Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note