Observability & Evaluation · Software component

Task Success Evaluator

Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:TaskSuccessEvaluator

An abstract evaluation component that judges whether, and how well, an agent run accomplished its task, producing success or partial-credit scores.

Responsibility. Decides task success for an agent run.

Also known as: Task completion scorer, Success rate evaluation, Goal achievement rate, Task completion rate tracker, Goal achievement evaluator

Variant of Response Scorer abstract

evaluatesis invoked byemits telemetry toreadsevaluatesspecializesis specialized byis specialized byis specialized byAgent Controller: evaluatesAgent ControllerEvaluation Harness: is invoked byEvaluation HarnessMetrics Collector: emits telemetry toMetrics CollectorEvaluation Dataset: readsEvaluation DatasetMCTS Planner: evaluatesMCTS PlannerResponse Scorer: specializesResponse ScorerTrajectory Matching Evaluator: is specialized byTrajectory Matching Eval…State Outcome Scorer: is specialized byState Outcome ScorerMilestone Evaluator: is specialized byMilestone Evaluator
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Milestone EvaluatorChoose for complex tasks admitting partial credit, where locating the failing intermediate step guides optimization.
State Outcome ScorerChoose for transactional domains where success has clear state-verifiable criteria (e.g., refund issued, inventory updated).
Trajectory Matching EvaluatorChoose only when a task has a single valid execution path; the text warns it penalizes valid alternative routes of stochastic agents.

Relationships

is invoked by dependency

reads dependency

emits telemetry to dynamic

evaluates assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Success Rate (SR) / Task Goal Completionpass@k and pass^k across independent trialsAction advancement scoringTool selection and parameter accuracy
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
False confidence about production readiness

Sources

  1. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
  2. Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
  3. Ch8.4: T. Nguyen, "Success Metrics and Multi-Dimensional Measurement," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.4. ISBN: 9798244538229.
  4. Ref1.01: NVIDIA, "NVIDIA NeMo Agent Toolkit overview," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 26, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html
  5. Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  6. Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note