Observability & Evaluation · Software component

State Outcome Scorer

Software componentObservability & EvaluationObservability & Evaluationarc:StateOutcomeScorer

A response scorer that judges task success by comparing environment or database state before and after the agent acts, independent of the interaction path taken.

Responsibility. Scores functional task success from resulting environment state.

Also known as: Stateful evaluation, Functional correctness scorer, Outcome validation, End-to-end functional correctness check, Goal database state comparison, Final State Evaluator

Variant of Task Success Evaluator abstract

When to choose. Choose for transactional domains where success has clear state-verifiable criteria (e.g., refund issued, inventory updated).

is target of alternativeToevaluatesalternative tospecializesinvokesis target of alternativeToinvokesreadsLLM Judge: is target of alternativeToLLM JudgeWeb Navigation Agent: evaluatesWeb Navigation AgentTrajectory Matching Evaluator: alternative toTrajectory Matching Eval…Task Success Evaluator: specializesTask Success EvaluatorBenchmark Environment: invokesBenchmark EnvironmentMilestone Evaluator: is target of alternativeToMilestone EvaluatorSimulated Web Environment: invokesSimulated Web EnvironmentEnvironment Snapshot: readsEnvironment Snapshot
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

reads dependency

evaluates assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Before/after state comparisonFunctional correctness (any valid path credited)pass^k reliability across repeated trialsOutcome-focused evaluationStateful evaluation
Technologies
tau-BenchWebArenatau-bench
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Penalizing valid alternative execution paths

Sources

  1. Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.
  2. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.