Observability & Evaluation · Software component
State Outcome Scorer
Software componentObservability & EvaluationObservability & Evaluationarc:StateOutcomeScorer
A response scorer that judges task success by comparing environment or database state before and after the agent acts, independent of the interaction path taken.
Responsibility. Scores functional task success from resulting environment state.
Also known as: Stateful evaluation, Functional correctness scorer, Outcome validation, End-to-end functional correctness check, Goal database state comparison, Final State Evaluator
Variant of Task Success Evaluator abstract
When to choose. Choose for transactional domains where success has clear state-verifiable criteria (e.g., refund issued, inventory updated).
Relationships
invokes dependency
reads dependency
evaluates assurance
alternative to variability
Design guidance
- SHOULD define goal states per task and tolerate conversation or navigation variation while assessing the outcome objectively.
Quantitative guidance
As stated by the sources; verify before use.
- tau-Bench pass^k measures reliability across k trials rather than single-attempt success (Ch3.2).
Classification
- Patterns
- Before/after state comparisonFunctional correctness (any valid path credited)pass^k reliability across repeated trialsOutcome-focused evaluationStateful evaluation
- Technologies
- tau-BenchWebArenatau-bench
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Penalizing valid alternative execution paths
Sources
- Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.